Key Takeaways
- GPU density: All four vendors ship 8-GPU SXM5 flagship platforms (Dell XE9680, HPE Cray XD670, Supermicro SYS-821GE-TNHR); Lenovo and HPE also sell lower-density, lower-cost PCIe platforms.
- Price band: An 8x H100/H200 SXM5 server runs roughly $200,000-$400,000+ depending on GPU memory (80GB H100 vs. 141GB H200) and support tier.
- Cooling is the real constraint. Two or three 8-GPU SXM5 servers already exceed the ~20-40kW practical ceiling of air cooling per rack — liquid cooling becomes mandatory, not optional, above that density.
- Networking matters for training, less for inference. Multi-node training clusters need InfiniBand NDR or RoCE v2 Ethernet fabric; single-node inference does not.
- Match GPU to workload, not budget alone. L40S (48GB, PCIe, air-cooled) fits inference; H100/H200 (80-141GB, SXM5, NVLink) fits training and large-batch serving.
- Buyers who need one or two GPUs of inference capacity should not buy a rack server at all — a workstation build is the right tier for that scale.
📍 In One Sentence
Enterprise GPU server buyers should match vendor to workload — Dell PowerEdge XE9680, HPE Cray XD670, and Supermicro SYS-821GE-TNHR compete directly on 8x H100/H200 SXM5 training platforms at $200,000-$400,000+, while HPE ProLiant DL380a Gen11 and Lenovo ThinkSystem SR675 V3 with PCIe L40S GPUs serve inference-only budgets at roughly a third to half the price.
💬 In Plain Terms
Buying one AI server for a company is a different purchase than building one gaming-style PC. These machines cost as much as a house, need a plan for cooling and electricity before they arrive, and come with a multi-year support contract. This guide compares the four companies that actually sell these rack-mounted AI servers — Dell, Lenovo, HPE, and Supermicro — so a buyer can pick the right size and vendor for the actual job (running AI models for employees vs. training new ones from scratch), not just the biggest spec sheet.
Quick Facts
- Dell PowerEdge XE9680: 6U, up to 8x H100/H200 SXM5, dual Intel Xeon Platinum, up to 4TB DDR5.
- Lenovo ThinkSystem SR675 V3: 3U, up to 8x GPU (H100/H200/L40S mix), dual AMD EPYC 9004/9005, up to 6TB DDR5, optional Neptune liquid cooling.
- HPE Cray XD670: 5U, up to 8x H100/H200 SXM5, dual Intel Xeon 4th Gen, InfiniBand NDR / Slingshot 11 / Ethernet fabric options.
- HPE ProLiant DL380a Gen11: 2U, up to 4 double-wide or 8 single-wide GPUs (L40S/H100 PCIe), dual Intel Xeon up to 64 cores.
- Supermicro SYS-821GE-TNHR: 8U, up to 8x H100/H200 SXM5, dual Intel Xeon 5th/4th Gen, up to 8TB DDR5.
- GPU VRAM: H100 = 80GB HBM3; H200 = 141GB HBM3e; L40S = 48GB GDDR6.
- Rack power: a single 8x H100/H200 SXM5 node draws roughly 10-12kW at full load — two of them already exceed the practical air-cooling ceiling for one rack.
Which GPU Server Should You Buy for Your Use Case?
The right vendor depends on workload and datacenter readiness, not brand preference. Training clusters, inference-only deployments, and liquid-cooling-ready facilities each point to a different platform.
- Best for large-scale multi-node training: Dell PowerEdge XE9680 — the broadest enterprise sales and system-integrator network for building out an InfiniBand-connected multi-rack cluster.
- Best for a liquid-cooling-ready datacenter: Lenovo ThinkSystem SR675 V3 — Neptune direct-to-chip liquid cooling cuts cooling cost materially at sustained high GPU utilization, with an AMD EPYC CPU option.
- Best for inference-only or a smaller first purchase: HPE ProLiant DL380a Gen11 — 2U, air-cooled, PCIe GPUs, roughly a third to half the cost of an 8x SXM5 training box.
- Best for configuration flexibility and price competitiveness: Supermicro SYS-821GE-TNHR — the widest build-to-order options via the system-integrator channel, often the most price-competitive route to an 8x H100/H200 SXM5 configuration.
- Skip rack servers entirely if: your actual need is 1-2 GPUs of inference capacity for a small team — a local LLM workstation build costs a fraction of the price and needs no datacenter cooling plan.
How Do the Four Vendors Compare on Specs and Price?
Dell, Lenovo, HPE, and Supermicro each field at least one 8-GPU SXM5 platform, but they diverge sharply on form factor, cooling options, and channel.
| Vendor / Model | Form Factor | Max GPUs | GPU Options | Cooling | Price Band |
|---|---|---|---|---|---|
| Dell PowerEdge XE9680 | 6U | 8 (SXM5, NVLink) | H100 80GB / H200 141GB | Air standard, liquid optional | ~$200K-$375K (8x GPU) |
| Lenovo SR675 V3 | 3U | 8 (PCIe or SXM) | H100 / H200 / L40S | Air, or Neptune liquid | Varies widely by config |
| HPE Cray XD670 | 5U | 8 (SXM5, NVLink) | H100 80GB / H200 141GB | Air standard | Quote-only |
| HPE ProLiant DL380a Gen11 | 2U | 4 double-wide / 8 single | L40S / H100 PCIe | Air only | Lower — quote-only |
| Supermicro SYS-821GE-TNHR | 8U | 8 (SXM5, NVLink) | H100 80GB / H200 141GB | Air standard | ~$200K-$320K (8x GPU) |
What Does the Dell PowerEdge XE9680 Offer?
The Dell PowerEdge XE9680 is a 6U rack server holding 8 NVIDIA HGX H100 or H200 SXM5 GPUs connected via NVLink, built specifically for large-model training and inference. It pairs the GPUs with two 4th- or 5th-generation Intel Xeon Scalable processors (up to 56 cores each), up to 32 DDR5 DIMM slots (4TB max, 4800 MT/s), and 10 PCIe Gen5 x16 slots for networking and storage expansion.
It ships standard with air cooling; Dell offers liquid-cooling options for datacenters running above the ~20-30kW-per-rack density where air cooling stops being practical.
Pricing is not published — reseller and system-integrator quotes for a fully configured 8x H100/H200 unit have ranged roughly $200,000-$375,000 depending on GPU memory tier, RAM, storage, and support level. Get a formal quote through Dell.com before budgeting.
What Does the Lenovo ThinkSystem SR675 V3 Offer?
The Lenovo ThinkSystem SR675 V3 is a 3U rack server supporting up to 8 double-wide or single-wide GPUs — including NVIDIA H100, H200, and L40S — paired with two 5th-generation AMD EPYC 9004/9005 processors and up to 6TB of DDR5-4800 memory across 24 DIMM slots.
The distinguishing feature is Lenovo Neptune, a direct-to-chip and hybrid liquid-to-air cooling system available as a configuration option — relevant for buyers whose facility already runs liquid cooling loops or plans to add one, since it materially reduces the cooling cost of sustained high-utilization GPU workloads versus air alone.
The SR675 V3 also supports mixed GPU configurations (H200 4-GPU NVLink builds, or L40S for inference-focused deployments), making it the most configuration-flexible platform in this comparison for buyers who want one chassis family to cover both training and inference tiers. Configure via Lenovo.com.
What Do the HPE Cray XD670 and ProLiant DL380a Gen11 Offer?
HPE sells two distinct GPU server tiers: the Cray XD670 for large-scale training, and the ProLiant DL380a Gen11 for inference and smaller deployments.
The Cray XD670 is a 5U chassis holding 8x NVIDIA H100 or H200 SXM5 GPUs with dual 4th-generation Intel Xeon processors. Its distinguishing feature is fabric choice: 8x PCIe Gen5 half-height slots supporting HPE Slingshot 11, InfiniBand NDR, or standard Ethernet — relevant for buyers already standardized on Slingshot from an existing HPE Cray supercomputing footprint.
The ProLiant DL380a Gen11 is a 2U server supporting 4 double-wide or 8 single-wide GPUs (L40S or H100 PCIe), up to 3TB DDR5, and PCIe 5.0 — the air-cooled, lower-density option for inference workloads or a first GPU purchase that doesn't justify an 8-GPU SXM5 platform. See both at HPE.com.
What Does the Supermicro SYS-821GE-TNHR Offer?
The Supermicro SYS-821GE-TNHR is an 8U rack server supporting up to 8 NVIDIA HGX H100 (80GB) or HGX H200 (141GB) GPUs, dual 4th- or 5th-generation Intel Xeon Scalable processors, and up to 8TB of DDR5-5600 memory across 32 DIMM slots — the highest maximum RAM capacity of the four platforms compared here.
Supermicro sells primarily through a system-integrator and reseller channel rather than direct enterprise account teams, which typically means more build-to-order flexibility (drive bays, networking cards, power supply redundancy) and, per current reseller listings, a competitive starting price for an 8x H100 configuration — starting-configuration quotes have ranged roughly $200,000-$320,000, with fully loaded configurations running higher.
Storage is a strength: up to 19 hot-swap 2.5" NVMe/SATA/SAS bays plus 2 M.2 slots, useful for buyers running large local datasets alongside inference or fine-tuning. See the base configuration at Supermicro.com.
Should You Buy H100, H200, or L40S GPUs?
H200 wins on memory and bandwidth, H100 is the more available and often cheaper SXM5 option, and L40S is the air-cooled, PCIe choice for inference-only deployments. All three are current NVIDIA datacenter GPUs as of September 2026; none is being discontinued imminently, so the choice is workload fit, not obsolescence risk.
- Choose H200 if: you're serving large-context workloads or training larger models where 80GB per GPU forces cross-GPU sharding you'd rather avoid.
- Choose H100 if: you need the widest vendor and reseller availability for a multi-node NVLink/InfiniBand training cluster and 80GB per GPU is enough for your model size.
- Choose L40S if: the workload is inference-only, the datacenter is air-cooled only, and 48GB per GPU covers your largest model at the quantization level you plan to run.
| GPU | VRAM | Memory Bandwidth | Best For |
|---|---|---|---|
| NVIDIA H100 SXM5 | 80GB HBM3 | 3.35 TB/s | Multi-node training, widest availability |
| NVIDIA H200 SXM | 141GB HBM3e | 4.8 TB/s | Large-context serving, bigger batch training |
| NVIDIA L40S | 48GB GDDR6 | 864 GB/s | Air-cooled PCIe inference, lower budget |
How Much Power and Cooling Does a GPU-Dense Rack Need?
A single 8-GPU H100/H200 SXM5 server draws roughly 10-12kW at full load — 8x 700W GPUs alone account for 5.6kW before CPUs, memory, and fans. Two or three of those servers in one rack already push past the practical ceiling of air cooling.
Industry figures put air cooling's practical limit at roughly 20-40kW per rack; liquid cooling (direct-to-chip or immersion) is needed above that, and can support 100-200kW+ per rack. For reference, NVIDIA's GB200 NVL72 rack-scale system draws roughly 120-130kW total — a data point on where AI rack density is heading, not a spec of any server compared here.
Practical implication for buyers: if you're racking more than one or two 8-GPU SXM5 servers per rack, plan for direct-to-chip liquid cooling (Lenovo Neptune, or a facility-level liquid loop) rather than assuming standard datacenter air handling will keep up.
Do You Need InfiniBand or Standard Ethernet?
Single-node inference deployments do not need InfiniBand — standard 100/200GbE Ethernet is enough. Multi-node training or fine-tuning clusters, where GPUs across servers need to synchronize gradients constantly, do need a dedicated high-bandwidth, low-latency fabric.
The two options for that fabric are InfiniBand NDR (400Gb/s per link, the traditional HPC choice, one NIC per GPU on flagship platforms) and RoCE v2 (RDMA over Converged Ethernet — e.g. NVIDIA Spectrum-X — which delivers similar throughput over a standard Ethernet fabric your network team may already operate).
- Use InfiniBand NDR if: you're building a dedicated multi-node training cluster and want the most mature, widely deployed RDMA fabric for that scale.
- Use RoCE v2 (Ethernet) if: your team already operates a converged Ethernet network and wants to avoid maintaining a separate InfiniBand fabric and skill set.
- Skip both if: you're running single-node inference — standard networking is sufficient and the extra fabric cost isn't justified.
How Do You Size a GPU Server Purchase Against Your Workload?
Size the purchase to the workload, not the biggest available configuration. Inference-only deployments and training/fine-tuning deployments have fundamentally different GPU, memory, and networking requirements.
- 1Classify the workload first
Why it matters: Inference-only (serving a fixed model to users) needs far less GPU memory and no multi-node fabric compared to training or fine-tuning, which needs to hold gradients and optimizer state in addition to model weights. - 2Estimate GPU memory need from model size and quantization
Why it matters: A 70B-parameter model at FP16 needs roughly 140GB of VRAM before overhead — that alone rules out a single-GPU L40S (48GB) and points toward multi-GPU H100/H200 sharding or a smaller/quantized model. - 3Decide single-node vs. multi-node
Why it matters: If one 8-GPU server's combined VRAM covers the model and concurrency target, skip InfiniBand/RoCE entirely and save the fabric cost; if not, budget for a dedicated networking fabric from the start. - 4Match cooling to rack density before ordering
Why it matters: Confirm with facilities whether the target rack can support liquid cooling before committing to more than one or two 8-GPU SXM5 servers per rack — retrofitting cooling after delivery is far more expensive than planning for it upfront. - 5Get a formal quote and confirm lead time
Why it matters: None of these vendors publish list pricing for 8-GPU configurations, and delivery lead times for GPU-dense servers have run several weeks to a few months depending on GPU allocation — budget the timeline, not just the price.
What Warranty and Support Tiers Should You Choose?
All four vendors offer tiered enterprise support beyond the base hardware warranty, but the tier names, response times, and included services differ — confirm current terms directly with the vendor before purchase, since programs change.
- Dell sells its ProSupport tiers (including options with faster, mission-critical response) alongside the XE9680 — ask specifically about GPU-server-qualified support, not the standard PowerEdge tier.
- Lenovo sells Premier Support tiers for the ThinkSystem line, with options for on-site response and proactive monitoring.
- HPE sells support through Pointnext Complete Care and offers HPE GreenLake as a consumption-based (pay-per-use) alternative to a capital purchase for buyers who want to avoid the six-figure upfront cost.
- Supermicro support terms vary more by reseller/system-integrator than the other three vendors, since much of its volume moves through that channel rather than direct enterprise sales — get the support terms in writing from your specific reseller, not just the base Supermicro warranty page.
- For any vendor: ask what happens on GPU failure specifically (replacement SLA, whether it requires shipping the whole node or just the GPU tray) — this is the failure mode most likely to actually happen on a GPU-dense server.
What Mistakes Do Enterprise Buyers Make?
- Buying 8-GPU SXM5 capacity for an inference-only workload. If you're only serving a fixed model to users, a 2U PCIe platform like the ProLiant DL380a Gen11 covers it at a fraction of the price and complexity.
- Ordering before confirming rack cooling capacity. A second or third 8-GPU SXM5 server in the same rack can push past air cooling's practical ceiling — confirm with facilities before the hardware arrives, not after.
- Skipping the networking fabric budget for a "we might scale later" cluster. Retrofitting InfiniBand or RoCE onto an already-deployed single-node fleet is more disruptive than budgeting for it in the original purchase.
- Treating the sticker price as the total cost. Support contracts, networking fabric, cooling retrofit, and power infrastructure upgrades routinely add 15-30% on top of the server hardware line item.
- Assuming H200 is always the right upgrade over H100. If your model and batch size fit comfortably in 80GB per GPU, H200's extra memory and cost buy you nothing — check actual VRAM need before paying the premium.
Frequently Asked Questions
How many H100 or H200 GPUs do I need for enterprise inference vs. training?
Inference-only deployments serving a single fixed model to a moderate number of concurrent users often fit on 1-4 GPUs and don't need an 8-GPU SXM5 platform at all. Training or fine-tuning large models (70B+ parameters) typically needs the full 8-GPU NVLink configuration to hold the model, gradients, and optimizer state across GPUs. Size the purchase from the workload, not a default of "buy the biggest platform."
What is the real total cost of an 8-GPU H100 rack server?
Reseller and system-integrator quotes for a fully configured 8x H100 unit have ranged roughly $200,000-$375,000 for the hardware alone, before support contracts, networking fabric, and any cooling infrastructure upgrade — those typically add another 15-30% on top. None of the four vendors publish list pricing; get a formal quote before budgeting.
Do I need liquid cooling for a GPU-dense rack?
If you're racking more than one or two 8-GPU H100/H200 SXM5 servers per rack, yes — each one draws roughly 10-12kW at full load, and air cooling's practical ceiling is around 20-40kW per rack. Below that density, standard air cooling can still work; confirm with facilities before ordering hardware.
Should I choose InfiniBand or Ethernet (RoCE) for networking?
For single-node inference, neither — standard Ethernet is enough. For multi-node training clusters, InfiniBand NDR is the more mature, widely deployed high-bandwidth RDMA fabric; RoCE v2 over Ethernet is the alternative if your network team wants to avoid running a separate InfiniBand fabric and skill set.
Dell vs. Lenovo vs. HPE vs. Supermicro — which vendor has the best enterprise support?
Dell, Lenovo, and HPE each sell tiered enterprise support (ProSupport, Premier Support, and Pointnext Complete Care respectively) through direct account teams. Supermicro sells primarily through system integrators and resellers, so support terms vary more by reseller than by a single Supermicro-wide tier — get support terms in writing from the specific reseller before purchase.
H100 vs. H200 vs. L40S — which GPU should I buy?
Choose H200 (141GB HBM3e) for large-context serving or training where 80GB per GPU forces cross-GPU sharding you'd rather avoid. Choose H100 (80GB HBM3) for the widest vendor and reseller availability in a multi-node training cluster where 80GB is enough. Choose L40S (48GB GDDR6, PCIe, air-cooled) for inference-only deployments on a lower budget.
Can I mix GPU models within the same rack or server?
Within a single server, no — an 8-GPU SXM5 platform is built and NVLink-connected around one GPU model (all H100 or all H200), not a mix. Within a rack, yes — you can run one server configured with H100/H200 for training next to another configured with L40S for inference, as long as each server's own cooling and power draw is accounted for separately.
What warranty and support tier should enterprise buyers choose?
At minimum, confirm the vendor's replacement SLA specifically for GPU failure (not just general hardware failure) — GPU failure is the most likely failure mode on a GPU-dense server, and replacement logistics (shipping a GPU tray vs. the whole node) vary by vendor and tier. Match the response-time tier to how much downtime the workload can actually tolerate; a mission-critical inference service needs faster response than a batch training job that can wait a day.
Is on-premises hardware cheaper than cloud GPU rental at enterprise scale?
It depends on utilization, not just sticker price — on-premises hardware has a high upfront cost but a low per-hour cost once running, while cloud rental has no upfront cost but a much higher per-hour rate. The crossover point is typically sustained, near-constant utilization; occasional or bursty workloads usually cost less to rent. See our cloud GPU rental guide for a detailed cost comparison.
How long does delivery take for an 8x H100 or H200 configuration?
Lead times for GPU-dense servers have varied from several weeks to a few months depending on GPU allocation and current demand — this is not a stock item most vendors keep on the shelf in an 8-GPU configuration. Confirm lead time as part of the formal quote, and budget the project timeline around it, not just the price.
Sources
- Dell PowerEdge XE9680 product page -- dell.com/en-us/shop/ipovw/poweredge-xe9680
- Lenovo ThinkSystem SR675 V3 Product Guide -- lenovopress.lenovo.com/lp1611-thinksystem-sr675-v3-server
- HPE Cray XD670 QuickSpecs -- hpe.com/us/en/hpe-cray-xd670.html
- HPE ProLiant DL380a Gen11 datasheet -- hpe.com/us/en/compute/hpe-proliant-compute/dl380a-gen11.html
- Supermicro SYS-821GE-TNHR datasheet -- supermicro.com/en/products/system/datasheet/SYS-821GE-TNHR
- NVIDIA H100/H200 Tensor Core GPU specifications -- nvidia.com