Advertisement
Data-Center GPU

NVIDIA RTX PRO 6000 Blackwell Server Edition

A passively cooled, 96GB Blackwell powerhouse built for dense AI-inference and rendering racks.

88/ 100
Excellent
Android Hire score
~$13,250
NVIDIA raised list pricing to $13,250 in 2026, up 55% from the $8,565 launch MSRP. Street runs roughly $11,400-$15,000 depending on channel.
96GB GDDR7 ECC
Memory
4 PFLOPS
AI (FP4)
Up to 600 W
Power
~1.6 TB/s
Bandwidth
NVIDIA RTX PRO 6000 Blackwell Server Edition

Image: NVIDIA

Best for

AI inference, rendering and virtualization in validated servers.

Standout

MIG, vGPU and Confidential Computing on one passive card.

Watch out

No NVLink, and it needs a server with strong front-to-back airflow.

By Aditya Singh

The NVIDIA RTX PRO 6000 Blackwell Server Edition is the data-center member of NVIDIA's RTX PRO 6000 Blackwell family (which also includes the Workstation and Max-Q editions). It pairs the full-fat Blackwell GB202 GPU with 96 GB of GDDR7 ECC memory in a passively cooled, server-optimized card designed to slot into validated OEM racks from Dell, HPE, Lenovo, Supermicro, ASUS and Cisco. It is, in effect, the modern successor to NVIDIA's L40 / L40S / A40 line.

Quick verdict: An outstanding universal data-center GPU for AI inference, rendering, simulation and virtualization. Its 96GB pool, FP4 throughput and MIG/vGPU/Confidential Computing make it a clear, much faster L40S replacement. It is not a large-scale training replacement for HBM-based HGX parts (no NVLink, and GDDR7 bandwidth is lower than HBM). Score: 88/100.

Jump to: Full specs · Price & where to buy · Server vs Workstation vs Max-Q · vs L40S, H100 & L4 · Performance · Power & cooling · Supported servers · MIG & vGPU

RTX PRO 6000 Blackwell Server Edition — Full Specifications

Every figure below is cross-checked against NVIDIA's official Server Edition product page and the Lenovo Press ThinkSystem product guide. Items NVIDIA does not officially publish for this SKU are explicitly flagged.

SpecificationDetail
GPU & Architecture
GPU chipNVIDIA GB202 (Blackwell)
Process nodeTSMC 4N-class (4nm) — flag: NVIDIA doesn't print an exact node
Streaming multiprocessors188 SMs
CUDA cores24,064
RT cores188 (4th gen) — 355 TFLOPS peak RT performance
Tensor cores752 (5th gen, FP4-capable)
Boost clock~2,617 MHz (third-party estimate — NVIDIA does not publish clocks for this SKU)
Memory
Memory size96 GB GDDR7 with ECC (clamshell configuration)
Memory bus512-bit
Memory bandwidthUp to 1,597 GB/s (~1.6 TB/s) — lower than Workstation/Max-Q
AI & Compute (Tensor)
FP4 Tensor (peak)4 PFLOPS
FP8 Tensor2 PFLOPS
FP16 / BF16 Tensor1 PFLOP
TF32 Tensor234 TFLOPS
FP32 (single precision)120 TFLOPS
Power & Cooling
Total board powerConfigurable, up to 600 W
CoolingPassive (server front-to-back airflow); single-slot FHXL liquid-cooled SKU also offered
Power connector16-pin (2x6+4 / 12V-2x6); cable kit varies by server vendor
Operating inlet temperature0°C to 50°C (storage −40°C to 75°C)
Form Factor & Connectivity
Form factorDual-slot, Full-Height Full-Length, 4.4″ H × 10.5″ L (air-cooled)
InterfacePCIe Gen 5 x16
Display outputs4× DisplayPort 2.1b — disabled by default (effectively headless)
NVLinkNot supported (PCIe peer-to-peer only)
Media engine4× NVENC (9th gen), 4× NVDEC (6th gen), 4× JPEG decoders, AV1 encode/decode
Enterprise Features
MIG (Multi-Instance GPU)Up to 4 isolated instances of 24 GB each
vGPUYes — NVIDIA vPC/vApps and RTX Virtual Workstation (vWS)
Confidential ComputingYes — hardware TEE for secure AI
Secure boot / ECCSecure boot with root of trust; ECC on all 96 GB
Availability & Price
LaunchAnnounced at GTC on March 18, 2025; broad OEM server availability from 2H 2025
Launch MSRP~$8,565
Price (July 2026)~$13,250 NVIDIA list; street ~$11,400–$15,000
OEM part numbersNVIDIA 900-2G153-0000-000; Lenovo 4X67B09287; HPE S6A73C

Price and Where to Buy (July 2026)

This is the single biggest change since launch, and the reason so many buyers land on this page. The Server Edition debuted at roughly $8,565 in March 2025. It does not sell anywhere near that today. Tom's Hardware reported that NVIDIA lifted list pricing to $13,250, a roughly 55% increase in about 16 months, driven overwhelmingly by the ongoing GDDR7 memory shortage — a 96 GB clamshell board is the most memory-exposed discrete GPU on the market.

ChannelPrice (July 2026)Notes
NVIDIA list / marketplace~$13,250Reference price after the 2026 increase
PNY (add-in-card partner)~$11,360Lowest widely quoted partner price
Newegg (retail)~$12,099Stock fluctuates heavily
NeweggBusiness (marketplace seller)~$14,989In stock, 6–8 day ship at time of writing
B&H Photo~$13,349+Frequently backordered
CDW / OEM quote (Dell, HPE, Lenovo)Quote-basedUsually bundled into a validated server BOM

Buying advice: if you need one or two cards for an existing chassis, retail/reseller channels are fine — but confirm the GPU is on your server's validated options list and that the vendor ships the correct 16-pin auxiliary power cable kit for your platform. For four or more cards, go through the OEM (Dell PowerEdge, HPE ProLiant, Lenovo ThinkSystem, Supermicro), where the card arrives pre-validated with the right airflow, riser and power configuration.

Rent before you buy. At $13k+ per card and volatile street pricing, cloud rental is the sane way to benchmark your workload first. The RTX PRO 6000 Server Edition is available on AWS EC2 G7e (from roughly $3.36/hr for a single GPU, scaling to ~$33/hr for 8×), Google Cloud G4 (around $4.50/hr per GPU), Azure NCv6-class instances, and neoclouds such as Vast.ai from about $1.42/hr. Prices move — treat these as order-of-magnitude, not quotes.

Server vs Workstation vs Max-Q — Same Silicon, Different Build

All three RTX PRO 6000 Blackwell variants share the same GB202 die, 24,064 CUDA cores, 752 Tensor cores, 188 RT cores and 96 GB of GDDR7 ECC. They differ in power delivery, cooling and where you can physically deploy them.

AttributeServer EditionWorkstation EditionMax-Q Edition
CoolingPassive (chassis airflow)Active dual flow-throughActive blower
PowerConfigurable, up to 600 W600 W300 W
Memory bandwidthUp to 1,597 GB/s1,792 GB/s1,792 GB/s
FP32120 TFLOPS~125 TFLOPS~110 TFLOPS (power-limited)
Display outputs4× DP 2.1b, disabled by default4× DP 2.1b, active4× DP 2.1b, active
SlotsDual-slot FHFL (single-slot liquid SKU)Dual-slotDual-slot
MIGUp to 4 × 24 GBUp to 4 × 24 GBUp to 4 × 24 GB
TargetValidated OEM servers / dense racksDesk-side tower workstationDensity-limited multi-GPU workstations

The Server Edition is the only variant with no onboard fans — the host server is its cooling system. Note the one real raw-spec trade-off: its memory bandwidth (up to 1,597 GB/s) is slightly lower than the Workstation and Max-Q editions (1,792 GB/s), a detail many spec aggregators still get wrong. That also costs it a handful of FP32 TFLOPS versus the Workstation card.

If you are choosing between them, read our companion breakdowns of the RTX PRO 6000 Blackwell Workstation Edition and the lower-power RTX PRO 6000 Blackwell Max-Q Workstation Edition.

RTX PRO 6000 Server Edition vs L40S, H100 and L4

The Server Edition's real competition is NVIDIA's own back catalogue. Here is how the key data-center PCIe cards line up:

SpecRTX PRO 6000 ServerL40SH100 PCIe (80GB)L4
ArchitectureBlackwell (GB202)Ada Lovelace (AD102)Hopper (GH100)Ada Lovelace (AD104)
Memory96 GB GDDR7 ECC48 GB GDDR6 ECC80 GB HBM2e24 GB GDDR6
Bandwidth~1,597 GB/s864 GB/s~2,000 GB/s300 GB/s
FP32120 TFLOPS91.6 TFLOPS51 TFLOPS30.3 TFLOPS
Low-precision peak4 PFLOPS (FP4)1.466 PFLOPS (FP8, sparse)~3.026 PFLOPS (FP8, sparse)0.485 PFLOPS (FP8, sparse)
FP4 supportYes (5th-gen Tensor)NoNoNo
Board powerUp to 600 W350 W350–400 W72 W
InterfacePCIe Gen 5 x16PCIe Gen 4 x16PCIe Gen 5 x16PCIe Gen 4 x16
NVLinkNoNoYes (600 GB/s bridge)No
RT cores / graphicsYes (355 TFLOPS RT)YesMinimalYes
Best forInference, rendering, VDI, mixed racksPrior-gen equivalentBandwidth-bound trainingLow-power video/inference edge

The practical read: against the L40S it is a straight upgrade — double the VRAM, nearly double the bandwidth, FP4 support the L40S simply does not have, and NVIDIA's own claim of up to 5× the LLM inference throughput. Against the H100 PCIe, it wins on capacity (96 GB vs 80 GB), graphics/RT and price-per-GB, but loses on raw memory bandwidth and, critically, on NVLink — which is what makes H100/H200/B200 the right answer for large distributed pretraining. Against L4 there is no contest on performance, but the L4's 72 W envelope still owns dense low-power video inference.

Target Workloads & Performance

NVIDIA's published comparison point for this card is the L40S, and the headline figures come from its own launch materials:

WorkloadClaimed uplift vs L40S
LLM inference throughputUp to 5×
Genomics (Smith-Waterman)Up to ~6.8×
Text-to-video generationUp to ~3.3×
Rendering / visualizationOver 2×
Inference concurrency (per rack)Scales with 8× cards per 4U chassis

Treat vendor uplift claims as best-case: most of the LLM inference gain comes from FP4 quantization plus the doubled memory pool, so a workload you refuse to quantize below FP16 will see far less than 5×. Where to expect the biggest real-world wins:

  • AI inference & serving: LLMs and multimodal/agentic AI. A 96 GB pool holds a 70B-class model at FP8, or a 120B-class model at FP4, on a single card with no tensor-parallel sharding — which alone removes a large chunk of latency and complexity.
  • Rendering & video: four NVENC and four NVDEC engines with AV1 make it a serious transcode and render-farm node, not just an AI part.
  • Virtualization / VDI: vGPU plus MIG (4 × 24 GB) for multi-tenant virtual workstations and isolated tenant workloads.
  • Scientific computing: genomics, drug discovery, CFD and data analytics all benefit from the capacity plus FP32 throughput.
  • Fine-tuning / smaller training: the large VRAM helps, but with no NVLink, multi-GPU scaling rides PCIe Gen 5 — better for inference-dense and single/few-GPU jobs than massive distributed pretraining.

Power and Cooling Requirements

This is where most failed deployments come from, so it is worth being blunt. The card is passive: it has no fans and it will thermally throttle or refuse to boot in a chassis that cannot push enough air through it.

  • Board power: configurable, up to 600 W. Operators commonly cap lower (in the 400–500 W range) to fit rack power and thermal budgets — you lose a modest slice of performance and gain a lot of density headroom.
  • Airflow: requires strong front-to-back server airflow. A 4U 8-GPU node at full tilt is a ~5 kW box before CPUs, so plan rack PDU capacity and hot-aisle containment accordingly.
  • Inlet temperature: rated for 0°C to 50°C operating inlet air; storage −40°C to 75°C.
  • Auxiliary power: a 16-pin (2x6+4) connector. The cable kit is server-specific — order it from the OEM, do not improvise with a desktop 12VHPWR adapter.
  • Liquid option: a single-slot, full-height full-length liquid-cooled SKU exists for the densest deployments, and is the route to more than 8 GPUs per chassis footprint.
  • Not for a desktop PC: a tower case has nothing like the static pressure this card needs. If you want one at a desk, buy the Workstation Edition.

Supported OEM Servers and Compatibility

The Server Edition ships almost exclusively through validated OEM platforms. Confirmed families include:

VendorExample platformsTypical GPU count
DellPowerEdge R770, XE-series GPU nodes2–8
HPEProLiant DL380a Gen12, DL385 Gen11up to 8 (4U)
LenovoThinkSystem SR650 V4 / SR650a V4, SR675 V32–8
SupermicroWorkload-optimized PCIe GPU systemsup to 8–10
ASUSIntel Xeon 6 and AMD EPYC 9005 4U RTX PRO Servers (with ConnectX-8 SuperNICs)8
CiscoUCS C-series GPU nodes2–8

It is also broadly certified on the hypervisor side — including VMware vSphere, Red Hat OpenShift/KVM and XenServer — and is offered as a first-class instance type on AWS, Google Cloud, Azure and Akamai's cloud. Before ordering, check three things: the card is on your specific server model's validated options list, the chassis has the riser slots and airflow rating for the GPU count you want, and the vendor is supplying the matching 16-pin power cable kit.

MIG, vGPU and Multi-Tenancy

Unlike the consumer Blackwell cards — the RTX 5090 class included — the Server Edition carries the full enterprise partitioning stack:

  • MIG: up to four hardware-isolated instances of 24 GB each, with independent memory and compute. Ideal for Kubernetes GPU sharing, per-tenant SLAs and packing several mid-size inference services onto one card.
  • vGPU: supported under NVIDIA vPC/vApps and RTX Virtual Workstation (vWS) — the route to multi-tenant virtual workstations for CAD, DCC and visualization users. Note that vGPU is licensed software with its own recurring cost; budget for it separately.
  • Confidential Computing: a hardware trusted execution environment protects model weights and data in use, which matters for regulated tenants and for anyone hosting third-party models.
  • Secure boot and ECC: root-of-trust secure boot, and ECC across the full 96 GB — non-negotiable for long-running production inference.

Pros and Cons

✅ Pros

  • 96 GB GDDR7 ECC — the largest VRAM pool on any discrete GPU, fits large models and scenes on one card
  • Strong FP4/FP8 inference (4 / 2 PFLOPS peak) — large gains over L40S
  • Passive design plus configurable power enables high rack density (up to 8 per 4U)
  • MIG (4 × 24 GB), vGPU and Confidential Computing for secure multi-tenant use
  • Broad OEM validation (Dell, HPE, Lenovo, Supermicro, ASUS, Cisco) and PCIe Gen 5
  • One universal SKU covering AI, rendering, simulation and VDI
  • Available on-demand from AWS, Google Cloud, Azure and neoclouds if you would rather rent

❌ Cons

  • No NVLink — limits high-bandwidth multi-GPU scaling for large training
  • ~1.6 TB/s GDDR7 bandwidth trails HBM rivals and even its own Workstation sibling
  • Requires a validated server with strong airflow — not a generic chassis, and not a desktop
  • Display outputs disabled by default (headless)
  • Price has risen roughly 55% since launch and remains volatile; vGPU adds licensing cost

The Bottom Line

Judged as a data-center product, the RTX PRO 6000 Blackwell Server Edition is excellent: category-leading VRAM capacity, superb inference and graphics-plus-AI versatility, dense passive deployment, and the full enterprise feature stack (MIG, vGPU, Confidential Computing). It loses a few points for the missing NVLink, GDDR7 bandwidth that trails HBM parts on the most bandwidth-bound workloads, and pricing that has drifted far above its launch MSRP. For inference, rendering and virtualization racks it remains the standout general-purpose choice — just benchmark it in the cloud before committing five figures per card. Our score: 88/100.

Frequently Asked Questions

How much does the NVIDIA RTX PRO 6000 Blackwell Server Edition cost in 2026?

It launched at roughly $8,565 in March 2025, but NVIDIA raised list pricing to about $13,250 — a ~55% increase driven mainly by the GDDR7 memory shortage. As of July 2026, street pricing runs roughly $11,400 to $15,000: around $11,360 from PNY, ~$12,099 at Newegg, ~$13,349 at B&H and ~$14,989 from marketplace sellers. Large orders are usually quoted through OEMs such as Dell, HPE and Lenovo as part of a server BOM.

What are the full specs of the RTX PRO 6000 Blackwell Server Edition?

It uses the Blackwell GB202 GPU with 188 SMs, 24,064 CUDA cores, 752 fifth-gen Tensor cores and 188 fourth-gen RT cores. Memory is 96 GB GDDR7 with ECC on a 512-bit bus at up to 1,597 GB/s. Peak performance is 4 PFLOPS FP4, 2 PFLOPS FP8, 1 PFLOP FP16/BF16, 234 TFLOPS TF32, 120 TFLOPS FP32 and 355 TFLOPS RT. It is a dual-slot FHFL PCIe Gen 5 x16 card with configurable board power up to 600 W.

What is the difference between the RTX PRO 6000 Server and Workstation editions?

They use the same GB202 GPU and 96GB GDDR7 ECC memory, but the Server Edition is passively cooled (relying on server chassis airflow), has configurable power up to 600W, slightly lower bandwidth at up to 1,597 GB/s, and display outputs disabled by default. The Workstation Edition has active dual flow-through fans, 600W, the full 1,792 GB/s bandwidth, and active display outputs for a desk-side tower.

How does the RTX PRO 6000 Server Edition compare to the L40S and H100?

Versus the L40S it is a straight upgrade: 96GB vs 48GB, ~1,597 GB/s vs 864 GB/s, PCIe Gen 5, FP4 Tensor support the L40S lacks, and NVIDIA claims up to 5x the LLM inference throughput. Versus the H100 PCIe it offers more capacity (96GB vs 80GB) and far better graphics/RT performance, but the H100's HBM2e bandwidth (~2 TB/s) and NVLink make it the better pick for large distributed training.

Can I use the RTX PRO 6000 Server Edition in a normal PC?

Not practically. It has no fans and relies on high-static-pressure server front-to-back airflow, its display outputs are disabled by default, and it needs a server-specific 16-pin power cable kit. It is designed for validated OEM servers from Dell, HPE, Lenovo, Supermicro, ASUS and Cisco. For a tower workstation, choose the Workstation Edition instead.

Which servers support the RTX PRO 6000 Blackwell Server Edition?

Validated platforms include the Dell PowerEdge R770 and XE GPU nodes, HPE ProLiant DL380a Gen12 and DL385 Gen11, Lenovo ThinkSystem SR650 V4 / SR650a V4 and SR675 V3, Supermicro's workload-optimized PCIe GPU systems, ASUS Intel Xeon 6 and AMD EPYC 9005 4U RTX PRO Servers, and Cisco UCS C-series. Dense 4U platforms typically hold up to eight cards. Always check your exact server model's validated options list.

Does the RTX PRO 6000 Server Edition support MIG and vGPU?

Yes to both. MIG allows up to four hardware-isolated instances of 24 GB each with independent memory and compute, which suits Kubernetes GPU sharing and multi-tenant inference. vGPU is supported via NVIDIA vPC/vApps and RTX Virtual Workstation (vWS), though vGPU is separately licensed software with a recurring cost.

How much power does the RTX PRO 6000 Server Edition draw, and does it need special cooling?

Board power is configurable up to 600 W, and many operators cap it lower (around 400-500 W) to fit rack power and thermal budgets. The card is passively cooled with no onboard fans, rated for 0°C to 50°C inlet air, and requires strong front-to-back server airflow. A single-slot FHXL liquid-cooled SKU is available for the densest deployments.

Does the RTX PRO 6000 Server Edition support NVLink?

No. This SKU does not support NVLink — multi-GPU communication is over PCIe Gen 5 peer-to-peer only. For large-scale distributed training, NVIDIA's HBM-based HGX H100/H200/B200 platforms remain the better fit.

Can I rent the RTX PRO 6000 Server Edition in the cloud instead of buying it?

Yes, and at current prices it is often the smarter first step. It is offered on AWS EC2 G7e instances from around $3.36/hour for a single GPU (scaling to roughly $33/hour for eight), Google Cloud G4 at about $4.50/hour per GPU, Azure NCv6-class instances, and neoclouds such as Vast.ai from around $1.42/hour. Rates move frequently, so treat these as ballpark figures.

NVIDIARTX PRO 6000BlackwellGraphics Cards
Advertisement