AMD EPYC 9354 Servers —from €299/month or €0.42/hour ⭐ 32 cores 3.25GHz / 768GB RAM / 2x3.84TB NVMe / 10Gbps 100TB
EN
Currency:
EUR – €
Choose a currency
  • Euro EUR – €
  • United States dollar USD – $
VAT:
OT 0%
Choose your country (VAT)
  • OT All others 0%

09.08.2026

NVIDIA RTX PRO 5000 Blackwell with 72 GB VRAM: Is the "Half-Flagship" Worth the Premium?

server one
HOSTKEY

NVIDIA is making things increasingly confusing with its Blackwell lineup. A look at the PCI.ids for released professional Blackwell GPUs reveals the following:

  # RTX PRO Blackwell Series (Desktop / Workstation)
  "10de:2bb5|RTX PRO 6000 Blackwell Server Edition"
  "10de:2bb4|RTX PRO 6000 Blackwell Workstation Edition"
  "10de:2bb1|RTX PRO 6000 Blackwell Max-Q"
  "10de:2bb3|RTX PRO 5000 Blackwell"
  "10de:2c31|RTX PRO 4500 Blackwell"
  "10de:2c37|RTX PRO 4000 Blackwell"
  "10de:2c36|RTX PRO 3000 Blackwell"
  "10de:2d30|RTX PRO 2000 Blackwell"
  "10de:2c35|RTX PRO 2000 Blackwell (Alternative)"
 
  # RTX PRO Blackwell Series (Server / Embedded)
  "10de:2c3a|RTX PRO 4500 Blackwell Server Edition"
  "10de:2c77|RTX PRO 5000 Blackwell Embedded GPU"
 
  # RTX PRO Blackwell Series (Laptop / Mobile)
  "10de:2c38|RTX PRO 5000 Blackwell Generation Laptop GPU"
  "10de:2f38|RTX PRO 3000 Blackwell Generation Laptop GPU"
  "10de:2db8|RTX PRO 1000 Blackwell Generation Laptop GPU"

Each version can also have multiple variants, as seen with our current test subject: the NVIDIA RTX PRO 5000 Blackwell with an expanded 72 GB of VRAM. We effectively have two versions:

  • RTX PRO 5000 Blackwell (48 GB GDDR7): The original base version. It features single-sided memory chips and supports MIG (Multi-Instance GPU) split into two 24 GB instances.
  • RTX PRO 5000 Blackwell (72 GB GDDR7): A special high-memory variant released later for inferencing larger AI models. This uses a clamshell configuration (double-sided memory chips). It supports MIG in a 2 x 36 GB mode.

In the current market, the base 48 GB model starts at approximately $5,000. The 72 GB version is estimated to start at $9,999 — effectively double the price.

NVIDIA and AMD GPU Servers
Dedicated and virtual servers with latest and previous generation GPUs on hourly billing

When compared to the flagship PRO 6000 Blackwell, the purchase seems questionable at first glance:

Feature

RTX PRO 5000 (72 GB)

RTX PRO 6000 (96 GB)

VRAM Capacity

72 GB GDDR7 ECC

96 GB GDDR7 ECC

CUDA Cores

14,080

24,064

Tensor Cores

440 (5th Gen)

752 (5th Gen)

Memory Bandwidth

1,344 GB/s

1,792 GB/s

FP32 Performance

72.2 TFLOPS

125 TFLOPS

TDP

300 W

600 W

Given the original MSRP of $8,000 for the PRO 6000 Blackwell, the existence of a 72 GB 5000-series version raises questions. However, current memory market realities mean this card is now retailing at twice its original launch price. Interestingly, while the RTX PRO 5000 Blackwell has half the power consumption, it isn't only "half" as performant — it actually delivers about 60% of the flagship's performance (a ~40% deficit). The trade-off is significantly lower requirements for power delivery and cooling.

The form factor remains standard: a dual-slot, full-height, full-length (FHFL) design with active blower-style cooling, identical to the RTX PRO 6000.

Installation and Testing

We secured a sample card and conducted comparative benchmarks. We had our team build a server, installed Ubuntu 24.04, and loaded drivers using our updated script.

We then ran our GPU benchmark suite. Unfortunately, we couldn't source an RTX PRO 6000 for a full cross-model comparison, so all comparisons were conducted against the DeepSeek R1 model.

GPU

VRAM

Model

Tokens/sec (avg)

Max Ctx

Load (s) avg

Gen (s) avg

NVIDIA RTX PRO 5000 Blackwell

72 GB

deepseek-r1:32b

50.77

128,000

4.32

50.69

NVIDIA RTX PRO 5000 Blackwell

72 GB

deepseek-r1:70b

25.08

96,000

6.67

99.12

NVIDIA RTX PRO 5000 Blackwell

72 GB

gpt-oss:120b

156.52

128,000

9.43

21.70

NVIDIA RTX PRO 5000 Blackwell

72 GB

qwen3-coder-next:q4_K_M

192.13

128,000

7.50

14.97

NVIDIA RTX 6000 PRO Blackwell (gen5)

96 GB

deepseek-r1:32b

58.73

128,000

2.44

42.91

NVIDIA RTX 6000 PRO Blackwell (gen5)

96 GB

deepseek-r1:70b

30.19

112,000

4.5

84.23

NVIDIA RTX A5000 (gen3)

24 GB

deepseek-r1:32b

25.77

12,000

11.49

94.10

2x AMD RADEON AI PRO R9700

64 GB (2x32)

deepseek-r1:32b

25.90

96,000

15.2

92.1

3x AMD RADEON AI PRO R9700

96 GB (3x32)

deepseek-r1:70b

12.77

76,000

17.0

210.1

3x AMD RADEON AI PRO R9700

96 GB (3x32)

gpt-oss:120b

60.75

128,000

21.05

57.34

3x AMD RADEON AI PRO R9700

96 GB (3x32)

qwen3-coder-next:q4_K_M

37.42

128,000

17.85

79.55

The 72 GB of VRAM allows the PRO 5000 Blackwell to handle large context windows (up to 128k tokens in our tests for mid-sized models) and run heavy models (up to 120B) without issues. Model loading speeds are exceptionally high, averaging between 4 and 9 seconds, which highlights excellent memory bandwidth.

Comparison with the Flagship (RTX PRO 6000 Blackwell, 96 GB)

Despite having inferior raw hardware specs, the RTX PRO 5000 demonstrated results remarkably close to the flagship in our tests, trailing by only 15–20%.

  • DeepSeek-R1 (32B): The speed difference is minimal. The RTX 5000 delivers 50.77 tokens/sec compared to 58.73 on the RTX 6000. Generation time differs by only about 8 seconds. Both cards maintain a max context of 128,000 tokens.
  • DeepSeek-R1 (70B): On larger models, the gap widens slightly: 25.08 vs. 30.19 tokens/sec (~17% deficit). Here, memory capacity becomes a factor. The RTX PRO 5000 is limited to a 96k token context, whereas the RTX PRO 6000's 96 GB of VRAM allows for up to 112k tokens.

Comparison with Previous Generation (RTX A5000, 24 GB)

The leap here is massive, driven by both the Blackwell architecture and the increased memory capacity. Note that we are not even considering optimized models like those using NVFP4 in this comparison.

On the deepseek-r1:32b model, the new RTX PRO 5000 is nearly twice as fast as the previous generation (50.77 vs. 25.77 tokens/sec). Response time dropped from 1.5 minutes (94s) to just 50 seconds. This improvement is due to both the architecture and the modernized memory type.

Furthermore, due to its modest 24 GB of VRAM, the older RTX A5000 "chokes" on heavy workloads, strictly limiting the context window to 12,000 tokens. The 72 GB on the RTX PRO 5000 allows for 10x more context (128k+ tokens), which fundamentally transforms performance when working with long documents or extensive chat histories.

The AMD Showdown (Radeon AI PRO R9700)

Even against multi-card AMD setups, a single RTX PRO 5000 shows absolute dominance in generation speed — likely due to superior software optimization for the NVIDIA/CUDA ecosystem. To be fair to AMD, you can purchase three R9700 cards for the price of one 48 GB PRO 5000 (current pricing), giving you 96 GB of VRAM total.

  • DeepSeek-R1 (32B): A single RTX 5000 is twice as fast as a dual-card AMD R9700 setup.
  • GPT-OSS (120B): The RTX 5000 delivers an incredible 156.52 tokens/sec, outperforming a triple-card AMD setup (60.75 tokens/sec) by 2.5x. NVIDIA responds in 21 seconds; AMD takes nearly a minute (57s).
  • Qwen3-Coder-Next (q4_K_M): This was a total blowout for the "Red Team." The RTX 5000 generates 192.13 tokens/sec — more than 5x faster than three AMD cards (37.42 tokens/sec). NVIDIA finishes in 15 seconds, while the AMD setup takes nearly 80 seconds.

AMD's strength lies purely in memory capacity and price-per-GB; in terms of raw speed, they are outperformed by a single mid-range Blackwell chip.

Verdict

If I were writing this article six months ago, my verdict would have been: "Just pay the extra and get the RTX PRO 6000 Blackwell."

However, given current market prices and memory shortages, the NVIDIA RTX PRO 5000 Blackwell is the ideal "sweet spot" for enthusiasts and professionals working with local LLMs. In practical inference tasks, it trails the flagship by only about 25% while consuming half the power — all while leaving the previous generation (A5000) in the dust by removing context length bottlenecks.

Additionally, thanks to MIG 2 support, you can split this card into two isolated virtual GPUs with 36 GB of VRAM each and guaranteed QoS, providing significant savings if you are running your own virtual GPU server infrastructure.

NVIDIA and AMD GPU Servers
Dedicated and virtual servers with latest and previous generation GPUs on hourly billing

Other articles

08.08.2026

Top 10 WordPress Plugins for Online Stores in 2026

Which plugin should you choose for a WordPress online store in 2026? This article compares 10 popular solutions for physical goods, digital products, subscriptions, and payments — from WooCommerce to Ecwid and WP Simple Pay.

08.08.2026

Building Our Own Programming Language Ranking Using GitHub Data in Anaconda and JupyterLab

We didn't argue with TIOBE or RedMonk — we built our own programming language ranking from GitHub data. The 2024–2025 numbers hold a few surprises: JavaScript leads, TypeScript surges, and Rust and Go win on project quality. We break down what's behind the numbers and where the distortions live.

08.08.2026

How to Revive Internal Documentation: An ONLYOFFICE Workspace Case Study

Documentation dies not because employees are lazy, but because it is inconvenient to use and people stop trusting it. We look at the "two pillars" of a good knowledge base: usability and control over how current it is. Using ONLYOFFICE Workspace as an example, we show how to turn chaos into a working process with templates, role-based access and review discipline.

31.07.2026

Large Models and the Cost per Million Tokens

Are Chinese models actually cheaper, or is it a trap? Learn how to avoid overpaying for tokens and why list prices don't tell the whole story regarding AI expenditures.

30.07.2026

All in One SEO (AIOSEO) for WordPress: Complete Plugin Review in 2026

All in One SEO (AIOSEO) is one of the oldest SEO plugins for WordPress with over 3 million active installations. We break down its features, pricing plans, pros, and cons in 2026.

Upload