NVIDIA says Vera Rubin boosts AI throughput per megawatt
Wed, 26th Aug 2026 (Today)
NVIDIA said its Vera Rubin NVL72 systems delivered up to 30 times higher throughput per megawatt than GB300 NVL72 systems on agentic AI workloads, based on the SemiAnalysis AgentX workload.
The figures address a growing part of the AI market in which software agents handle complex, multi-step tasks that generate far more tokens than standard chatbot requests. Citing OpenRouter data, NVIDIA said agentic AI workloads consume 15 times more tokens than a simple chat request because agents query databases, search documents, call tools and invoke sub-agents before producing an answer.
That shift has made long-context processing and energy efficiency more important for AI infrastructure operators. NVIDIA said each step in an agentic workflow adds to the context window for the next step, increasing computational load and raising the importance of throughput per megawatt and the cost of producing tokens.
NVIDIA also said Vera Rubin NVL72 delivered up to 35 times lower cost per million tokens than GB300 NVL72. It said those gains would allow operators of power-constrained AI facilities to complete more agentic work within the same energy budget.
The benchmark used recorded agentic coding sessions from SemiAnalysis AgentX, retaining context growth, tool calls and sub-agent spawning. NVIDIA said the early results are pending SemiAnalysis review and do not yet include Vera CPU performance for tool calling.
Workload shift
Agentic AI workloads differ from chat and document summarisation systems because they can span hundreds of thousands of input tokens, with wide variation in input and output lengths. That makes single-request inference tests less useful for gauging real-world performance, particularly for operators trying to size systems for sustained use.
NVIDIA said the Blackwell platform had already shown strong results on AgentX across models including Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro. It added that GB300 NVL72 delivered up to 15 times better throughput per megawatt than systems based on the Hopper architecture on the DeepSeek V4 Pro model.
NVIDIA said Vera Rubin extends that performance lead across what it described as the full Pareto curve, with the largest improvement reported on DeepSeek V4 Pro. It attributed the gains to a larger scale-up GPU domain and software designed alongside the hardware.
System design
NVIDIA attributed the latest performance figures to a series of inference techniques used for long-running AI agents. These include disaggregated serving, which separates context processing from response generation, and rate matching, which aligns the speed of systems handling those distinct tasks.
It also pointed to large-scale expert parallelism for mixture-of-experts models, distributed key-value caching across the GPU domain, and cache offloading to host and storage tiers for less active context. NVIDIA said those methods reduce recomputation and improve the handling of long sessions.
Other methods include routing requests to GPUs that already store the relevant context, as well as fused CUDA kernels that combine computation and inter-GPU communication in a single pass. NVIDIA said Rubin GPUs also use fifth-generation Tensor Cores, a third-generation Transformer Engine and NVFP4 quantisation to reduce memory use and increase throughput.
The NVL72 scale-up domain remains central to both Vera Rubin and Grace Blackwell designs. NVIDIA said its sixth-generation NVLink technology and NVLink switches provide 10 times higher packet rates and three times lower latency than Ethernet alternatives.
Power economics
NVIDIA also highlighted its DSX MaxLPS power management technology, which it said can provision up to 40% more GPUs within the same megawatt budget by managing power across the GPU, rack and workload levels. The claim points to a broader industry concern over electricity costs and the difficulty of expanding AI infrastructure where grid capacity is limited.
NVIDIA tied the technical claims directly to AI factory economics, arguing that throughput per megawatt influences revenue while token cost affects profit margins. That framing reflects pressure on cloud providers, model developers and large enterprises to justify the rising cost of running increasingly complex AI services.
Beyond the GPU system itself, NVIDIA said the wider Vera Rubin platform includes the Vera CPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX and ConnectX-9 SuperNIC. It added that Vera Rubin is in full production and scaling across its partner ecosystem.
"These early results for Vera Rubin NVL72 demonstrate NVIDIA's accelerated pace of innovation."