A 30-fold efficiency claim can move an infrastructure plan. It should not move one alone. NVIDIA’s new Vera Rubin result measures accepted agent work per megawatt, but the company says the analysis is early and still awaits an outside review.
NVIDIA Vera Rubin agent efficiency is the headline from the company’s August 23 disclosures. NVIDIA reports that one Vera Rubin NVL72 rack can deliver up to 30 times more throughput per megawatt than a GB300 NVL72 system on a DeepSeek V4 Pro agent workload evaluated with SemiAnalysis AgentX. It also reports up to 35 times lower cost per million tokens.
Those are NVIDIA figures from an early implementation. SemiAnalysis has not yet published the promised independent review, and NVIDIA says the current result excludes tool-calling work performed by the Vera CPU. The useful reading is therefore narrower: the rack is showing a large company-measured gain on one agent-oriented test, not a settled 30× advantage for every production system.
The metric is better than tokens per second
Agent workloads spend time reasoning, calling tools, waiting on services, checking results, and retrying failed branches. Raw output speed can look excellent while little useful work reaches completion. AgentX instead emphasizes accepted work under a power budget. That puts the scarce data-center resource—megawatts—closer to the outcome an operator buys.
The denominator still matters. “Per megawatt” can hide differences in rack configuration, utilization, cooling assumptions, host CPUs, networking, or which external tool time is counted. A fair buyer test needs the complete system boundary and the same task acceptance rule on both systems.
What NVIDIA actually compared
NVIDIA identifies Vera Rubin NVL72 and GB300 NVL72 as the two rack-scale systems. The model is DeepSeek V4 Pro, and the workload uses SemiAnalysis AgentX. NVIDIA describes the Vera Rubin stack as being in full production and says the result will be reviewed by SemiAnalysis.
| Claim | Status | What is still needed |
|---|---|---|
| Up to 30× throughput per megawatt | NVIDIA-reported | Independent review, full configuration, and run variance |
| Up to 35× lower cost per million tokens | NVIDIA-reported | Price inputs, utilization, depreciation, power, and cooling assumptions |
| DeepSeek V4 Pro on AgentX | Named workload | Task list, acceptance criteria, prompts, and repeatability details |
| Vera CPU tool calling | Excluded from current result | End-to-end measurement with tools included |
The 35× cost number is not a purchase quote
Cost per million tokens is convenient because model providers publish it. An owned AI factory has a different ledger. Hardware price, financing, utilization, power contracts, cooling, networking, maintenance, software, and operator time all enter the calculation.
A rack that is 35 times cheaper under one assumed utilization can be far less impressive when demand is bursty or the application spends most of its wall time waiting for a database. Ask for cost per accepted task at your expected duty cycle, plus the idle-power and partial-load curves. Procurement should not substitute a token denominator for a business outcome.
Groq 3 LPX adds a different speed option
NVIDIA also announced Groq 3 LPX, reporting 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context window. The company calls that four times the nearest alternative and says Groq 3 LPX is in full production.
This is a separate serving result, not support for the 30× AgentX claim. Fast sequential generation can materially improve voice, interactive coding, and long reasoning loops. It still needs quality parity, time-to-first-token, batch-size, concurrency, context-length, and total-system-power measurements. The model and serving stack must be treated as one tested unit.
A buyer-side acceptance test
- Select 50 to 100 real tasks with a written acceptance test and known failure cost.
- Freeze model version, prompts, tool permissions, context, retry policy, and stopping rules.
- Measure accepted tasks, not attempted tasks or generated tokens.
- Record wall time, time to first token, output rate, tool wait time, retries, and human review.
- Measure rack input power and facility overhead at idle, partial load, and sustained load.
- Repeat runs across several days and report median, tail latency, and variance.
- Price the full system at realistic utilization and include migration, staffing, networking, and maintenance.
Our review of Fable 5’s NanoGPT Speedrun shows why the search harness and time budget belong beside the model name. Our guide to NVIDIA NeMo Switchyard applies the same principle to routing: trust, cost, and data boundaries change with the complete stack.
My verdict: wait for the denominator
NVIDIA is asking a more useful question than “How many tokens can the chip emit?” Accepted agent work per megawatt points toward the constraint that increasingly shapes AI deployments. The reported scale of the gain also makes the result worth testing now.
I would use the claim to secure an evaluation slot, not to approve a purchase. The decision changes when SemiAnalysis publishes its review, the complete system boundary is visible, and your own tasks show better cost per accepted result. Until then, 30× is a promising company-reported ceiling.
Read the primary disclosures
- Read NVIDIA’s Vera Rubin agent-efficiency announcement.
- Review the developer benchmark explanation.
- Compare the Vera Rubin and Groq 3 LPX production announcement.
Which acceptance test would you require before treating 30× per megawatt as a planning number?
Checked August 24, 2026. Throughput, cost, output-speed, production-status, comparison, and exclusion figures are NVIDIA statements. SemiAnalysis’s independent review was not public when checked.