Skip to main content

NVIDIA Says DSX Raised AI Token Throughput 24% Without More Power

4 min read

NVIDIA reports 24% more B200 token throughput within the same power budget. A new alliance with Google and Emerald AI now targets flexible-grid standards.

NVIDIA Says DSX Raised AI Token Throughput 24% Without More Power

September 16 update: Google, NVIDIA and Emerald AI form AEMA

Update, September 16, 2026: Emerald AI, Google and NVIDIA have announced the AI Energy Management Alliance, or AEMA. The group proposes performance-based standards for connecting flexible AI data centers to power grids. Its initial framework focuses on response speed, response duration, ride-through capability, contingency obligations, operational data sharing, risk-adjusted interconnection and cost allocation.

AEMA turns one deployment idea into a standards effort

The original DSX analysis below covers a company-reported deployment where software coordinated an AI fleet inside a fixed site power budget. AEMA addresses the next boundary: how a utility or grid operator should evaluate a data center that promises to change demand when the grid needs support. Instead of treating every facility as an inflexible peak load, the alliance wants interconnection rules to recognize measured flexibility.

AEMA proposalQuestion a grid operator must answerEvidence needed
Response speedHow quickly can load change after a signal?Metered response under realistic operating conditions
Response durationHow long can the facility sustain the change?Repeated events across workload and weather conditions
Ride-throughCan the site remain stable during grid disturbances?Protection studies and witnessed tests
Data sharingWhat telemetry can the operator trust?Defined schemas, timing, security and audit history
Cost allocationWho pays for upgrades and residual risk?Transparent planning scenarios and enforceable obligations
AEMA identifies these policy areas. The evidence column is MustHave.ai analysis of what implementation would require.

The alliance is not yet proof of grid adoption

AEMA is an industry alliance and proposed framework, not a regulator-approved interconnection standard. Utilities, system operators and regulators will decide whether the performance definitions are measurable, enforceable and fair to other customers. The practical milestone is a tariff or interconnection agreement that attaches clear penalties and verification rules to promised flexibility.

This update strengthens the strategic relevance of DSX without changing the status of NVIDIA's throughput claims. The 24% B200 result and the 40% Vera Rubin projection remain company reported and should still be tested on each operator's workload and electrical design.

NVIDIA says its DSX infrastructure software helped Lambda run 19 HGX B200 nodes inside the electrical budget of 16 full-power nodes. The company reports token throughput rising from roughly 4 million to 5 million tokens per second, a 24% gain without increasing the site power envelope.

The result appears in NVIDIA’s September 15 account of how DSX coordinates AI factory power. It is a company-reported deployment result, not an independent benchmark. Its value is the operating model: optimize useful tokens across a power-constrained fleet instead of assuming every server must draw peak power simultaneously.

The reported arithmetic

MetricReported baselineReported DSX result
HGX B200 nodes inside the budget16 full-power nodes19 coordinated nodes
Token throughputAbout 4M tokens/secondAbout 5M tokens/second
Throughput changeBaseline+24%
Performance per wattBaseline+23%
All figures are NVIDIA and Lambda reported.

Why a fleet can beat full-power operation

GPU power and application throughput are not perfectly linear. A node near its maximum power limit may consume a disproportionate share of the final watts for a smaller throughput increase. If software can coordinate power caps, workload priority and thermal headroom across many nodes, the site may run more machines at a more efficient point.

The strategy resembles filling a fixed elevator with more well-packed boxes instead of allowing fewer boxes to occupy the maximum floor space. The gain depends on model type, batch size, latency target, hardware generation and how evenly work can be shifted.

Tokens per watt needs a quality and latency boundary

A raw token is not a complete unit of customer value. Operators should pair tokens per watt with time to first token, inter-token latency, response quality, request completion and service-level violations. A system that emits more low-value tokens or misses latency targets can look efficient while reducing product quality.

The 4 MW to 3 MW event tests a different capability

NVIDIA also reports a grid-response demonstration in which an AI load dropped from 4 MW to 3 MW automatically while priority inference continued. Santa Clara utility Silicon Valley Power had previously confirmed the pilot program. The utility page confirms the project, not NVIDIA’s later performance result.

The Vera Rubin figure is a projection

NVIDIA says DSX could enable up to 40% more Vera Rubin NVL72 capacity at the same megawatt level in suitable environments. That is forward-looking. It should not be merged with the B200 deployment result. Hardware topology, cooling, power distribution and workload flexibility will determine whether a site approaches the projection.

A buyer validation plan

  1. Record the site power cap, node count, GPU model and cooling conditions.
  2. Run the same model mix under fixed-power and coordinated-power policies.
  3. Measure accepted requests, tokens, latency, quality and power at the meter.
  4. Test a grid curtailment event with a declared priority workload.
  5. Repeat during high ambient temperature and partial node failure.

Our NVIDIA JAX MoE analysis shows why a named workload and hardware label must stay attached to performance claims. The NIM and Nemotron throughput guide covers a separate B200 inference claim.

The practical verdict

NVIDIA DSX reframes an AI data center as a power-scheduled production system. The reported 24% throughput gain is promising, but the transferable lesson is to optimize accepted, quality-controlled work per site watt and verify the result on the buyer’s own workload.

Primary sources

Checked September 15, 2026. Performance figures and projections are company reported. MustHave.ai has not independently audited the deployment.

Leave a comment

Your email address will not be published. Required fields are marked *