September 16 update: Google, NVIDIA and Emerald AI form AEMA
Update, September 16, 2026: Emerald AI, Google and NVIDIA have announced the AI Energy Management Alliance, or AEMA. The group proposes performance-based standards for connecting flexible AI data centers to power grids. Its initial framework focuses on response speed, response duration, ride-through capability, contingency obligations, operational data sharing, risk-adjusted interconnection and cost allocation.
AEMA turns one deployment idea into a standards effort
The original DSX analysis below covers a company-reported deployment where software coordinated an AI fleet inside a fixed site power budget. AEMA addresses the next boundary: how a utility or grid operator should evaluate a data center that promises to change demand when the grid needs support. Instead of treating every facility as an inflexible peak load, the alliance wants interconnection rules to recognize measured flexibility.
| AEMA proposal | Question a grid operator must answer | Evidence needed |
|---|---|---|
| Response speed | How quickly can load change after a signal? | Metered response under realistic operating conditions |
| Response duration | How long can the facility sustain the change? | Repeated events across workload and weather conditions |
| Ride-through | Can the site remain stable during grid disturbances? | Protection studies and witnessed tests |
| Data sharing | What telemetry can the operator trust? | Defined schemas, timing, security and audit history |
| Cost allocation | Who pays for upgrades and residual risk? | Transparent planning scenarios and enforceable obligations |
The alliance is not yet proof of grid adoption
AEMA is an industry alliance and proposed framework, not a regulator-approved interconnection standard. Utilities, system operators and regulators will decide whether the performance definitions are measurable, enforceable and fair to other customers. The practical milestone is a tariff or interconnection agreement that attaches clear penalties and verification rules to promised flexibility.
This update strengthens the strategic relevance of DSX without changing the status of NVIDIA's throughput claims. The 24% B200 result and the 40% Vera Rubin projection remain company reported and should still be tested on each operator's workload and electrical design.
NVIDIA says its DSX infrastructure software helped Lambda run 19 HGX B200 nodes inside the electrical budget of 16 full-power nodes. The company reports token throughput rising from roughly 4 million to 5 million tokens per second, a 24% gain without increasing the site power envelope.
The result appears in NVIDIA’s September 15 account of how DSX coordinates AI factory power. It is a company-reported deployment result, not an independent benchmark. Its value is the operating model: optimize useful tokens across a power-constrained fleet instead of assuming every server must draw peak power simultaneously.
The reported arithmetic
| Metric | Reported baseline | Reported DSX result |
|---|---|---|
| HGX B200 nodes inside the budget | 16 full-power nodes | 19 coordinated nodes |
| Token throughput | About 4M tokens/second | About 5M tokens/second |
| Throughput change | Baseline | +24% |
| Performance per watt | Baseline | +23% |
Why a fleet can beat full-power operation
GPU power and application throughput are not perfectly linear. A node near its maximum power limit may consume a disproportionate share of the final watts for a smaller throughput increase. If software can coordinate power caps, workload priority and thermal headroom across many nodes, the site may run more machines at a more efficient point.
The strategy resembles filling a fixed elevator with more well-packed boxes instead of allowing fewer boxes to occupy the maximum floor space. The gain depends on model type, batch size, latency target, hardware generation and how evenly work can be shifted.
Tokens per watt needs a quality and latency boundary
A raw token is not a complete unit of customer value. Operators should pair tokens per watt with time to first token, inter-token latency, response quality, request completion and service-level violations. A system that emits more low-value tokens or misses latency targets can look efficient while reducing product quality.
The 4 MW to 3 MW event tests a different capability
NVIDIA also reports a grid-response demonstration in which an AI load dropped from 4 MW to 3 MW automatically while priority inference continued. Santa Clara utility Silicon Valley Power had previously confirmed the pilot program. The utility page confirms the project, not NVIDIA’s later performance result.
The Vera Rubin figure is a projection
NVIDIA says DSX could enable up to 40% more Vera Rubin NVL72 capacity at the same megawatt level in suitable environments. That is forward-looking. It should not be merged with the B200 deployment result. Hardware topology, cooling, power distribution and workload flexibility will determine whether a site approaches the projection.
A buyer validation plan
- Record the site power cap, node count, GPU model and cooling conditions.
- Run the same model mix under fixed-power and coordinated-power policies.
- Measure accepted requests, tokens, latency, quality and power at the meter.
- Test a grid curtailment event with a declared priority workload.
- Repeat during high ambient temperature and partial node failure.
Our NVIDIA JAX MoE analysis shows why a named workload and hardware label must stay attached to performance claims. The NIM and Nemotron throughput guide covers a separate B200 inference claim.
The practical verdict
NVIDIA DSX reframes an AI data center as a power-scheduled production system. The reported 24% throughput gain is promising, but the transferable lesson is to optimize accepted, quality-controlled work per site watt and verify the result on the buyer’s own workload.
Primary sources
Checked September 15, 2026. Performance figures and projections are company reported. MustHave.ai has not independently audited the deployment.