An operator who cannot get another megawatt from the utility has only one lever left: extract more sellable output from the power already contracted. That is the argument NVIDIA built its AI Infra Summit keynote around this week, and the company came with a customer number to prove the lever moves โ Lambda fit 19 servers into a power budget normally reserved for 16, and got 24% more tokens per second out of the cluster for it.
Key takeaways
- Lambda's DSX MaxLPS deployment lifted cluster-wide token throughput from roughly 4 million to 5 million tokens per second and performance per watt by 23% โ the first production validation of the software on Blackwell servers.
- NVIDIA reported up to 30x higher throughput per megawatt and up to 45x lower cost per million tokens for Vera Rubin NVL72 against GB300 NVL72, measured on replayed agentic coding sessions rather than single requests.
- A separate demonstration with Emerald AI had an AI facility answering hundreds of live demand signals from a municipal utility without disturbing priority jobs.
Why the yardstick is changing
Peak FLOPS made sense as a buying signal when silicon supply was the constraint. It makes much less sense in a building where the interconnection queue, not the purchase order, decides how much compute can run. NVIDIA's framing โ set out by Ian Buck, its vice president of hyperscale and high-performance computing, in Tuesday's summit announcement โ is that validated agentic tokens per megawatt is the figure that actually maps to revenue.
The benchmark community is moving the same way. SemiAnalysis built AgentX around recorded real-world agentic coding sessions, preserving context growth, tool-call latency and sub-agent spawning instead of timing isolated prompts. NVIDIA puts one such session at roughly 15 times the token volume of an ordinary chat request, with input and output lengths varying widely between calls. That is not a workload a per-request suite describes well, which is why NVIDIA positions AgentX as a complement to MLPerf inference testing rather than a replacement.
Against that yardstick, NVIDIA's published AgentX figures for Vera Rubin NVL72 running DeepSeek V4 Pro reach 30x the throughput per megawatt of the prior GB300 generation, with cost per million tokens down as much as 45x. Both are stated as ceilings, not typical results.
What the Lambda result shows in practice
DSX MaxLPS is the piece doing the work in Lambda's case. Rather than provisioning each rack for its worst-case draw and stranding whatever it does not use, the software tracks consumption across GPUs and racks continuously and redirects available headroom to the jobs that need it. Training and serving have different power profiles, so a mixed estate leaves a great deal of that headroom idle by default.
Three extra nodes inside an unchanged power envelope is a modest-sounding win that compounds at scale, and it needs no new substation, cabling or permit. NVIDIA expects the margin to widen on its next generation, projecting up to 40% additional GPU capacity and up to 35% higher token throughput for Vera Rubin NVL72 sites under the same megawatt ceiling. For latency-bound agent work, the company is also pairing Vera Rubin with Groq 3 LPX, reporting 2,529 output tokens per second per user on a 100K-context Qwen 3.8 27B run.
The grid becomes part of the architecture
The demonstration with the longest tail had little to do with chips. Through Silicon Valley Power's flexible-load interconnection programme, Emerald AI showed a facility cutting draw automatically in response to utility signals โ hundreds of them โ while critical jobs kept running. Emerald intends to build its Conductor power-management product on NVIDIA DSX Flex, which reads load-shedding requests, demand-response events and pricing changes and acts within a hierarchy the operator defines in advance.
The point of the exercise is regulatory as much as technical. Utilities ration interconnection capacity on the assumption that a data centre is an inflexible block of demand; a facility that can prove it will yield on request has a different case to make in the queue. That argument runs in parallel with the capital NVIDIA has been marshalling for AI factory buildouts, and it addresses the constraint money alone cannot clear.
How much to bank on the numbers
Every multiple here comes from a vendor or a partner, measured on workloads the vendor chose, and the headline comparisons pit a current generation against its immediate predecessor. Lambda's is the outlier worth weighting: it is a production cluster, the deltas are small enough to be plausible, and other operators running Blackwell can check it. The 30x and 45x AgentX claims sit on a public dashboard, which at least makes them contestable.
Partner news rounded out the slate โ Amazon's Annapurna Labs is co-developing NVHBM custom memory, d-Matrix is attaching Raptor XPUs to Vera CPUs over NVLink Fusion, and Pinterest has Blackwell and Dynamo behind conversational visual search. With NVIDIA saying Vera Rubin is in full production, independent measurement should arrive soon enough to settle the rest.
FAQ
What is tokens per megawatt?
It is throughput normalised by electrical draw rather than by chip count or peak FLOPS. Where a utility connection rather than capital caps expansion, it tells an operator how much sellable output each contracted megawatt produces, which is closer to a revenue figure than a hardware spec.
How much did DSX MaxLPS improve Lambda's cluster?
Lambda ran 19 nodes in the power envelope usually allocated to 16 full-power nodes, raising cluster-wide token throughput roughly 24%, from about 4 million to 5 million tokens per second, and performance per watt by 23%. NVIDIA describes it as the first validation of MaxLPS on Blackwell servers.
Is the SemiAnalysis AgentX benchmark different from MLPerf?
Yes. AgentX replays whole agent trajectories, including tool calls and sub-agent spawning, while MLPerf Inference scores per-request performance across a wide range of models and scenarios. NVIDIA presents the two as complementary, since they answer different questions about the same hardware.






