Nvidia said Monday that its Groq 3 LPX rack has entered full production. The announcement came at the Hot Chips conference in Palo Alto, California. It is the first commercial product built from the Groq assets Nvidia acquired in December. That deal cost $20 billion and remains the largest purchase in the company's history.
The timing is not accidental. Nvidia reports quarterly earnings on Wednesday. Investors will be looking for evidence that the acquisition is converting into shipping hardware rather than roadmap slides.
Why decode latency became the bottleneck
Agentic systems produce text one token at a time. Each token waits on the one before it. Small delays compound across long chains of tool calls. Nvidia calls this the decode phase, and it is exactly where the LPX architecture is aimed.
Rubin GPUs still handle the heavy context processing. The LPX accelerators take over the latency-sensitive generation step. Nvidia frames the split as codesign rather than replacement.
Nvidia senior director Dion Harris told reporters the point is matching the right processor to the right part of the workload, not retiring the GPU.
Harris also pointed to a commercial angle. Cloud providers can package low-latency output as a premium service tier. That gives them a reason to buy specialized silicon instead of simply adding more general-purpose GPUs.
The numbers Nvidia is quoting
Nvidia packs 256 Groq 3 accelerators into a single rack. The chips connect over direct chip-to-chip links. Each die carries 500 megabytes of on-chip SRAM, the design choice that removes most memory-related stalls.
Samsung manufactures the Groq silicon. TSMC continues to build Nvidia's GPUs. Splitting fabrication across two foundries inside one rack-scale product is an unusual arrangement.
On performance, Nvidia cites a benchmark from Artificial Analysis. Running the open Gemma 4 31B model at a 100,000-token context, the rack delivered 3,400 output tokens per second. Nvidia says that is four times the nearest alternative platform.
Those figures come from a vendor-selected benchmark on a vendor-selected model. Independent results at other context lengths have not been published. Buyers pricing long-context agent workloads should read the 4x claim as a ceiling rather than a typical outcome.
Who is deploying first
Nebius is the first AI cloud to adopt the LPX rack. It will run alongside Vera CPUs and Rubin GPUs inside the company's Token Factory. Harris said those racks come online later this year.
CoreWeave is taking a different piece of the announcement. It has moved Spectrum-X Multiplane into production to connect Vera Rubin racks. That networking layer splits each server connection into parallel planes and scales to 512,000 GPUs without adding a third network tier.
SpaceXAI committed to Vera CPUs for its next generation of agentic AI. The company says it intends to run the architecture in orbit as well as on the ground.
A crowded low-latency field
Nvidia is not alone in this niche. AMD said earlier this year it would pair its rack-scale systems with Cerebras silicon for the same class of work. OpenAI's Ultrafast mode runs on Cerebras and advertises 750 tokens per second.
The competitive question has shifted. It is no longer who trains the largest model. It is who can serve tokens fast enough, at a price that survives an agent running for hours. Jensen Huang said in March that he would allocate a quarter of coding-oriented data center space to Groq chips. The rest, he said, stays Vera Rubin.






