AI Newsway

Nvidia Puts Its $20 Billion Groq Bet Into Full Production

The LPX rack targets decode latency, the bottleneck that makes AI agents feel slow

|3 min read0
AI Summary
Nvidia announced at Hot Chips on Monday that its Groq 3 LPX rack is in full production, the first commercial product from the $20 billion Groq acquisition closed in December. Each rack holds 256 Samsung-built Groq 3 accelerators with 500MB of on-chip SRAM, handling the latency-sensitive decode phase while Rubin GPUs process context, and Nvidia cites 3,400 output tokens per second on Gemma 4 31B at 100,000-token context. Nebius is the first cloud to deploy it later this year.
Rack-scale data center infrastructure of the kind Nvidia's Groq 3 LPX systems are designed to slot into alongside Vera Rubin NVL72 deployments
Rack-scale data center infrastructure of the kind Nvidia's Groq 3 LPX systems are designed to slot into alongside Vera Rubin NVL72 deployments

Nvidia said Monday that its Groq 3 LPX rack has entered full production. The announcement came at the Hot Chips conference in Palo Alto, California. It is the first commercial product built from the Groq assets Nvidia acquired in December. That deal cost $20 billion and remains the largest purchase in the company's history.

The timing is not accidental. Nvidia reports quarterly earnings on Wednesday. Investors will be looking for evidence that the acquisition is converting into shipping hardware rather than roadmap slides.

Why decode latency became the bottleneck

Agentic systems produce text one token at a time. Each token waits on the one before it. Small delays compound across long chains of tool calls. Nvidia calls this the decode phase, and it is exactly where the LPX architecture is aimed.

Rubin GPUs still handle the heavy context processing. The LPX accelerators take over the latency-sensitive generation step. Nvidia frames the split as codesign rather than replacement.

Nvidia senior director Dion Harris told reporters the point is matching the right processor to the right part of the workload, not retiring the GPU.

Harris also pointed to a commercial angle. Cloud providers can package low-latency output as a premium service tier. That gives them a reason to buy specialized silicon instead of simply adding more general-purpose GPUs.

The numbers Nvidia is quoting

Nvidia packs 256 Groq 3 accelerators into a single rack. The chips connect over direct chip-to-chip links. Each die carries 500 megabytes of on-chip SRAM, the design choice that removes most memory-related stalls.

Samsung manufactures the Groq silicon. TSMC continues to build Nvidia's GPUs. Splitting fabrication across two foundries inside one rack-scale product is an unusual arrangement.

On performance, Nvidia cites a benchmark from Artificial Analysis. Running the open Gemma 4 31B model at a 100,000-token context, the rack delivered 3,400 output tokens per second. Nvidia says that is four times the nearest alternative platform.

Those figures come from a vendor-selected benchmark on a vendor-selected model. Independent results at other context lengths have not been published. Buyers pricing long-context agent workloads should read the 4x claim as a ceiling rather than a typical outcome.

Who is deploying first

Nebius is the first AI cloud to adopt the LPX rack. It will run alongside Vera CPUs and Rubin GPUs inside the company's Token Factory. Harris said those racks come online later this year.

CoreWeave is taking a different piece of the announcement. It has moved Spectrum-X Multiplane into production to connect Vera Rubin racks. That networking layer splits each server connection into parallel planes and scales to 512,000 GPUs without adding a third network tier.

SpaceXAI committed to Vera CPUs for its next generation of agentic AI. The company says it intends to run the architecture in orbit as well as on the ground.

A crowded low-latency field

Nvidia is not alone in this niche. AMD said earlier this year it would pair its rack-scale systems with Cerebras silicon for the same class of work. OpenAI's Ultrafast mode runs on Cerebras and advertises 750 tokens per second.

The competitive question has shifted. It is no longer who trains the largest model. It is who can serve tokens fast enough, at a price that survives an agent running for hours. Jensen Huang said in March that he would allocate a quarter of coding-oriented data center space to Groq chips. The rest, he said, stays Vera Rubin.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

OpenAI Says Its Own Models Helped Tape Out Jalapeño in Nine Months
Tech & Business

OpenAI Says Its Own Models Helped Tape Out Jalapeño in Nine Months

AI-generated kernels beat OpenAI expert-written versions by up to 1.8x, as Jalapeño posts its first InferenceX benchmark results.

Seung Jung22 days ago
Amazon Triples Its Nvidia Order to 2 Million GPUs
Tech & Business

Amazon Triples Its Nvidia Order to 2 Million GPUs

Amazon is adding 2 million more Nvidia GPUs to AWS just five months after committing to 1 million, even as it scales its own Trainium and Graviton silicon.

Seung Jung21 days ago
Nvidia Spent $27 Billion Without Filing a Single Merger Notice. The DOJ Wants to Know Why.
Tech & Business

Nvidia Spent $27 Billion Without Filing a Single Merger Notice. The DOJ Wants to Know Why.

Antitrust enforcers have opened their first real examination of the deal structure that has replaced the acquisition in AI: Nvidia has received a formal Justice...

Seung Jung4 days ago
Huang Defends Nvidia's 70% Growth Call: 'We Put In $1 and $100 Comes Back'
Tech & Business

Huang Defends Nvidia's 70% Growth Call: 'We Put In $1 and $100 Comes Back'

Jensen Huang reaffirmed Nvidia's 70% growth guidance, implying $680 billion in revenue, and answered circular-deal critics with '$1 in, $100 back.'

Seung Jung6 days ago
NVIDIA Makes Tokens Per Megawatt the Metric, and Lambda Puts a Number on It
Tech & Business

NVIDIA Makes Tokens Per Megawatt the Metric, and Lambda Puts a Number on It

At AI Infra Summit, NVIDIA reframed AI infrastructure around tokens per megawatt, backed by Lambda's 24% throughput gain and new agentic benchmark results.

Seung Jung17 hours ago
Starcloud Raised $250M for Orbital AI Compute. Its Hardest Problem Is Booking a Rocket.
Tech & Business

Starcloud Raised $250M for Orbital AI Compute. Its Hardest Problem Is Booking a Rocket.

The Series A extension doubles Starcloud to a $2.3B valuation, but Falcon 9 winds down in 2028 and Starship has yet to prove rapid reuse.

Seung Jung25 days ago