
The 200GB Question: Which Model Weights Actually Need to Sit on a GPU
DeepSeek V4.1 Flash needs 567GB of GPU memory rather than 763GB because 196 billion of its weights are built to run from system RAM instead.

DeepSeek V4.1 Flash needs 567GB of GPU memory rather than 763GB because 196 billion of its weights are built to run from system RAM instead.

Amazon is adding 2 million more Nvidia GPUs to AWS just five months after committing to 1 million, even as it scales its own Trainium and Graviton silicon.

Nvidia's Groq 3 LPX rack is in full production, commercializing its largest-ever acquisition and targeting the decode latency that slows AI agents.

SpaceX's $60 billion all-stock purchase of AI coding startup Cursor took effect August 14, giving the editor access to Musk's GPU fleet.

A new verifier re-tested 2,638 AI-generated GPU kernels already marked correct and found 39.5% broken beyond any tolerance argument.

Modular shipped Mojo 1.0 in release 26.5, promising API stability after three years of churn, with the compiler still set to be open-sourced in 2026.