
Mercury 2.5 Ships at 1,107 Tokens a Second and Four Cents a Million
A phone agent startup says Mercury cut its worst-case response from minutes to one second. Inception Labs' new diffusion model is built around that kind of number.

A phone agent startup says Mercury cut its worst-case response from minutes to one second. Inception Labs' new diffusion model is built around that kind of number.

Nvidia and Palantir published an 86.7% versus 55.5% accuracy gap favoring a 30B fine-tuned model over a 550B general one on Nvidia's own materials allocation.

Antitrust enforcers have opened their first real examination of the deal structure that has replaced the acquisition in AI: Nvidia has received a formal Justice...

Jensen Huang reaffirmed Nvidia's 70% growth guidance, implying $680 billion in revenue, and answered circular-deal critics with '$1 in, $100 back.'

Amazon is adding 2 million more Nvidia GPUs to AWS just five months after committing to 1 million, even as it scales its own Trainium and Graviton silicon.

Nvidia has reportedly agreed to acquire Hugging Face for $12.9 billion, a multiple of roughly 80x revenue that buys the developer graph, not the income statement.

AI-generated kernels beat OpenAI expert-written versions by up to 1.8x, as Jalapeño posts its first InferenceX benchmark results.

Perplexity and Nvidia shipped Portable Computer, an agent running Qwen or PPLX 27B on local RTX hardware with no metered tokens until the user allows a cloud step.

Nvidia's Groq 3 LPX rack is in full production, commercializing its largest-ever acquisition and targeting the decode latency that slows AI agents.

The Series A extension doubles Starcloud to a $2.3B valuation, but Falcon 9 winds down in 2028 and Starship has yet to prove rapid reuse.

Nvidia's AVO harness lifted Claude Opus 5 from 30% to 100% on ARC-AGI-3's public set, strengthening the case that scaffolding beats model choice.

Etched raised $700M at a $21B valuation led by Jane Street, doubling its July price. The startup splits inference into prefill and decode with separate silicon.