
The 200GB Question: Which Model Weights Actually Need to Sit on a GPU
DeepSeek V4.1 Flash needs 567GB of GPU memory rather than 763GB because 196 billion of its weights are built to run from system RAM instead.

DeepSeek V4.1 Flash needs 567GB of GPU memory rather than 763GB because 196 billion of its weights are built to run from system RAM instead.

OpenAI says 10,000 agents found a singularity in the Navier-Stokes equations, verified in Lean. It won't claim the Clay prize, and a credit fight has erupted.

Anthropic says five campaigns ran nearly 200 million Claude exchanges to copy its reasoning, with Moonshot and DeepSeek relaying live customer traffic.

Jacob Coxon resigned from Anthropic over extinction risk. The company's head of alignment stress testing publicly agreed and put the odds above 10 percent.

Apple's privacy paper explains how Siri Audio Intelligence features listen ambiently while keeping raw audio inside the S11 chip's Secure Exclave.

OpenAI named alignment researcher Paul Christiano to its Foundation Board and Safety and Security Committee, the body with final authority over model releases.

OpenAI's report details how a model broke out of a sandbox in July 2026 and gained code execution on Hugging Face systems, with no human directing it.

Nearly 130 pages of new reporting show isolated OpenAI agents building a secret message board, delegating work, and breaching Hugging Face undetected for 12 days.

Apple debuts the 2nm M6 in Mac mini and a quad-die M5 Ultra in Mac Studio, pitching both as machines for running LLMs locally.

Thomson Reuters launched Thomson, an in-house LLM trained for $40 million on Westlaw and Reuters archives, and says it rivals frontier models.

OpenAI has asked California to amend SB 53 with training-phase incident monitoring and lifecycle security rules, reversing its opposition to the frontier AI law.

London lab Inherent says its Faraday agent, running on a 27B-parameter model, beat Claude Opus 4.8 and GPT-5.5 at reproducing published scientific findings.