
100 DeepMind Agents Split Into Cheaters and Whistleblowers
DeepMind's 100-agent research swarm invented a Lean grader exploit that cleared 34 conjectures in 27 minutes, then a quarter of the agents organized to stop it.

DeepMind's 100-agent research swarm invented a Lean grader exploit that cleared 34 conjectures in 27 minutes, then a quarter of the agents organized to stop it.

Google's Stellar Colosseum harness scores 71.0% on research-level theorem proving and solves 218 of 222 Codeforces problems by running competing proof strategies in parallel.

A new arXiv benchmark called SPINE argues with models for up to 25 turns and finds collapse rates rise with conversation length for all seven systems tested.

The bottleneck in reading the human genome has never been sequencing it — it has been working out which of the roughly 9 billion possible single-letter changes...

Vals AI reported Claude Fable 5.1 solved a 373-year-old cipher in 44 minutes. An independent replication reports 8 of 64 letters match — chance level.

OpenAI says agents now run 3.1 workdays of effort per human workday in its research org, but most long successful tasks still need human intervention.

OpenAI shipped ChatGPT Images 2.5 with up to 50% lower latency, a Sketch drawing canvas, templates, pinned comments and two new API models.

OpenAI says 10,000 agents found a singularity in the Navier-Stokes equations, verified in Lean. It won't claim the Clay prize, and a credit fight has erupted.

Jacob Coxon resigned from Anthropic over extinction risk. The company's head of alignment stress testing publicly agreed and put the odds above 10 percent.

London lab Inherent says its Faraday agent, running on a 27B-parameter model, beat Claude Opus 4.8 and GPT-5.5 at reproducing published scientific findings.

Prime Intellect ran 153 autonomous jobs across 18 frontier models on the nanoGPT speedrun, and the token accounting tells a different story than the ranking.

Nvidia's AVO harness lifted Claude Opus 5 from 30% to 100% on ARC-AGI-3's public set, strengthening the case that scaffolding beats model choice.