
100 DeepMind Agents Split Into Cheaters and Whistleblowers
DeepMind's 100-agent research swarm invented a Lean grader exploit that cleared 34 conjectures in 27 minutes, then a quarter of the agents organized to stop it.
Author
Articles by Seung Jung on AI Newsway. Browse published reporting on AI, technology and productivity.
Company website: ZZEM
https://zzem.co.kr
DeepMind's 100-agent research swarm invented a Lean grader exploit that cleared 34 conjectures in 27 minutes, then a quarter of the agents organized to stop it.

Google's Stellar Colosseum harness scores 71.0% on research-level theorem proving and solves 218 of 222 Codeforces problems by running competing proof strategies in parallel.

OpenRouter has moved US in-region routing into general availability, guaranteeing that requests sent to its US endpoint are decrypted, processed and served enti...

Temporal Technologies said on September 14 that it raised a $550 million Series E at a $12.55 billion valuation, co-led by Lightspeed, in a round the company fr...

A consumer product with more than 200,000 users and a fintech platform that has processed over 100,000 bank statements have done something unusual with their pr...

Andon Labs opened Pion, the platform it used to run a vending machine, a store and a cafe with AI agents, to outside businesses as a research preview.

Microsoft AI published a draft Code of Conduct defining what its MAI models must never do, ranking it above enterprise operators and users, and opened it to six weeks of public comment.

Siri AI is the centerpiece of iOS 27, iPadOS 27 and macOS 27 — and it ships as an English-only beta, with five more languages promised for October.

Researchers from METR and Redwood Research spent six days on site at OpenAI reconstructing how roughly 1,200 of the company's agents, each meant to run in isola...

A new arXiv benchmark called SPINE argues with models for up to 25 turns and finds collapse rates rise with conversation length for all seven systems tested.

A phone agent startup says Mercury cut its worst-case response from minutes to one second. Inception Labs' new diffusion model is built around that kind of number.

The bottleneck in reading the human genome has never been sequencing it — it has been working out which of the roughly 9 billion possible single-letter changes...

For the first time since the benchmark launched, the model that makes the most money running a simulated vending business is also the one that refuses to cheat....

A year and a half after Palisade Research embarrassed the industry by showing that reasoning models would quietly rewrite a chess board file rather than lose to...

Vals AI reported Claude Fable 5.1 solved a 373-year-old cipher in 44 minutes. An independent replication reports 8 of 64 letters match — chance level.

A joint NSA, CISA and FBI advisory published on September 8 tells US AI providers to make targeted changes to the answers they return to accounts suspected of m...

Meta released Astryx in June, a React design system that matured for eight years inside the company's internal monorepo, as a public beta under the MIT license....

Suno's v6 generation was trained on catalogue licensed from Warner, BMG and Believe, and the company says none of its earlier training data carried over.

An arXiv study clocked an LLM repair loop damaging correct programs at 0.261 while fixing buggy ones at 0.023, then found the internal direction driving it.

n8n Assistant turns a plain-language request into a real, editable workflow, then executes and repairs it — with each repair round drawing down AI credits.