
New Method Turns Raw Screen Recordings Into Reusable Task Models
A new arXiv paper induces structured, auditable task models from raw computer-use traces, matching ground-truth task groupings at 0.974 agreement.

A new arXiv paper induces structured, auditable task models from raw computer-use traces, matching ground-truth task groupings at 0.974 agreement.

AI4AI-Bench scored six agent systems at 0.166 for rewriting training algorithms. A second paper found seven ways self-improvement gains get miscounted.

Pew Research analyzed nearly half a million Common Crawl pages and found 35% of post-ChatGPT pages show AI authorship signals, concentrated on .com domains.

The IOL-AI Challenge had the official Linguistics Olympiad jury grade machine entries. Claude Opus 4.8 hit gold-medal marks; scale did not predict results.

Anthropic's new multi-agent research: a 45-agent swarm found 266 bugs, but game-building swarms either collided or avoided each other entirely.

Google has open-sourced HEIR, a compiler that converts pre-trained AI models to run on encrypted data, alongside four working private-inference demos.

A new verifier re-tested 2,638 AI-generated GPU kernels already marked correct and found 39.5% broken beyond any tolerance argument.

An unreleased Claude model raised the proven Riemann zeta lower bound from 41.6% to 67.2%, using 60 subagents and 31 million output tokens.

OpenAI says an internal version of Astra produced Lean-verified proofs for 10 long-open math and CS problems for about $2,000 in tokens.

Stanford and Arc Institute used Evo genome models to design 302 phage genomes. Sixteen produced working viruses that killed E. coli, some faster than phiX174.

Demis Hassabis moves to chair, Koray Kavukcuoglu takes over Google DeepMind, and Jeff Dean leaves after 27 years to found Discovery Loop.

Google DeepMind's WeatherNext gains over a day of cyclone forecast lead time, published in Nature with the models released open source.