
New Benchmark Finds AI Agents Barely Improve Training Algorithms
AI4AI-Bench scored six agent systems at 0.166 for rewriting training algorithms. A second paper found seven ways self-improvement gains get miscounted.

AI4AI-Bench scored six agent systems at 0.166 for rewriting training algorithms. A second paper found seven ways self-improvement gains get miscounted.

DeepReinforce's Ornith-1.5 posts 86.1 on Terminal-Bench 2.1 against Claude Opus 4.8's 85.0, and ships a 9B variant small enough to run on a phone.

The IOL-AI Challenge had the official Linguistics Olympiad jury grade machine entries. Claude Opus 4.8 hit gold-medal marks; scale did not predict results.

Artificial Analysis scored Z.ai's GLM-5.3 at 60 on its Intelligence Index, well above the 35 median, at $4.40 per million output tokens. The catch is verbosity.

A new verifier re-tested 2,638 AI-generated GPU kernels already marked correct and found 39.5% broken beyond any tolerance argument.

Alibaba's Qwen team released Qwen3.8 as open weights, pairing a 2.4-trillion-parameter MoE flagship with a compact 27B vision-language model.

xAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol while holding pricing flat at $2/$6 per million tokens.

DeepSeek has moved its V4 Pro model to general availability, and the release is drawing attention less for raw capability than for what that capability now cost...

Google's Gemini 3.7 Flash lands three weeks after 3.6 Flash, scoring 65.3% on DeepSWE and shipping at half the price through the end of 2026.

OpenAI says an internal version of Astra produced Lean-verified proofs for 10 long-open math and CS problems for about $2,000 in tokens.

Alibaba's Qwen3.8 Max leads the Artificial Analysis agentic index, winning through long-horizon persistence rather than top reasoning scores.