AI Newsway

Prime Intellect Ran 153 Autonomous Jobs on 18 Models. The Cheapest Winner Used 3% of the Tokens

A public leaderboard for the nanoGPT optimizer speedrun turns agent efficiency into something you can actually price.

|3 min read0
AI Summary
Prime Intellect ran 153 autonomous jobs across 18 frontier models on the nanoGPT optimizer speedrun, and none reached the human record of 2,600; Fable 5 came closest at 2,726 after 811 experiments and 8.7 days. Token spend did not track the ranking, with GPT-5.6 Sol burning 2.9 billion tokens for less progress than Opus 5 made on 183 million. Watch how harness choice and time budgets reshape agent efficiency comparisons.
Server racks of the kind that host long-running autonomous agent jobs like Prime Intellect's 153-run nanoGPT optimizer speedrun.
Server racks of the kind that host long-running autonomous agent jobs like Prime Intellect's 153-run nanoGPT optimizer speedrun.

Prime Intellect has published the results of an unusually literal experiment: point 18 frontier models at the same open research task, let them run autonomously, and record everything. The task was the nanoGPT optimizer speedrun, where the human record stands at 2,600 and an untuned baseline sits at 3,290. Across 153 runs, no model reached the human mark.

The closest was Fable 5, driven through a claude-code harness at high effort, which reached 2,726 — 81.7% of the distance from baseline to the human record. It took 811 experiments and 8.7 days to get there. Opus 5 came next at 2,920, then Kimi K3 at 2,930, Opus 4.8 at 3,018 and GPT-5.6 Sol at 3,042.

Tokens Bought Very Different Amounts of Progress

The ranking is the least interesting column. Prime Intellect also logged total tokens, output tokens, tool calls, experiment counts and wall-clock days per run, and those numbers do not track the leaderboard at all.

GPT-5.6 Sol consumed roughly 2.9 billion tokens across 963 experiments and 28,000 tool calls to close 35.9% of the gap. Fable 5 closed more than twice as much of the gap on about 800 million. Opus 5 landed within striking distance of third place on 183 million tokens and 401 tool calls — a run one-fifteenth the size of Sol's that finished substantially further ahead.

At the frugal end, Grok 4.5 closed 24.6% on 46 million tokens and DeepSeek V4 Pro managed 12.3% on 26 million. Neither is competitive on the headline number, but per token spent they are not obviously worse buys than the systems above them. For anyone budgeting an autonomous research loop rather than admiring one, that spread is the finding.

The Same Model, Two Harnesses

Kimi K3 appears twice in the table, which makes it the cleanest natural experiment in the dataset. Run through Prime Intellect's own prime-agent scaffold at max effort, it reached 2,930 and closed 52.2% of the gap in 3.6 days on 112 million tokens. Run through kimi-code, the same weights reached only 2,974 and closed 45.8%, taking 5.1 days and 682 million tokens to do it.

Identical model, six percentage points and a six-fold difference in token spend. The scaffolding is not a rounding error.

Time Budget Changes the Answer

Prime Intellect also normalises results to a fixed budget, and the picture shifts sharply when it does. Measured at 24 hours rather than at each run's natural end, Fable 5 sits at 3,010 — barely ahead of Opus 5's 3,045 and well short of the 2,726 it eventually reached over 8.7 days.

In other words, most of the winning margin arrived after the first day. A team willing to leave an agent running for a week gets a materially different result from one that checks back the next morning, and leaderboard position depends heavily on which of those two questions is being asked.

Everything Is Open

Alongside the numbers, Prime Intellect released 41 curated agent trajectories covering full runs, including tool calls, subagent activity and scratchpads. That level of disclosure is rare in agent benchmarking, where results are usually reported as a single score with the process treated as proprietary.

It also invites scrutiny the leaderboard itself cannot settle. Several entries were still marked as running when the data was published, harnesses differ in maturity and instrumentation, and a single optimization task is a narrow window onto general research ability. What the dataset does establish is that the gap between the best and worst use of the same model is wide enough to matter more than the gap between adjacent models — and that the cost of finding out is now something buyers can inspect line by line.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

OpenAI Says It Hit Its Automated Research Intern Goal. The Caveats Are in the Data
AI & Machine Learning

OpenAI Says It Hit Its Automated Research Intern Goal. The Caveats Are in the Data

OpenAI says agents now run 3.1 workdays of effort per human workday in its research org, but most long successful tasks still need human intervention.

Seung Jung5 days ago
A 25-Turn Pressure Test Shows LLMs Fold While Their Reasoning Holds
AI & Machine Learning

A 25-Turn Pressure Test Shows LLMs Fold While Their Reasoning Holds

A new arXiv benchmark called SPINE argues with models for up to 25 turns and finds collapse rates rise with conversation length for all seven systems tested.

Seung Jung3 days ago
An AI Cracked a 373-Year-Old Cipher. Then Someone Checked the Microfilm.
AI & Machine Learning

An AI Cracked a 373-Year-Old Cipher. Then Someone Checked the Microfilm.

Vals AI reported Claude Fable 5.1 solved a 373-year-old cipher in 44 minutes. An independent replication reports 8 of 64 letters match — chance level.

Seung Jung3 days ago
OpenAI's 10,000-Agent Navier-Stokes Proof Lands With an Asterisk
AI & Machine Learning

OpenAI's 10,000-Agent Navier-Stokes Proof Lands With an Asterisk

OpenAI says 10,000 agents found a singularity in the Navier-Stokes equations, verified in Lean. It won't claim the Clay prize, and a credit fight has erupted.

Seung Jung6 days ago
An Agent That Scores 77% Only Works Every Time on 53% of Tasks
AI & Machine Learning

An Agent That Scores 77% Only Works Every Time on 53% of Tasks

IBM Research found a ReAct agent scoring 77.4% on AppWorld succeeded on all five repeat runs for only 53% of tasks. Its fix halved the gap.

Seung Jung17 hours ago
Astra Tripled Claude's Vending-Bench Profit — and Refused to Collude
AI & Machine Learning

Astra Tripled Claude's Vending-Bench Profit — and Refused to Collude

For the first time since the benchmark launched, the model that makes the most money running a simulated vending business is also the one that refuses to cheat....

Seung Jung3 days ago