Prime Intellect has published the results of an unusually literal experiment: point 18 frontier models at the same open research task, let them run autonomously, and record everything. The task was the nanoGPT optimizer speedrun, where the human record stands at 2,600 and an untuned baseline sits at 3,290. Across 153 runs, no model reached the human mark.
The closest was Fable 5, driven through a claude-code harness at high effort, which reached 2,726 — 81.7% of the distance from baseline to the human record. It took 811 experiments and 8.7 days to get there. Opus 5 came next at 2,920, then Kimi K3 at 2,930, Opus 4.8 at 3,018 and GPT-5.6 Sol at 3,042.
Tokens Bought Very Different Amounts of Progress
The ranking is the least interesting column. Prime Intellect also logged total tokens, output tokens, tool calls, experiment counts and wall-clock days per run, and those numbers do not track the leaderboard at all.
GPT-5.6 Sol consumed roughly 2.9 billion tokens across 963 experiments and 28,000 tool calls to close 35.9% of the gap. Fable 5 closed more than twice as much of the gap on about 800 million. Opus 5 landed within striking distance of third place on 183 million tokens and 401 tool calls — a run one-fifteenth the size of Sol's that finished substantially further ahead.
At the frugal end, Grok 4.5 closed 24.6% on 46 million tokens and DeepSeek V4 Pro managed 12.3% on 26 million. Neither is competitive on the headline number, but per token spent they are not obviously worse buys than the systems above them. For anyone budgeting an autonomous research loop rather than admiring one, that spread is the finding.
The Same Model, Two Harnesses
Kimi K3 appears twice in the table, which makes it the cleanest natural experiment in the dataset. Run through Prime Intellect's own prime-agent scaffold at max effort, it reached 2,930 and closed 52.2% of the gap in 3.6 days on 112 million tokens. Run through kimi-code, the same weights reached only 2,974 and closed 45.8%, taking 5.1 days and 682 million tokens to do it.
Identical model, six percentage points and a six-fold difference in token spend. The scaffolding is not a rounding error.
Time Budget Changes the Answer
Prime Intellect also normalises results to a fixed budget, and the picture shifts sharply when it does. Measured at 24 hours rather than at each run's natural end, Fable 5 sits at 3,010 — barely ahead of Opus 5's 3,045 and well short of the 2,726 it eventually reached over 8.7 days.
In other words, most of the winning margin arrived after the first day. A team willing to leave an agent running for a week gets a materially different result from one that checks back the next morning, and leaderboard position depends heavily on which of those two questions is being asked.
Everything Is Open
Alongside the numbers, Prime Intellect released 41 curated agent trajectories covering full runs, including tool calls, subagent activity and scratchpads. That level of disclosure is rare in agent benchmarking, where results are usually reported as a single score with the process treated as proprietary.
It also invites scrutiny the leaderboard itself cannot settle. Several entries were still marked as running when the data was published, harnesses differ in maturity and instrumentation, and a single optimization task is a narrow window onto general research ability. What the dataset does establish is that the gap between the best and worst use of the same model is wide enough to matter more than the gap between adjacent models — and that the cost of finding out is now something buyers can inspect line by line.






