A three-week experiment by CodeScene put a number on agentic legacy modernization that is hard to ignore: about $4,000 in tokens to carry 300,000 lines of C from a Code Health score of 5.6 to a flawless 10.0. The number travels well. The condition attached to it travels badly, and that condition is the more useful finding for anyone planning the same work.
That condition is an oracle. Every change the agents made was checked against a replay-trace harness that diffed the game's rollback state hash one frame at a time. The codebase under the knife was an open-source decompilation of Street Fighter III: 3rd Strike β chosen, per the CodeScene write-up, partly because the two engineers play it and could tell if it broke. A deterministic fighting game is one of the few kinds of software where that kind of verification is even possible.
Key takeaways
- Token spend for the uplift came to roughly $4,000, which CodeScene frames as half a monthly developer salary against a pre-agent estimate of 12 to 18 months of expert work.
- The safety net was behavioral, not test-based: a replay harness comparing rollback state hashes frame by frame, backed by a CodeHealth score the agents could optimize against.
- Figures circulating as results β 70% fewer AI-induced defects, 45% less token waste β are projections carried over from earlier CodeScene research and were not measured here.
What the $4,000 actually covered
The raw activity log is substantial: 2,903 commits, 726 files touched, 252,055 lines modified. Scope matters for reading those totals. This was two engineers working part-time, not a staffed team, and the spend covers tokens alone rather than salaries or review time.
Feedback came through an MCP server that put CodeScene's CodeHealth metric in reach of the agents, so a transformation could be scored instead of merely attempted. Pairing a deterministic quality number with a deterministic correctness check is what let the loop run largely unattended β remove either half and the throughput figure stops meaning much.
The playbook is the finding, not the score
A perfect metric reading is the least transferable part of this experiment. The durable artifact is a playbook of 22 recipes and 82 supporting notes that the agents assembled as they went, naming the transformation shapes they kept encountering and writing down the preconditions for each.
Some entries are straight out of the standard repertoire β Extract Function, Guard Clauses, Parameter Object. Three are unlikely to appear in any refactoring catalogue, because they describe this codebase specifically. Shared Index Range folds together loops whose only difference is where they start and stop. Action Parameter absorbs duplicated control structures that vary mainly in which function gets called. Uniform Step Table replaces a scatter of heterogeneous calls with table-driven dispatch.
Negative results were kept as well. Plenty of attempts left the metric untouched, and some pushed it down; those were written into the playbook next to the wins, which is closer to lab notebook discipline than to a code cleanup.
Model tiering showed up clearly in the run. Anthropic's Claude Opus handled the bulk after early experimentation, and the team reported Claude Code with Opus outperforming Codex with Sol at the specific job of capturing and documenting emerging patterns. Runs on the smaller Sonnet and Terra models tended to stall β files would plateau at what looked like a local optimum with code smells still in place.
Where the objections land
As InfoQ documented, the LinkedIn argument that followed was not about whether the run happened. It was about what it licenses anyone else to conclude. Supporters pointed out that trace replay clears a much higher bar than a green test suite, which certifies only that the tests still pass.
The scope challenges were pointed. One tech lead wanted to know whether the output was merged at all, and whether open-source game code stands in for production software carrying revenue. Daniel Webb, CTO at NeoSee and one of the two engineers, said it landed on main in a fork by way of 54 pull requests β a real merge, though not into an upstream project with outside maintainers.
Other critiques went after the vocabulary. Describing a codebase as perfect drew pushback on its own. So did the new recipes: if DRY is about duplicated knowledge rather than duplicated characters, then a recipe keyed on loop bounds may be collapsing a distinction worth keeping. And because Claude Code and Codex are each tuned to their own models, it is genuinely unclear how the credit splits between model capability and harness design. Architecture went unscored throughout, which leaves open the case where health metrics look clean and structural debt surfaces a year later.
What has not been measured
Adam Tornhill, who founded CodeScene, framed the run against thirty years of large-system work and called it his first encounter with superhuman AI performance at scale. He also stressed that automated tests and equivalence checks are absolutely essential safeguards β a caveat that quietly sets the entry price, since unhealthy codebases are typically unhealthy in part because they lack exactly those things.
The two most quotable figures in circulation are forecasts. Roughly 70% fewer AI-induced defects and roughly 45% less token waste come from extrapolating CodeScene's earlier Code Red research, which found healthy code 10x faster to evolve and carrying 15x fewer defects on average. Only the $4,000 and the three weeks were observed in this project.
Testing the rest is the point of a planned Lund University study, which will hand students both versions of the codebase β one at 5.6, one at 10.0 β and have them build features with frontier models to compare cost and quality directly. That is a meaningful check to run, particularly against evidence that LLM bug-fixers break working code more often than they mend broken code.
FAQ
Was the refactored code actually merged?
It was merged to main, but on a fork rather than into the upstream project. Daniel Webb put the total at 54 pull requests. Reviewers in the discussion argued that getting comparable work accepted by an active upstream project with independent maintainers would be a far harder test.
Does the $4,000 figure include engineering time?
It does not. That sum is token spend only, across three weeks of part-time effort from two engineers. CodeScene sets it beside half a monthly developer salary, and estimates the same modernization would have consumed 12 to 18 months of expert developer time in the pre-agent era.
Can this approach work on a typical legacy codebase?
The scoring half transfers anywhere; the verification half usually does not. Frame-by-frame replay depends on a deterministic game loop, and ordinary business systems rarely offer an equivalent. Absent some comparable correctness oracle, an agent can drive the quality metric upward with nothing confirming that behavior stayed intact.






