A Microsoft research team has published a multi-agent coding harness that deletes the part most rivals build first: the orchestrator. In Agensh, released with code on GitHub, every worker runs the same five-step loop and the same prompt, differing only by worker ID. There is no lead session handing out work. The paper scales that design to 1,024 concurrent workers β and the numbers show both why the idea holds and where it stops paying.
Key takeaways
- On the five hardest ProgramBench tasks, mean final test-pass rate rose from 19.31% with one agent to 20.68%, 26.52% and 28.78% with 8, 32 and 128 agents, all using GPT-5.6-sol (high) under a six-hour budget.
- On pandoc alone, one agent scored 33.89%, 128 agents scored 50.94%, and 1,024 agents scored 55.06% β an eightfold increase in workers for 4.12 percentage points.
- Coordination runs on ordinary infrastructure: a Gitea repository as the shared workspace, Mattermost for messages, and an append-only log of typed findings.
How the harness replaces the orchestrator
Each worker gathers context on the shared goal and peer progress, claims a sub-task and announces its scope, acts on it with tools, verifies the result against acceptance criteria, then merges into the shared workspace and starts again. Overlapping claims are settled by the workers themselves through direct messages rather than by a planner.
Three services carry the coordination. The shared workspace is a Git platform β Gitea in the experiments β where workers hold private branches and integrate into main, so merge conflicts surface as ordinary conflicts. The message interface is Mattermost, with a team channel delivered at the start of each loop iteration and higher-priority direct messages delivered at the end of every infrastructure tool call. The shared context is an append-only store of typed entries: OBSERVED behavior, confirmed FACTs, failed approaches marked FAIL, active CLAIMs, and PATCH_SUMMARY records of completed changes, with a grep tool for history beyond recent memory.
The cooperation loop lives in the worker prompt, not the runtime, which is what lets the authors describe Agensh as layered above a single-agent harness rather than replacing it. Their runs used Copilot underneath; the paper argues Claude Code or others could slot in through a thin adapter. That is a pointed contrast with the harnesses it cites: Claude Code's agent teams still revolve around a fixed lead session, Codex routes sub-agent results back to a parent, and Kimi Agent Swarm trains an orchestrator to decompose and schedule work.
What the scaling curve actually shows
ProgramBench asks an agent organization to reproduce a reference program's behavior from scratch in six hours with internet access disabled. The team took the five hardest of its 200 instances β FFmpeg, GROMACS, pandoc, PHP-src and ctags β repositories running to thousands of files and millions of bytes of source.
Holding the model and harness fixed, 128 agents beat one agent by 9.47 percentage points, roughly a 49% relative gain. Larger groups also got there sooner: on pandoc, 128 agents cleared a 30% test-pass rate at the 30-minute checkpoint, 32 agents at 60 minutes and 8 agents at 90 minutes, while the single agent stayed below that line for the first two hours. For work under a hard deadline, that latency result may matter more than the final score.
The 1,024-agent run is the headline and the caveat. It adds 21.17 points over a single worker, but only 4.12 over 128. The paper reports no token or dollar accounting for any configuration, so the cost of the last 896 workers is left to the reader β and on a benchmark where the ceiling is a full behavioral reproduction, 55.06% is still a partial one.
The interesting part is the behavior, not the score
Because every worker gets an identical prompt, anything resembling structure has to emerge. The trajectories show it doing so. In GROMACS, workers announced a module interface and then implemented against it independently. In FFmpeg, one worker narrowed its scope after a peer flagged an overlap. In PHP-src, peers approved a contribution, another worker produced a counterexample, the approval was withdrawn, and the fix was reviewed again before merging.
At the largest scale, roles harden. On pandoc two workers invented an integration protocol β author tests a branch, sends the commit hash to a peer for validation and merge β which other workers then copied and later revised to hand the whole cycle to the reviewing peer. Multiple workers acted as integrators, and one worker contacted several candidates, took the first valid responder and cancelled the rest. That redundancy is what kept a failed worker from stalling the organization, a more useful property than any single number in the table.
It also raises the same governance question that other swarm results have: an organization that writes its own review protocol is harder to audit than one following a planner's instructions. Agensh makes the artifacts legible β commits, messages, typed findings β but nobody approved the workflow the agents settled on.
FAQ
Is Agensh open source?
The code is published at github.com/microsoft/Agensh and the paper carries arXiv's perpetual non-exclusive license. The harness depends on Gitea and Mattermost for its shared workspace and messaging, both of which are self-hostable.
Which model did the experiments use?
All reported runs used GPT-5.6-sol at high reasoning effort, with Copilot as the underlying single-agent harness and a six-hour budget per task. The authors say the cooperation loop is harness-agnostic, but they do not report results on other models.
Does more agents always mean better results?
Not proportionally. Scores rose at every step from 1 to 1,024 agents, but the gain from 128 to 1,024 was 4.12 percentage points on pandoc, against 17.05 points from 1 to 128. The paper presents agent count as a scaling dimension, not a free one.






