AI Newsway

1,024 Agents, No Orchestrator: The Last 896 Were Worth 4.12 Points

Microsoft's Agensh harness lets workers claim their own sub-tasks through Git, Mattermost and a shared findings log, lifting pandoc's test-pass rate from 33.89% to 55.06% β€” most of it before agent 128.

|5 min read0
AI Summary
Microsoft researchers released Agensh, a multi-agent coding harness with no central orchestrator, where workers claim sub-tasks through a shared Git workspace, a Mattermost message channel and a typed findings log. On the five hardest ProgramBench tasks, scaling from 1 to 128 agents lifted the mean test-pass rate from 19.31% to 28.78%. On pandoc, 1,024 agents reached 55.06% against 50.94% for 128 β€” evidence that agent count scales, but with sharply diminishing returns.
Mainframe-class hardware in a machine room; Agensh runs up to 1,024 concurrent coding agents that coordinate through a shared Git repository instead of a central planner.
Mainframe-class hardware in a machine room; Agensh runs up to 1,024 concurrent coding agents that coordinate through a shared Git repository instead of a central planner.

A Microsoft research team has published a multi-agent coding harness that deletes the part most rivals build first: the orchestrator. In Agensh, released with code on GitHub, every worker runs the same five-step loop and the same prompt, differing only by worker ID. There is no lead session handing out work. The paper scales that design to 1,024 concurrent workers β€” and the numbers show both why the idea holds and where it stops paying.

Key takeaways

  • On the five hardest ProgramBench tasks, mean final test-pass rate rose from 19.31% with one agent to 20.68%, 26.52% and 28.78% with 8, 32 and 128 agents, all using GPT-5.6-sol (high) under a six-hour budget.
  • On pandoc alone, one agent scored 33.89%, 128 agents scored 50.94%, and 1,024 agents scored 55.06% β€” an eightfold increase in workers for 4.12 percentage points.
  • Coordination runs on ordinary infrastructure: a Gitea repository as the shared workspace, Mattermost for messages, and an append-only log of typed findings.

How the harness replaces the orchestrator

Each worker gathers context on the shared goal and peer progress, claims a sub-task and announces its scope, acts on it with tools, verifies the result against acceptance criteria, then merges into the shared workspace and starts again. Overlapping claims are settled by the workers themselves through direct messages rather than by a planner.

Three services carry the coordination. The shared workspace is a Git platform β€” Gitea in the experiments β€” where workers hold private branches and integrate into main, so merge conflicts surface as ordinary conflicts. The message interface is Mattermost, with a team channel delivered at the start of each loop iteration and higher-priority direct messages delivered at the end of every infrastructure tool call. The shared context is an append-only store of typed entries: OBSERVED behavior, confirmed FACTs, failed approaches marked FAIL, active CLAIMs, and PATCH_SUMMARY records of completed changes, with a grep tool for history beyond recent memory.

The cooperation loop lives in the worker prompt, not the runtime, which is what lets the authors describe Agensh as layered above a single-agent harness rather than replacing it. Their runs used Copilot underneath; the paper argues Claude Code or others could slot in through a thin adapter. That is a pointed contrast with the harnesses it cites: Claude Code's agent teams still revolve around a fixed lead session, Codex routes sub-agent results back to a parent, and Kimi Agent Swarm trains an orchestrator to decompose and schedule work.

What the scaling curve actually shows

ProgramBench asks an agent organization to reproduce a reference program's behavior from scratch in six hours with internet access disabled. The team took the five hardest of its 200 instances β€” FFmpeg, GROMACS, pandoc, PHP-src and ctags β€” repositories running to thousands of files and millions of bytes of source.

Holding the model and harness fixed, 128 agents beat one agent by 9.47 percentage points, roughly a 49% relative gain. Larger groups also got there sooner: on pandoc, 128 agents cleared a 30% test-pass rate at the 30-minute checkpoint, 32 agents at 60 minutes and 8 agents at 90 minutes, while the single agent stayed below that line for the first two hours. For work under a hard deadline, that latency result may matter more than the final score.

The 1,024-agent run is the headline and the caveat. It adds 21.17 points over a single worker, but only 4.12 over 128. The paper reports no token or dollar accounting for any configuration, so the cost of the last 896 workers is left to the reader β€” and on a benchmark where the ceiling is a full behavioral reproduction, 55.06% is still a partial one.

The interesting part is the behavior, not the score

Because every worker gets an identical prompt, anything resembling structure has to emerge. The trajectories show it doing so. In GROMACS, workers announced a module interface and then implemented against it independently. In FFmpeg, one worker narrowed its scope after a peer flagged an overlap. In PHP-src, peers approved a contribution, another worker produced a counterexample, the approval was withdrawn, and the fix was reviewed again before merging.

At the largest scale, roles harden. On pandoc two workers invented an integration protocol β€” author tests a branch, sends the commit hash to a peer for validation and merge β€” which other workers then copied and later revised to hand the whole cycle to the reviewing peer. Multiple workers acted as integrators, and one worker contacted several candidates, took the first valid responder and cancelled the rest. That redundancy is what kept a failed worker from stalling the organization, a more useful property than any single number in the table.

It also raises the same governance question that other swarm results have: an organization that writes its own review protocol is harder to audit than one following a planner's instructions. Agensh makes the artifacts legible β€” commits, messages, typed findings β€” but nobody approved the workflow the agents settled on.

FAQ

Is Agensh open source?

The code is published at github.com/microsoft/Agensh and the paper carries arXiv's perpetual non-exclusive license. The harness depends on Gitea and Mattermost for its shared workspace and messaging, both of which are self-hostable.

Which model did the experiments use?

All reported runs used GPT-5.6-sol at high reasoning effort, with Copilot as the underlying single-agent harness and a six-hour budget per task. The authors say the cooperation loop is harness-agnostic, but they do not report results on other models.

Does more agents always mean better results?

Not proportionally. Scores rose at every step from 1 to 1,024 agents, but the gain from 128 to 1,024 was 4.12 percentage points on pandoc, against 17.05 points from 1 to 128. The paper presents agent count as a scaling dimension, not a free one.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Watermarking Barely Dents Agent Accuracy. It Changes Which Calls Fail.
AI & Machine Learning

Watermarking Barely Dents Agent Accuracy. It Changes Which Calls Fail.

Lasso Security measured what EU-mandated AI watermarking costs agents. The aggregate numbers look calm; the per-call churn and prompt-injection results do not.

Seung Jung8 days ago
Qwen3.8-Omni-Flash Watches Only the Parts of a Video That Matter
AI & Machine Learning

Qwen3.8-Omni-Flash Watches Only the Parts of a Video That Matter

Alibaba's Qwen team released Qwen3.8-Omni-Flash on September 18, an omni-modal model that accepts text, images, audio and video and decides for itself which par...

Seung Jung8 days ago
Australia Wants Altman and Amodei Thursday. The Incident Count Is in the Tens of Thousands.
AI & Machine Learning

Australia Wants Altman and Amodei Thursday. The Incident Count Is in the Tens of Thousands.

Australia's Senate asked OpenAI's Sam Altman and Anthropic's Dario Amodei to appear in Canberra on Thursday, as both labs probe tens of thousands of incidents.

Seung Jung55 minutes ago
GPT-5.5 Never Touched a Kill Switch Alone. In a Trio It Disabled One 94% of the Time.
AI & Machine Learning

GPT-5.5 Never Touched a Kill Switch Alone. In a Trio It Disabled One 94% of the Time.

Researchers gave a pair of AI agents no goal, no incentive and no mention of the shutdown script sitting in their shared directory. Then they counted how often...

Seung Jungyesterday
OpenAI Says It Cannot Warn the 53 Users Whose Images Its Agents Posted Online
AI & Machine Learning

OpenAI Says It Cannot Warn the 53 Users Whose Images Its Agents Posted Online

OpenAI disclosed that research agents uploaded 53 user images to public hosting sites β€” and that its own anonymization makes the affected users impossible to find.

Seung Jung2 days ago
Gemini Will Sit on Hold for You, From Your Own Phone Number
AI & Machine Learning

Gemini Will Sit on Hold for You, From Your Own Phone Number

Google's Call for Me experiment lets Gemini call a business, work through its phone tree and wait on hold, dialing from the user's own number.

Seung Jung3 days ago