When a single AI agent tries to build a complex piece of software, it almost always loses its way. The agent holds the entire goal in its context window alongside the latest line of code, and as the project grows, the two pull against each other until the agent drifts, forgets, or starts over. Cursor, the company behind one of the most widely used AI coding editors, believes it has solved that problem not by giving agents more memory, but by dividing the work between two kinds of intelligence. Its new agent swarm architecture assigns a small number of powerful frontier models to the job of planning and a large fleet of cheaper, faster models to the job of execution. The result is a system that wrote a complete, working Rust implementation of SQLite from scratch in under four hours, using as little as one-fifteenth the cost of a solo frontier model, and producing code that was 85 percent more compact.
Why a lone agent fails on long-running tasks
The core insight behind Cursor’s swarm is that the difficulty of long-running software projects is not a reasoning problem but a context problem. A single agent, whether it is a frontier model like GPT-5.5 or a specialized coder, must traverse the entire task tree while keeping both the high-level goal and the immediate subtask in its active memory. The agent’s context window fills with irrelevant branches, partial solutions, and stale decisions. Over time, the agent begins to produce code that contradicts earlier choices, generates duplicate implementations, or simply stalls.
Cursor’s swarm breaks that cycle by splitting the agent into two distinct roles. Planner agents, which use the most powerful frontier models available, recursively break a large goal into smaller, independent tasks. Worker agents, which use cheaper and faster models, carry out those tasks without ever needing to see the full plan. The result is a task tree that adapts as work progresses, because planners can revise the tree based on what workers report, while workers never have to worry about the big picture.
In this architecture, planners do not write code, and workers do not plan. The separation is enforced at the system level, not by prompting. Cursor says this role split solves the context problem that has plagued every prior attempt to build autonomous software agents. A lone agent has to keep the entire tree in mind; a swarm spreads the tree across agents that each hold only a single node.
When Git breaks under agent speed
Cursor’s earlier browser swarm, which the company demoed in early 2026, reached about 1,000 commits per hour. That pace was already fast enough to strain Git. But the new swarm reaches 1,000 commits per second, a rate that creates failure modes that human teams never encounter. Cursor had to build its own version control system because Git simply could not handle the throughput.
The most dramatic failure was what Cursor calls “split-brain design.” Two planner agents, working independently, would unknowingly build the same idea in different parts of the codebase and implement it in different ways. When the agents did know about each other, a different problem emerged: contention. Planners would block each other with competing edits, each trying to lock the same file or function, and the system would grind to a halt.
Cursor solved the split-brain problem with shared design documents. Every time a planner made a significant decision, it recorded that decision in a document that all agents could read. Code tied to a decision linked back to the document through a reference that was checked at compile time. If two planners tried to implement the same feature, the compile-time reference would flag the conflict. When merge conflicts still occurred, a neutral agent stepped in and resolved them automatically.
Workers also learned to flag bloated files. Rather than try to fix a file that had grown too large, a worker would signal an external agent to split the file into smaller modules. Because agents had learned not to touch core code while working in existing codebases that include human developers, Cursor took the opposite approach in the swarm: it let agents break things on purpose. An agent could patch code outside its assigned area, and the compiler would carry the change through the entire system. The result was a fluid, self-healing codebase that could reorganize itself without waiting for a human to approve a refactor.
Multiple review angles and a self-maintained field guide
Reliability in a multi-agent system comes from diverse perspectives, not from perfect reasoning. Cursor tested several review approaches to catch errors before they propagated. One reviewer received the worker’s full transcript, including every prompt and output. Another reviewer saw only the final output. A third reviewer saw only the codebase, with no context about what the worker was trying to do. No single perspective caught everything, but the uncorrelated perspectives combined to produce high reliability. When one reviewer missed a bug, another caught it.
Cursor also introduced a “field guide,” a knowledge folder that agents themselves maintain. The field guide has a fixed line limit, so agents must be selective about what they record. Every agent receives the field guide’s contents at startup. Since model weights are frozen, the field guide captures surprising findings and useful shortcuts that later agents can exploit. Over time, the swarm becomes more efficient because it learns from its own mistakes without retraining the model.
The SQLite benchmark: a test of pure reasoning
Cursor gave the swarm a genuinely hard problem: build a Rust implementation of SQLite from scratch. The swarm received only the 835-page SQLite manual. Source code, test suites, the SQLite binary, and internet access were all withheld. The benchmark was sqllogictest, a test suite that contains millions of SQL queries and known answers. The swarm did not know the test suite existed.
Four configurations were tested against the old swarm architecture. The solo runs used GPT-5.5 and Grok 4.5. The hybrid runs used Opus 4.8 and Fable 5 as planners, with Cursor’s own Composer 2.5 as the worker model. The new system beat the old one in every configuration. After four hours, the new runs scored between 73 and 85 percent on sqllogictest, while the old runs scored between 11 and 77 percent. Every configuration of the new system later reached 100 percent.
The Grok 4.5 runs illustrated why the old architecture fell behind. The old swarm produced 68,000 commits in two hours, about 70 times as many as the new one. Most of that activity was wasted work. The old run accumulated more than 70,000 merge conflicts, while the new run stayed below 1,000 throughout the test. The most contested file in the old run recorded 7,771 conflicts from 1,173 agents, compared to 47 conflicts in the new run. The same split-brain problem appeared in the package structure. The old run broke the project into 54 Rust crates with three separate SQL packages. The new run settled on nine crates early and never changed its mind.
The code size differences were even more striking. In the Fable 5 configuration, the old swarm needed 64,305 lines of engine code, while the new one needed only 9,908 lines. In the Opus configuration, the old system produced 19,013 lines and scored 97 percent. The new system reached 100 percent with just 4,645 lines. At the same or better test scores, the new architecture cut codebase size by as much as 85 percent.
Why cheaper worker models created the largest savings
Total costs ranged from $1,339 for the Opus hybrid to $10,565 for GPT-5.5 running alone. Workers accounted for at least 69 percent of the tokens in every run and usually more than 90 percent. But planner tokens cost more, so the cost split looked very different. In the Opus hybrid, the planner produced only a small share of the tokens but accounted for two thirds of the total bill.
The worker model created the largest cost gap. In the GPT-5.5 solo run, workers alone cost $9,373. In the run using Opus and Composer, the entire worker fleet cost $411 at comparable quality. The difference comes down almost entirely to model pricing. Composer 2.5 benchmarks at the level of Opus 4.7 and GPT-5.5 but costs just $0.50 per million input tokens and $2.50 per million output tokens. The model is based on Kimi K2.5, according to Cursor founder Michael Truell.
Cursor argues that only a few parts of a large task need the intelligence of a frontier model, specifically task breakdown and key design decisions. Once a frontier planner resolves the ambiguity, cheaper models can follow its plan. However, the hybrid runs showed that planner quality still mattered. The Fable 5 planner used fewer planning tokens than Opus, but its workers needed far more tokens to finish the job. The Fable run cost more overall, even though the planner itself was cheaper. The lesson is that a good planner saves more than its own cost by keeping workers on track.
What is the most cost-effective way to run a coding swarm?
Based on Cursor’s experiments, the most cost-effective approach is to use a powerful frontier model as the planner and a very cheap but capable model as the worker. In the SQLite benchmark, the Opus planner with Composer 2.5 workers cost $1,339, which was the lowest total cost among all configurations that achieved 100 percent on the test suite. The GPT-5.5 solo run cost $10,565, and the Fable 5 solo run cost $20,057. The key is that the planner must be good enough to produce a clear, unambiguous plan; otherwise, the workers waste tokens figuring out what to do.
The swarm as a probabilistic compiler
Cursor describes swarms as a kind of probabilistic compiler that translates intent into executable work one step at a time. The company says accurately describing that intent was the main constraint in the experiment. When the plan is clear, even a cheap model can execute it. When the plan is vague, the whole system degrades. Cursor published the codebase from the Opus solo run as minisqlite on GitHub, so developers can see exactly what the swarm produced.
Runs like these are no longer limited to lab experiments. A prerelease version of Fable 5 handled most of Bun’s rewrite from Zig to Rust. Sixty-four instances wrote more than a million lines of code in 11 days for about $165,000. That is still a significant expense, but it is orders of magnitude cheaper than what a human team would cost to produce the same volume of code.
Production use today looks different. A study published in late 2025 found that 68 percent of agents used in production completed no more than ten steps before a human intervened. For 47 percent, the limit was fewer than five. The gap between what is possible in a controlled benchmark and what is safe in production is still wide. But Cursor’s swarm shows that the bottleneck is not the model’s ability to write code. It is the system’s ability to keep the work organized, the context small, and the conflicts resolved. Once that system is in place, cheaper models can handle the vast majority of the actual coding work, and the cost of building software with AI drops by an order of magnitude.