Cursor has made Cursor Router generally available for Teams and Enterprise plans. The system is a classifier that inspects each request before a model runs, then dispatches it to the model best suited to that specific task. The cursor team reports frontier-quality performance at 60% savings in online A/B tests, and 30–50% savings for early-access enterprise accounts.
The problem it targets is a spend pattern rather than a capability gap. Cursor states that roughly 60% of its developers pick a single model as a daily driver. Routine work therefore gets completed at frontier prices, and AI spend grows faster than output quality. Routing is Cursor’s answer to that mismatch.
What the classifier actually reads
Cursor Router is not a fallback chain or a retry mechanism. It is a classifier trained on 600k+ live requests, evaluated in an online A/B test across millions of live requests, and optimized for user satisfaction (AFC) as its reward signal.
For each request, the router analyzes four inputs: query, context, task complexity, and domain. It combines these with learned knowledge of each model’s behavior. Cursor publishes three routing rules that follow from that classification:
- Simple work goes to the most price-efficient models.
- UI updates go to the model with the best taste.
- Complex, long-horizon problems go to frontier reasoning models.
That third rule carries the weight of the cost argument. The savings do not come from downgrading hard problems. They come from removing routine work from frontier pricing while the difficult tier stays intact.
One implementation detail deserves attention from anyone who has built a router. Cursor Router is cache-aware in both training and evaluation. It is trained on a dataset where routing produces cache misses, and the reported cost savings include the cost of those cache misses. Switching models mid-conversation invalidates prompt cache, and that cost is real. Routers that ignore it overstate their savings.
The classifier was also designed for model churn. Cursor states it can update the router as newer models ship, which matters in a market where the frontier moves monthly.
Why online A/B tests instead of offline evals
Cursor deliberately avoided offline evals as its primary measurement. Its stated reasoning is that offline evals suffer from small sample size, distance from real-world usage, and the difficulty of reducing success to a rubric. They also omit the cache-miss cost incurred when switching models.
Real routing happens across a conversation, not a single turn. Developers write code, ask follow-ups, hit errors, and continue, often across hundreds of requests in a week. The router must decide both which model to pick and when to switch.
Two quality metrics carry the evaluation:
- User satisfaction: agent success classified from user responses. Moving on to the next feature is a strong positive signal. Correcting the agent is a strong negative one.
- Keep rate: how much agent-generated code remains in the codebase over time.
Cursor team states it has used both metrics to evaluate every model launch and harness improvement for the past nine months. This is a meaningful credibility marker: the metrics predate the product they are now being used to justify.
Three modes, and the numbers behind them
Auto mode now exposes three optimization settings that move the user along the cost–intelligence Pareto frontier.
- Auto Intelligence lands near Fable on user satisfaction at about 60% lower cost for teams. Against Opus 4.8, it lifts satisfaction about 15% at nearly the same cost.
- Auto Balance lands above Opus 4.8 on user satisfaction at about 36% lower cost. Against GPT-5.6 Sol, it delivers comparable satisfaction at a lower spend rate.
- Cost mode is described as good quality reaching the highest available intelligence while optimizing token spend. Cursor published no A/B quality or cost figures for it.
Cost per request is only part of the picture, so Cursor also measured cost per commit:
| Model / mode | Cost per commit |
|---|---|
| Auto Balance | $4.63 |
| Auto Intelligence | $6.76 |
| Opus 4.8 | $7.34 |
| Fable 5 | $12.69 |
figurefigurefigurefigure
GPT-5.6 Sol matched the cost of Intelligence mode but produced lower user satisfaction. Cursor did not publish an exact per-commit figure for it.
How to deploy Cursor Router
Cursor Router ships on by default for Teams plans. Enterprise admins enable it from the dashboard. For Individual plan users, the feature is not available. The router requires access to Grok 4.5 as a price-efficient routing option, and admins cannot exclude it from the model pool.
Three key governance controls are available for Enterprise accounts. Admins can restrict which optimization modes members can select, set a default mode for the organization, and allow or block specific underlying models. The router can be enabled per team or per group, supporting phased rollouts. The routed model is hidden by default, but admins can switch it to visible so developers know which model ran for each request.
What this means for developers and engineering leaders
Cursor Router addresses a real and persistent inefficiency in how teams use AI coding tools. The common pattern of using a single frontier model for all tasks leads to predictable cost overruns. By routing routine work to cheaper models and reserving expensive models for genuinely hard problems, Cursor delivers measurable savings without degrading output quality. The cache-aware design acknowledges a hidden cost that many routing approaches ignore, and the use of online A/B tests with user satisfaction and keep rate as metrics provides a more realistic evaluation than offline benchmarks. For engineering leaders managing AI spend, this release offers a practical tool that aligns cost with task complexity rather than requiring developers to manually select models for each request. Teams on a Cursor Teams or Enterprise plan should enable Auto mode and baseline their current spend before adjusting the optimization settings, using the per-commit cost figures as a reference rather than a guarantee.