The economics of artificial intelligence in software development are undergoing a brutal, real-world stress test. As engineering teams rush to integrate agentic coding tools into their daily workflows, many are discovering a painful paradox: the very technology designed to accelerate output is also capable of devouring entire annual IT budgets in a matter of weeks. The soaring cost of tokens—the computational currency that powers large language models—has introduced a new, volatile variable into project planning. Yet, for a growing number of enterprises including Replit, Kilo Code, and Symbotic, this challenge is not a reason to retreat but a signal to refine their strategies. At the forefront of this transition, leaders from these three companies argue that the era of the human engineer as a primary writer of code is rapidly ending, replaced by a model where humans orchestrate, review, and manage fleets of AI agents. The central question is no longer whether AI coding agents work, but how to make them work without breaking the bank.
At Kilo Code, the shift is stark. According to co-founder Emilie Schario, engineers within the company are now actively reading or writing code themselves only about one percent of the time. The remaining 99 percent of the work is executed by autonomous agents. This fundamental change is forcing development teams to confront a new set of operational questions: Which parts of the system are safe to hand over to an AI? Who is responsible for cleaning up when a model makes a mistake? How do you support an architecture that uses models from multiple providers? And, perhaps most pressingly, do skyrocketing token bills represent genuine productivity gains or just a burned IT budget?
When 99 Percent of Coding Is Handed to Agents
The transformation described by Schario at VB Transform 2026 represents a generational shift in software engineering. At Kilo Code, the human role has evolved from craftsman to conductor. “Unless something’s really broken or debugging, 99% of the time engineers are not reading or writing code anymore,” Schario said. This is not a pilot program or a limited experiment; it is the operational reality for a company building agentic coding tools for the enterprise.
This level of delegation demands a high degree of trust, which Schario argues must be built on a foundation of clear guidance, shared skills, and robust protocols. The ability to give an agent a task and trust its output is not automatic. It requires a structured system. Model Context Protocol (MCP), a standard for how AI models access and use context, is a critical piece of this puzzle. By empowering models with the right context, companies can significantly improve the quality of agent output and reduce the need for human intervention.
For Jared Go, distinguished engineer for AI and cloud at Symbotic, a warehouse automation company, the current moment is about directing the focus of AI with precision. Rather than letting an agent run wild, Go provides strict criteria. “These are my criteria,” he explained. “Let’s look at it from the lens of security, elegance, clean, concise code, water tightness.” When the constraints are clear, AI can perform the heavy lifting, reducing the criticality of human code review.
However, Go acknowledges that human involvement becomes essential further down the line, particularly when product decisions are at stake. This is where the distinction between greenfield and brownfield development becomes crucial.
Greenfield Is Easy, Brownfield Is the Real Fight
The performance gap between creating new code and modifying existing code is one of the most significant technical barriers to mass AI agent adoption. “Greenfield is so easy for agents,” Go noted. Building a brand new codebase from scratch, with a clean slate and no legacy constraints, is a task where AI agents excel. The challenge, as any experienced engineer knows, lies in the other direction. “Brownfield we all know is where the actual challenge lies,” Go said.
Writing, updating, or maintaining existing code requires a deep understanding of context, historical decisions, and often, messy dependencies. Agents frequently struggle with this. They may introduce breaking changes, fail to understand the intent behind a legacy workaround, or simply generate code that does not fit the existing architecture. This is why Symbotic, despite heavy AI use, maintains a human loop for oversight.
Replit’s Approach: Human on the Loop, Not in the Loop
While Kilo Code and Symbotic have embraced agentic workflows, Replit has taken a more measured, but no less ambitious, path. Amol Jain, head of product engineering at Replit, describes a system that has “gone very agentic” but remains “conservative” in its approach to AI code review.
Replit’s internal tool is designed for a self-driving software engineering experience. Developers give a task to a fleet of agents, which then handle end-to-end planning, implementation, and testing. “It’s a fleet of agents that run in their own cloud virtual machines (VMs) with access controls behind token proxies so they’re secure,” Jain said.
The key innovation is in the review process. An agent reviews each pull request (PR) and assigns it a risk score. Low-risk PRs are considered safe for self-merging by their author. Higher-risk PRs are flagged for human reviewers, who read the code and provide feedback. “The idea was human on the loop, not human in the loop,” Jain explained. This distinction is critical. It means humans are not bottlenecks, but overseers who only step in when the system deems it necessary.
Jain shared a powerful example of the system in action. An engineer was struggling to reproduce and solve a “very gnarly bug” deep in Replit’s systems. The bug was sent to an AI manager agent, which made an unusual decision: it told the bug to go to sleep. The manager agent then spun up a group of underlying agents to investigate. These agents found the root cause. The manager then spun up a second group of agents that found a fix. Six hours later, the AI had a pull request ready for a bug that had puzzled human engineers for a considerable time.
The Rise of Multi-Model Architectures
One of the most significant shifts in the AI coding landscape is the move away from single-provider lock-in. As customers become more sophisticated, they are demanding the ability to choose the best model for each specific task. This trend is driving a new architectural paradigm: the multi-model gateway.
Kilo Code now supports over 500 models in its gateway, allowing its customers to pick and choose with unprecedented granularity. “Your software that you’re using to do agentic engineering should be decoupled from the model that you’re using to do it,” Schario said. This decoupling is the foundation of cost control and performance optimization.
A common workflow, according to Schario, involves using an expensive, frontier-tier model for the architecture and planning phase of a project. This is where reasoning and complex logic are most critical. Once the plan is set, the company switches to a much less expensive open-weight model to execute the bulk of the coding work.
The routing logic itself is becoming a sophisticated piece of engineering. It must factor in data retention policies, regional provider limitations, the specific keys a customer has brought, and the type of commits being made. “It’s factoring in what’s important to you, what limitations you’ve set, what data retention policies you’ve established, what keys you’ve brought in, what commits you might have … into that routing decision,” Schario said.
Replit takes a similar view, positioning itself as the decision-maker on behalf of the user. “We are essentially making the decisions on users’ behalf of what model to use when, in what capacity, to minimize cost and maximize capability,” Jain said. This approach requires deep insight into the cost-versus-capability spectrum of hundreds of models, a task that most individual enterprises are not equipped to handle.
What Is Tokenmaxxing and How Does It Burn Budgets?
A new term has entered the lexicon of engineering leadership: tokenmaxxing. It refers to the practice or phenomenon of consuming an extremely high volume of tokens, often on premium models, leading to runaway costs. For many enterprises, tokenmaxxing has become a management crisis.
Schario described hearing from customers who have accidentally spent their entire AI budget for the year in a short period. “I accidentally spent my whole AI budget for the year … so what do I do now?” is a question she has heard directly. The answer, for Kilo Code, is to guide them back to the multi-model workflow: use expensive models for planning, then switch to open-weight models for execution.
Visibility into usage is critical. Schario noted a particular engineer at her own company who has a “heavy foot” and is constantly at the top of the usage board. “I regularly have to nudge, ‘What are you doing there?’” she said. It is easy to look at a $600 daily bill and react with alarm, but the context is critical. When you consider the amount of work completed for that cost, the return can still be justified.
Symbotic took a structural approach to the problem. The company built an internal tool that gives managers visibility into pull requests and usage trends per employee. They then set per-month cost tiers. Managers can move users up or down a tier as they see fit. “Having a cap and seeing how many people went up in cap this month makes a big difference when you’re trying to corral these costs and make things efficient,” Go said.
A major shock to Symbotic’s cost structure came when Cursor, a coding tool the company uses heavily, ended a legacy discount. The company had been grandfathered into a flat per-request rate, even for frontier models. When Cursor moved everyone to full pricing, it forced a company-wide “reckoning on efficiency,” according to Go. Engineers began sharing tips on which models worked best for specific tasks, like C# code, and which were a waste of money.
The Right Metric: Cost Per Pull Request
Schario argues that the traditional metrics of developer productivity are insufficient in the age of AI agents. Instead, she is focusing on a new key performance indicator. “Cost per pull request is the metric that I’m paying attention to right now,” Schario said. “It feels like the closest proximity for how I can measure value.” This metric directly ties the expense of tokens to the output of finished, reviewed code. It transforms the cost conversation from a raw spend number to a value-per-unit calculation.
“The spend with no return on that spend is the problem,” Schario said. The challenge for enterprises is to ensure that every token burned is yielding a measurable return in terms of features built, bugs fixed, or code refactored.
When the Cost Problem Moves Out of Engineering
The challenge of runaway AI costs is not confined to engineering teams. As platforms like Replit broaden access to agents across the organization, the risk of unexpected spending spreads. Jain recalled a case where a user on the support side of the business “blew through an insane amount of money.” When the team investigated, they discovered the support agent had been running an automation on a premium model—Jain referred to it as “GPT 5.5 Pro Max”—which was far more powerful and expensive than the task required.
This case highlights a growing governance problem. “At least till that point, the ROI was rather clear,” Jain said. “We could see engineering productivity 3X, so no one had questioned it yet.” When AI agents are used outside of engineering, the ROI is often less visible and the cost discipline is weaker.
To prevent this, Jain emphasized the need for sensible defaults, model routing, and visibility tools that are “not anti-productive.” The most critical rule is simple. “Most tasks do not need the frontier,” he said. Defaulting to the cheapest capable model, rather than the most powerful available model, is the single most effective way to control costs.
The era of unlimited, unmonitored AI spending is over. Enterprises are learning that AI coding agents are not a magic bullet but a powerful new tool that requires careful financial and operational governance. The companies that succeed will be those that treat agentic AI as a managed portfolio of compute resources, optimizing for cost per pull request and routing work to the right model at the right time. The future of software engineering is not about humans writing less code, but about humans becoming far more strategic about how they spend their AI budget.