{"id":75069,"date":"2026-08-04T23:43:47","date_gmt":"2026-08-05T03:43:47","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=75069"},"modified":"2026-08-04T23:43:47","modified_gmt":"2026-08-05T03:43:47","slug":"ai-coding-agent-costs","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/ai-coding-agent-costs\/","title":{"rendered":"AI coding agents blow through budgets; Replit, Kilo Code, Symbotic manage"},"content":{"rendered":"<p>The economics of artificial intelligence in software development are undergoing a brutal, real-world stress test. As engineering teams rush to integrate agentic coding tools into their daily workflows, many are discovering a painful paradox: the very technology designed to accelerate output is also capable of devouring entire annual IT budgets in a matter of weeks. The soaring cost of tokens\u2014the computational currency that powers large language models\u2014has introduced a new, volatile variable into project planning. Yet, for a growing number of enterprises including <a href=\"https:\/\/replit.com\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Replit<\/a>, Kilo Code, and <a href=\"https:\/\/www.symbotic.com\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Symbotic<\/a>, this challenge is not a reason to retreat but a signal to refine their strategies. At the forefront of this transition, leaders from these three companies argue that the era of the human engineer as a primary writer of code is rapidly ending, replaced by a model where humans orchestrate, review, and manage fleets of AI agents. The central question is no longer whether <a href=\"https:\/\/overcentral.com\/en\/supabase-evals-benchmark-ai-coding-agents\/\" title=\"Supabase Releases Open Source Benchmark for AI Coding Agents\" data-iacss-internal=\"1\">AI coding agents<\/a> work, but how to make them work without breaking the bank.<\/p>\n<p>At Kilo Code, the shift is stark. According to co-founder Emilie Schario, engineers within the company are now actively reading or writing code themselves only about one percent of the time. The remaining 99 percent of the work is executed by autonomous agents. This fundamental change is forcing development teams to confront a new set of operational questions: Which parts of the system are safe to hand over to an AI? Who is responsible for cleaning up when a model makes a mistake? How do you support an architecture that uses models from multiple providers? And, perhaps most pressingly, do skyrocketing token bills represent genuine productivity gains or just a burned IT budget?<\/p>\n<h2>When 99 Percent of Coding Is Handed to Agents<\/h2>\n<p>The transformation described by Schario at VB Transform 2026 represents a generational shift in software engineering. At Kilo Code, the human role has evolved from craftsman to conductor. \u201cUnless something&#8217;s really broken or debugging, 99% of the time engineers are not reading or writing code anymore,\u201d Schario said. This is not a pilot program or a limited experiment; it is the operational reality for a company building agentic coding tools for the enterprise.<\/p>\n<p>This level of delegation demands a high degree of trust, which Schario argues must be built on a foundation of clear guidance, shared skills, and robust protocols. The ability to give an agent a task and trust its output is not automatic. It requires a structured system. Model Context Protocol (MCP), a standard for how AI models access and use context, is a critical piece of this puzzle. By empowering models with the right context, companies can significantly improve the quality of agent output and reduce the need for human intervention.<\/p>\n<p>For Jared Go, distinguished engineer for AI and cloud at Symbotic, a warehouse automation company, the current moment is about directing the focus of AI with precision. Rather than letting an agent run wild, Go provides strict criteria. \u201cThese are my criteria,\u201d he explained. \u201cLet&#8217;s look at it from the lens of security, elegance, clean, concise code, water tightness.\u201d When the constraints are clear, AI can perform the heavy lifting, reducing the criticality of human code review.<\/p>\n<p>However, Go acknowledges that human involvement becomes essential further down the line, particularly when product decisions are at stake. This is where the distinction between greenfield and brownfield development becomes crucial.<\/p>\n<h3>Greenfield Is Easy, Brownfield Is the Real Fight<\/h3>\n<p>The performance gap between creating new code and modifying existing code is one of the most significant technical barriers to mass AI agent adoption. \u201cGreenfield is so easy for agents,\u201d Go noted. Building a brand new codebase from scratch, with a clean slate and no legacy constraints, is a task where AI agents excel. The challenge, as any experienced engineer knows, lies in the other direction. \u201cBrownfield we all know is where the actual challenge lies,\u201d Go said.<\/p>\n<p>Writing, updating, or maintaining existing code requires a deep understanding of context, historical decisions, and often, messy dependencies. Agents frequently struggle with this. They may introduce breaking changes, fail to understand the intent behind a legacy workaround, or simply generate code that does not fit the existing architecture. This is why Symbotic, despite heavy AI use, maintains a human loop for oversight.<\/p>\n<h2>Replit\u2019s Approach: Human on the Loop, Not in the Loop<\/h2>\n<p>While Kilo Code and Symbotic have embraced agentic workflows, Replit has taken a more measured, but no less ambitious, path. Amol Jain, head of product engineering at Replit, describes a system that has \u201cgone very agentic\u201d but remains \u201cconservative\u201d in its approach to <a href=\"https:\/\/overcentral.com\/en\/ai-code-writing-reshapes-software-supply-chain-security\/\" title=\"AI Code Writing Reshapes Software Supply Chain Security\" data-iacss-internal=\"1\">AI code<\/a> review.<\/p>\n<p>Replit\u2019s internal tool is designed for a self-driving software engineering experience. Developers give a task to a fleet of agents, which then handle end-to-end planning, implementation, and testing. \u201cIt&#8217;s a fleet of agents that run in their own cloud virtual machines (VMs) with access controls behind token proxies so they&#8217;re secure,\u201d Jain said.<\/p>\n<p>The key innovation is in the review process. An agent reviews each pull request (PR) and assigns it a risk score. Low-risk PRs are considered safe for self-merging by their author. Higher-risk PRs are flagged for human reviewers, who read the code and provide feedback. \u201cThe idea was human on the loop, not human in the loop,\u201d Jain explained. This distinction is critical. It means humans are not bottlenecks, but overseers who only step in when the system deems it necessary.<\/p>\n<p>Jain shared a powerful example of the system in action. An engineer was struggling to reproduce and solve a \u201cvery gnarly bug\u201d deep in Replit\u2019s systems. The bug was sent to an AI manager agent, which made an unusual decision: it told the bug to go to sleep. The manager agent then spun up a group of underlying agents to investigate. These agents found the root cause. The manager then spun up a second group of agents that found a fix. Six hours later, the AI had a pull request ready for a bug that had puzzled human engineers for a considerable time.<\/p>\n<h2>The Rise of Multi-Model Architectures<\/h2>\n<p>One of the most significant shifts in the <a href=\"https:\/\/overcentral.com\/en\/ai-coding-agents-trigger-security-rules\/\" title=\"AI Coding Agents Trigger Endpoint Security Rules Meant for Attackers\" data-iacss-internal=\"1\">AI coding<\/a> landscape is the move away from single-provider lock-in. As customers become more sophisticated, they are demanding the ability to choose the best model for each specific task. This trend is driving a new architectural paradigm: the multi-model gateway.<\/p>\n<p>Kilo Code now supports over 500 models in its gateway, allowing its customers to pick and choose with unprecedented granularity. \u201cYour software that you&#8217;re using to do agentic engineering should be decoupled from the model that you&#8217;re using to do it,\u201d Schario said. This decoupling is the foundation of cost control and performance optimization.<\/p>\n<p>A common workflow, according to Schario, involves using an expensive, frontier-tier model for the architecture and planning phase of a project. This is where reasoning and complex logic are most critical. Once the plan is set, the company switches to a much less expensive open-weight model to execute the bulk of the coding work.<\/p>\n<p>The routing logic itself is becoming a sophisticated piece of engineering. It must factor in data retention policies, regional provider limitations, the specific keys a customer has brought, and the type of commits being made. \u201cIt&#8217;s factoring in what&#8217;s important to you, what limitations you&#8217;ve set, what data retention policies you&#8217;ve established, what keys you&#8217;ve brought in, what commits you might have \u2026 into that routing decision,\u201d Schario said.<\/p>\n<p>Replit takes a similar view, positioning itself as the decision-maker on behalf of the user. \u201cWe are essentially making the decisions on users&#8217; behalf of what model to use when, in what capacity, to minimize cost and maximize capability,\u201d Jain said. This approach requires deep insight into the cost-versus-capability spectrum of hundreds of models, a task that most individual enterprises are not equipped to handle.<\/p>\n<h2>What Is Tokenmaxxing and How Does It Burn Budgets?<\/h2>\n<p>A new term has entered the lexicon of engineering leadership: tokenmaxxing. It refers to the practice or phenomenon of consuming an extremely high volume of tokens, often on premium models, leading to runaway costs. For many enterprises, tokenmaxxing has become a management crisis.<\/p>\n<p>Schario described hearing from customers who have accidentally spent their entire AI budget for the year in a short period. \u201cI accidentally spent my whole AI budget for the year \u2026 so what do I do now?\u201d is a question she has heard directly. The answer, for Kilo Code, is to guide them back to the multi-model workflow: use expensive models for planning, then switch to open-weight models for execution.<\/p>\n<p>Visibility into usage is critical. Schario noted a particular engineer at her own company who has a \u201cheavy foot\u201d and is constantly at the top of the usage board. \u201cI regularly have to nudge, &#8216;What are you doing there?&#8217;\u201d she said. It is easy to look at a $600 daily bill and react with alarm, but the context is critical. When you consider the amount of work completed for that cost, the return can still be justified.<\/p>\n<p>Symbotic took a structural approach to the problem. The company built an internal tool that gives managers visibility into pull requests and usage trends per employee. They then set per-month cost tiers. Managers can move users up or down a tier as they see fit. \u201cHaving a cap and seeing how many people went up in cap this month makes a big difference when you&#8217;re trying to corral these costs and make things efficient,\u201d Go said.<\/p>\n<p>A major shock to Symbotic\u2019s cost structure came when Cursor, a coding tool the company uses heavily, ended a legacy discount. The company had been grandfathered into a flat per-request rate, even for frontier models. When Cursor moved everyone to full pricing, it forced a company-wide \u201creckoning on efficiency,\u201d according to Go. Engineers began sharing tips on which models worked best for specific tasks, like C# code, and which were a waste of money.<\/p>\n<h3>The Right Metric: Cost Per Pull Request<\/h3>\n<p>Schario argues that the traditional metrics of developer productivity are insufficient in the age of AI agents. Instead, she is focusing on a new key performance indicator. \u201cCost per pull request is the metric that I&#8217;m paying attention to right now,\u201d Schario said. \u201cIt feels like the closest proximity for how I can measure value.\u201d This metric directly ties the expense of tokens to the output of finished, reviewed code. It transforms the cost conversation from a raw spend number to a value-per-unit calculation.<\/p>\n<p>\u201cThe spend with no return on that spend is the problem,\u201d Schario said. The challenge for enterprises is to ensure that every token burned is yielding a measurable return in terms of features built, bugs fixed, or code refactored.<\/p>\n<h2>When the Cost Problem Moves Out of Engineering<\/h2>\n<p>The challenge of runaway AI costs is not confined to engineering teams. As platforms like Replit broaden access to agents across the organization, the risk of unexpected spending spreads. Jain recalled a case where a user on the support side of the business \u201cblew through an insane amount of money.\u201d When the team investigated, they discovered the support agent had been running an automation on a premium model\u2014Jain referred to it as \u201cGPT 5.5 Pro Max\u201d\u2014which was far more powerful and expensive than the task required.<\/p>\n<p>This case highlights a growing governance problem. \u201cAt least till that point, the ROI was rather clear,\u201d Jain said. \u201cWe could see engineering productivity 3X, so no one had questioned it yet.\u201d When AI agents are used outside of engineering, the ROI is often less visible and the cost discipline is weaker.<\/p>\n<p>To prevent this, Jain emphasized the need for sensible defaults, model routing, and visibility tools that are \u201cnot anti-productive.\u201d The most critical rule is simple. \u201cMost tasks do not need the frontier,\u201d he said. Defaulting to the cheapest capable model, rather than the most powerful available model, is the single most effective way to control costs.<\/p>\n<p>The era of unlimited, unmonitored AI spending is over. Enterprises are learning that AI coding agents are not a magic bullet but a powerful new tool that requires careful financial and operational governance. The companies that succeed will be those that treat agentic AI as a managed portfolio of compute resources, optimizing for cost per pull request and routing work to the right model at the right time. The future of software engineering is not about humans writing less code, but about humans becoming far more strategic about how they spend their AI budget.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>The economics of artificial intelligence in software development are undergoing a brutal, real-world stress test. As engineering teams rush to integrate agentic coding tools into their daily workflows, many are discovering a painful paradox: the very technology designed to accelerate output is also capable of devouring entire annual IT budgets in a matter of weeks. [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":84029,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/75069.png","fifu_image_alt":"AI coding agents blow through budgets; Replit, Kilo Code, Symbotic manage","footnotes":""},"categories":[31],"tags":[],"class_list":["post-75069","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/75069.png","fifu_image_alt":"AI coding agents blow through budgets; Replit, Kilo Code, Symbotic manage","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/75069","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=75069"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/75069\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/84029"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=75069"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=75069"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=75069"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}