{"id":81006,"date":"2026-09-11T20:35:15","date_gmt":"2026-09-12T00:35:15","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=81006"},"modified":"2026-09-11T20:35:15","modified_gmt":"2026-09-12T00:35:15","slug":"harnessdev-llm-harness-generalization-study-81006","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/harnessdev-llm-harness-generalization-study-81006\/","title":{"rendered":"HarnessDev Reveals Only 34 of 64 LLM Harness Changes Generalize"},"content":{"rendered":"<p>Agent harnesses are the invisible scaffolding that determines whether a large language model succeeds or fails in the real world, yet until now the field had no systematic way to evaluate them. A harness is the code that wraps the model: the execution loop, tool definitions, context management, state handling, recovery mechanisms, and verification logic. The same model with the same weights can produce wildly different results depending on which harness it sits inside. On the Terminal-Bench 2.1 leaderboard, for example, GPT-5 solves 35.2 percent of tasks inside the Terminus 2 harness but 49.6 percent inside Codex CLI \u2014 an absolute gap of 14.4 points with no change to the model itself. Most benchmarks fix the harness and treat it as irrelevant infrastructure. Now a team of researchers from <a href=\"https:\/\/www.bytedance.com\" target=\"_blank\" rel=\"sponsored noopener noreferrer\" data-iacss-external=\"1\">ByteDance Seed<\/a>, Singapore University of Technology and Design, Georgia Institute of Technology, M&#8209;A&#8209;P, and <a href=\"https:\/\/tokenwave.ai\" target=\"_blank\" rel=\"sponsored noopener noreferrer\" data-iacss-external=\"1\">TokenWave.AI<\/a> has flipped that assumption entirely. Their new framework, <a href=\"https:\/\/harnessdev.com\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">HarnessDev<\/a>, makes the harness the object of evaluation: the artifact under scrutiny is the runnable code the model writes, not the answer it produces. In a finding that will reshape how the industry thinks about LLM agent reliability, HarnessDev reveals that only 34 of 64 changes made during harness evolution actually generalize, exposing a fragile feedback loop between a model&#8217;s self-built infrastructure and its performance on unseen tasks.<\/p>\n<h2>What Is HarnessDev? Evaluating the Scaffold, Not the Answer<\/h2>\n<p>HarnessDev is a two-stage benchmark designed to test whether a large language model can create and improve its own agent harness \u2014 the code that governs how the model interacts with tools, handles errors, manages context, and verifies results. The core insight is that the harness is a separable artifact: its quality can be measured independently of the model that runs inside it. The framework provides a shared weak seed harness \u2014 a bare set of passive primitives for file I\/O, search, process execution, and logging, with no loop, planner, verifier, retry logic, or stopping rule \u2014 and asks each participating model to build a full harness from that seed. The resulting harness can then be executed by any LLM, not only the one that wrote it, making it possible to isolate the quality of the code from the quality of the model that generated it.<\/p>\n<h2>Two Stages: Creation and Evolution<\/h2>\n<h3>Creation: From a Seed That Scores Zero<\/h3>\n<p>In the Creation stage, every participating model receives the same weak seed harness. Unmodified, that harness scores zero on every benchmark because it contains no policy for deciding when to act, what tool to call, how to manage context, or when to stop. The creator also receives a task-family specification, a short design tutorial, and one to three development cases drawn from the family. With these materials, the model must build a complete harness, which is then frozen before it is allowed to see any hidden test instances. The harness is graded on capability (task success rate) and efficiency (token consumption by the execution model; creator tokens are not counted).<\/p>\n<h3>Evolution: Revising With Execution Feedback<\/h3>\n<p>In the Evolution stage, each creator starts from its own frozen Creation harness and revises it using execution feedback from a fixed set of 100 SWE-bench Pro tasks and all 89 Terminal-Bench 2.1 tasks. Each official candidate must complete both evaluations as a pair, with a budget of ten pairs and at most two five-task probes between pairs. Every official version is later scored on 630 held-out SWE-Pro instances that the creator never sees. Progress is measured as the change in held-out score \u2014 the true test of generalization. The Evolution stage is designed to simulate how a development team would iteratively improve an agent harness: run tests, inspect failures, fix bugs, and deploy. But as the results show, this loop is far from reliable.<\/p>\n<h2>The Setup: Six Creator LLMs and Five Benchmarks<\/h2>\n<p>The researchers tested six creator LLMs: Opus 4.8, GPT&#8209;5.5, Gemini 3.1 Pro, DeepSeek V4 Pro, Qwen 3.7 Max, and Seed 2.0 Pro. Each model worked inside a fixed development environment (Claude Code 2.1.177, except GPT&#8209;5.5 which used Codex 0.144.3). Creation spanned four domains across five benchmarks totaling 2,207 instances: SWE-bench Pro public split (731 tasks), Terminal-Bench 2.1 (89), MLE-bench (75), EQ-Bench3 (46), and BrowseComp (1,266). Each creator built three harnesses per benchmark, reported as an average of three (avg@3). Harnesses were evaluated under two conditions: Self-Eval, where each harness was run by its own creator model, and Unified-Eval, where all harnesses were run by a single fixed executor (Gemini 3.1 Pro) to isolate harness quality from model capability.<\/p>\n<h2>Creation Results: Gaps, Surprises, and Dead Code<\/h2>\n<p>Under Self-Eval, Opus 4.8 posted the highest average score at 67.8, against a human-engineered reference of 86.2 \u2014 a gap of 18.4 points. The gap varied dramatically by domain. On code tasks, Opus 4.8 reached 69.3 on SWE-Pro versus the 80.0 reference, while Gemini 3.1 Pro led Terminal-Bench at 68.8 against 88.8. The widest gap appeared in search: the best BrowseComp score was 52.6 (GPT&#8209;5.5) against a reference of 92.2. Writing was the only domain where a model exceeded the reference: Opus 4.8 scored 84.6 on EQ-Bench3, above the 83.7 reference. In ML experimentation, both Opus 4.8 (32.9) and Gemini 3.1 Pro (32.4) beat the 24.0 MLE-bench reference, though the medal rate metric for MLE-bench is not directly comparable to task success rates used elsewhere.<\/p>\n<p>Code volume did not predict quality. The 18 code harnesses added a total of 17,111 net lines, yet Gemini added the fewest (1,006) and led Terminal-Bench. Self-test count barely correlated with score (Spearman 0.13 to 0.26); only revision calls reached 0.57. Much of the generated machinery was inert. Of 108 code component instances, 72 triggered in real runs and 18 never fired \u2014 all of them state and memory components. Eleven of the 18 harnesses defined a State class, yet no checkpoint event appeared across 26,679 trajectories. Among writing harnesses, 124 of 587 features were dead code. This suggests that models are generating structural code they believe is necessary, but that code rarely executes in practice.<\/p>\n<h2>Cost and Executor Transfer: The Hidden Fragility<\/h2>\n<p>Token costs varied enormously. On MLE-bench, GPT&#8209;5.5 hit a 19.1 medal rate with 29.3 million tokens, while DeepSeek V4 Pro achieved a similar 19.6 rate using 208.4 million tokens \u2014 a roughly 7.1x difference for comparable performance. Swapping the executor from the creator model to a fixed Gemini 3.1 Pro reshuffled rankings completely. Qwen 3.7 Max gained 17.6 points on BrowseComp and 12.9 on MLE-bench, suggesting its own executor was a bottleneck. But Opus 4.8, which dominated under Self-Eval, saw its SWE-Pro score collapse from 69.3 to 33.0 under Gemini, partly because one of its harnesses hard-coded a 120-step limit that was tuned for its own executor. The Opus search harness\u2019s duplicate-query rate jumped from 10.1 percent to 88.2 percent after the switch, revealing that the harness had implicitly relied on the creator model\u2019s ability to filter duplicates \u2014 a capability Gemini did not replicate.<\/p>\n<h2>Evolution Results: Only 34 of 64 Changes Generalize<\/h2>\n<p>The Evolution stage produced nine lineages (five using self-runtime feedback, four using fixed-Gemini feedback) with 73 official versions and 64 adjacent switches \u2014 changes from one official version to the next. All five self-runtime creators improved on held-out tasks, from +1.43 to +4.44 points, with a mean of +3.11. Under fixed Gemini, only Opus improved; GPT&#8209;5.5 regressed by 10.32 points. Progress was anything but monotonic. Of the 64 switches, 8 regressed on both benchmarks, 16 regressed on one, 27 gained only within the noise band (about \u00b14.75 pair-score points), and only 2 showed clear positive evidence beyond noise. Feedback and held-out scores moved in the same direction only 34 of 64 times (53.1 percent) \u2014 barely better than a coin flip. Only 2 of the 9 declared final versions were actually the held-out optimal version. A single commit could vary by about \u00b14.75 points, meaning much of the observed \u201cimprovement\u201d could be noise. Of 169 new functions or classes added during evolution, 25 had no caller.<\/p>\n<p>The clearest win came from Opus 4.8, which noticed that 99 of 100 runs reported success while only 48 actually passed, traced the discrepancy to premature completion, and added a completion gate. Failure diagnosis was otherwise the weakest step: the dedicated trajectory inspection interface was called only twice across all lineages. This suggests that LLMs, when given execution logs, are not yet good at diagnosing why a harness failed.<\/p>\n<h2>What Does the 34-of-64 Result Mean for the Industry?<\/h2>\n<p>The finding that only 34 of 64 evolutionary changes moved feedback and held-out scores in the same direction has profound implications for any organization that relies on benchmark-driven iteration of LLM agents. The standard practice \u2014 run a test, observe a failure, fix the code, deploy \u2014 assumes that the feedback signal from a fixed set of development tasks is a reliable proxy for performance on unseen tasks. HarnessDev shows that this assumption fails <a href=\"https:\/\/overcentral.com\/en\/google-hollywood-ai-licensing-79386\/\" title=\"Google Needs Hollywood More Than Studios Need AI\" data-iacss-internal=\"1\">more than<\/a> 40 percent of the time. A change that improves scores on the development set can just as easily degrade performance on held-out tasks, and vice versa. The noise band of roughly \u00b14.75 points means that small improvements, even if statistically significant on the development set, are likely spurious.<\/p>\n<p>This fragility is compounded by the fact that harnesses are deeply coupled to their executor. A harness that performs well when paired with its creator model can break completely under a different executor, as the Opus 4.8 example demonstrates. The hard-coded 120-step limit and the implicit reliance on the creator\u2019s duplicate-filtering behavior are not bugs in the conventional sense \u2014 they are adaptive optimizations that happen to be non-portable. In a world where agent deployments often switch underlying models (e.g., from GPT-4 to Claude or Gemini), these hidden dependencies create brittle systems that fail silently.<\/p>\n<p>The dead-code problem also raises questions about over-engineering. The fact that 11 out of 18 harnesses define a State class but never checkpoint, or that 124 of 587 writing features are never invoked, suggests that LLMs are producing structural complexity that <a href=\"https:\/\/overcentral.com\/en\/ai-search-moves-cognitive-load-does-not-remove-it\/\" title=\"AI Search Moves Cognitive Load, Does Not Remove It\" data-iacss-internal=\"1\">does not<\/a> contribute to task success. More code does not mean better code. The Gemini harness, which added the fewest lines, led Terminal-Bench. Simplicity, when aligned with the task domain, appears to be a stronger predictor of harness quality than verbosity.<\/p>\n<h2>How HarnessDev Works: An Illustrated Breakdown<\/h2>\n<p>The paper includes an interactive explainer that walks through the seed-to-harness process, the Creation scores across domains, the effect of swapping the executor, and the Evolution outcomes. The explainer visualizes how a weak seed \u2014 containing only passive primitives for CLI, file operations, and execution, with no policy \u2014 is built up by a creator model that adds six control modules: the execution loop, tool policy, context management, state and memory, lifecycle handling, and verification. In the paper\u2019s sample runs, state and memory is the weakest module, mirroring the overall finding that models declare state classes but rarely use them. The explainer also shows the creation scores as bar charts against human references, the effect of switching from Self-Eval to Unified-Eval (green for gains, amber for losses), and a segmented bar of the 64 evolution switches color-coded by outcome \u2014 from regress-on-both to clear-gain-beyond-noise.<\/p>\n<p><strong>Featured snippet:<\/strong> HarnessDev evaluates whether an LLM can create and improve its own agent harness \u2014 the code that wraps the model for execution, tools, state, and verification. In the Creation stage, each model starts from a minimal seed that scores zero and must build a complete harness using task specifications and development cases. In the Evolution stage, the model revises its harness using execution feedback from a fixed set of tasks, with each version later scored on held-out instances to measure generalization. The framework reveals that only 34 of 64 evolutionary changes generalize, and that harness quality is highly sensitive to the executor model, with hard-coded and implicit dependencies causing large performance swings when the executor changes.<\/p>\n<h2>Implications for Agent Infrastructure and Future Research<\/h2>\n<p>The HarnessDev results arrive at a moment when the industry is rapidly shipping agentic systems \u2014 coding assistants, browser automation tools, data science agents \u2014 all of which depend on harnesses that are hand-crafted by engineers or, increasingly, generated by LLMs themselves. The finding that only 53.1 percent of evolutionary steps move in the intended direction should give pause to anyone using automated code generation to build production agent infrastructure. It suggests that the feedback loop between test performance and real-world generalization is noisier and more fragile than commonly assumed. The team\u2019s recommendation to use a fixed executor during evolution to isolate harness improvements from model improvements is a practical takeaway, but the fact that even under a fixed executor only Opus improved while GPT&#8209;5.5 regressed by double digits indicates that the problem runs deeper.<\/p>\n<p>The dead-code finding also points to a need for better code pruning and verification during harness generation. If models are adding complex state management that never executes, they are wasting tokens and cognitive overhead without gaining capability. The strong performance of the minimalist Gemini harness suggests that the field may need to incentivize parsimony in harness design, perhaps through token cost metrics that penalize dead code.<\/p>\n<p>Future research could extend HarnessDev to multi-model evaluation, where harnesses are stress-tested across a range of executors with varying capabilities. It could also explore whether models can be trained to diagnose failures more effectively \u2014 the trajectory inspection interface was used only twice, indicating that even when feedback is available, models do not exploit it. The interactive explainer embedded in the paper provides a user-friendly window into these dynamics, but the raw data is clear: building a reliable harness is harder than building a reliable answer.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Agent harnesses are the invisible scaffolding that determines whether a large language model succeeds or fails in the real world, yet until now the field had no systematic way to evaluate them. A harness is the code that wraps the model: the execution loop, tool definitions, context management, state handling, recovery mechanisms, and verification logic. [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":83241,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/81006.png","fifu_image_alt":"HarnessDev Reveals Only 34 of 64 LLM Harness Changes Generalize","footnotes":""},"categories":[31],"tags":[],"class_list":["post-81006","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/81006.png","fifu_image_alt":"HarnessDev Reveals Only 34 of 64 LLM Harness Changes Generalize","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/81006","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=81006"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/81006\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/83241"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=81006"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=81006"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=81006"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}