The core loop of an AI agent looks deceptively simple. An LLM calls tools, receives results, and repeats. For short tasks that resolve in a handful of steps, that loop holds together. Give it a migration job that runs for an hour and executes 200 tool calls, and the breakdown arrives in two predictable forms. The AWS Samples design guide for autonomous cloud coding agents names them directly: shallow agents suffer from context overflow, get distracted through goal loss, and fail to maintain state over extended periods. The layer that fixes this is not the model. It is the harness — the orchestration shell that AWS describes as managing everything but the LLM itself.
This article opens that layer for examination. Compaction, memory strategy, context budgeting, and todo-state recitation form the machinery that transforms a shallow loop into a deep, durable agent. We examine how LangChain Deep Agents, Claude Code, Manus, OpenAI Codex, and Amazon Bedrock AgentCore implement each mechanism, including the actual thresholds they ship and the trade-offs they accept.
Why a Bigger Context Window Does Not Solve Agent Context Overflow
The obvious engineering response to context overflow is to provision a larger window. The evidence accumulated across multiple independent evaluations says this helps far less than intuition suggests. Chroma’s Context Rot report tested 18 LLMs including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3, and found that performance grows increasingly unreliable as input length increases, even on simple retrieval tasks. Anthropic’s context engineering guide explains the underlying mechanism: attention creates n² pairwise relationships for n tokens, so every added token depletes a finite “attention budget.” Context is a resource with diminishing returns, not a bucket that can be enlarged indefinitely.
For an agent executing a multi-hour loop, this dynamic is worse than it appears from a single-pass benchmark. Manus reports that a typical complex task requires around 50 tool calls, and that the input-to-output token ratio runs near 100:1. Every observation lands in context and stays there by default. The original instruction — the task goal, the constraint, the user’s priority — drifts toward the middle of the window as new observations pile in. That middle region is precisely where recall degrades fastest across every major model family. Goal loss is not a model bug that a larger context window can patch. It is the expected outcome of an unmanaged context on any task long enough to fill a window beyond its reliable working capacity.
Mechanism 1: Context Budgeting and Offloading
The first job of a competent harness is deciding what never enters the window at all. LangChain’s Deep Agents ships two offloading rules with hard, production-tested numbers. When a tool response exceeds 20,000 tokens, it is written to the filesystem and replaced in context with a file path plus a preview of the first 10 lines. When session context crosses 85% of the model’s window, older write and edit tool calls — whose full file contents already reside on disk — are truncated to a pointer. Only after offloading exhausts every available compression does the harness fall back to summarization, which is lossy by design.
Claude Code applies the same budgeting discipline to what loads before the first prompt even fires. Auto memory is capped at the first 200 lines or 25 kilobytes. MCP tool schemas stay deferred by default, with only tool names listed in the initial context; full schemas load on demand through a tool search mechanism. After compaction, any file that needs to be re-read and exceeds 5,000 tokens arrives as a path reference rather than inline content. The context window simulation published in the Claude Code documentation makes the payoff concrete: a research subagent reads 6,100 tokens worth of files and returns a 420-token distilled result to its parent agent.
That subagent pattern is budgeting at the architecture level rather than the token level. Anthropic’s guide notes that each subagent may burn tens of thousands of tokens exploring a topic independently, but returns only a distilled summary — often 1,000 to 2,000 tokens — to the coordinator. The Amazon Bedrock AgentCore walkthrough builds exactly this topology: a coordinator spawns three browser subagents in parallel, each isolated in its own MicroVM, and an analyst subagent receives only their structured findings. AWS reports a 4-to-6-minute expected runtime for the full pipeline, and notes that sequential processing would take up to three times longer.
For readers evaluating their own agent architecture, the practical question is straightforward: what percentage of your context window is consumed by raw tool output that could be replaced with a pointer or a preview? Most teams running code-agent workflows find that number sits between 40 and 70 percent of total context, which means offloading alone can double the effective working capacity of a fixed window without any model change.
Mechanism 2: Compaction — Summarization with Intent Preservation
When offloading is no longer sufficient, the harness must summarize. Compaction is the practice of taking a conversation state that approaches the window limit, condensing it through an LLM call, and reinitiating a new context with the summary. It is also the point where goal loss most often occurs, because a lossy summary can drop the one constraint that mattered three turns ago.
The implementations differ sharply in what they promise to retain. Claude Code’s compaction prompt explicitly preserves architectural decisions, unresolved bugs, and implementation details while discarding redundant tool outputs and verbose logs. Immediately after compaction, the agent re-reads up to five of the files modified most recently, reloads the configuration rules matching those files, and re-injects any invoked skill bodies — capped at 5,000 tokens per skill and 25,000 total across all skills. The documentation is explicit that detailed instructions from early in the conversation may be lost, which is why persistent rules belong in the project-root CLAUDE.md file, which is re-injected from disk on every compaction cycle. Users can steer the process with commands like /compact focus on the auth bug fixcodecodecodecodecodecode or shift the trigger point with /autocompactcodecodecodecodecodecode.
Deep Agents made goal preservation a structural feature rather than a prompt engineering artifact. Its summary is a structured document with dedicated fields for session intent, artifacts created, and next steps. The LangChain team added those fields after forced-summarization experiments showed measurable improvement in task completion rates. The full original transcript is also written to the filesystem, so any fact that was summarized away can be recovered later through a read_file call — a safety valve that turns a lossy compression into a reversible operation.
Compaction has moved into the API layer as well. OpenAI’s Responses API offers server-side compaction through a context_managementcodecodecodecodecodecode parameter with a compact_thresholdcodecodecodecodecodecode, plus a standalone /responses/compactcodecodecodecodecodecode endpoint that returns a compacted context window containing an opaque, encrypted compaction item. OpenAI instructs developers to pass that returned window unchanged into the next API call. Codex relies on this mechanism to sustain long-running coding tasks that would otherwise exhaust the model’s context budget mid-task. The Claude Developer Platform exposes a compact_20260112codecodecodecodecodecode context-management edit with custom instructions and a pause_after_compactioncodecodecodecodecodecode option that allows inserting content before the model continues. When custom instructions are written for that edit, they replace the default prompt entirely — so a compaction prompt is a real engineering artifact requiring careful design, not a configuration toggle.
The critical insight for practitioners is that compaction is not a single operation. It is a pipeline: compress, verify what was lost, reload critical state from disk, and continue. Implementations that skip the verification and reload steps produce agents that appear to work on short tasks and silently fail on long ones.
Mechanism 3: Todo-State and Recitation
Compaction protects the goal at the moment of summarization. Todo-state protects it on every turn in between. Manus described the mechanism plainly: its agent creates a todo.md file and rewrites it step by step, checking items off as they complete. Rewriting the list recites the objectives into the end of the context, pushing the global plan into the model’s recent attention span and reducing the “lost in the middle” drift that plagues long conversations. No architecture change is required. It is natural language used to bias the model’s own attention allocation.
The evidence on todo-state is not one-sided, and the nuanced results are worth examining closely. Deep Agents shipped a write_todos tool by default until version 0.7 in July 2026, when LangChain made TodoListMiddleware opt-in after running evaluations across three task categories. Those evals showed slightly better reward and lower cost with todos disabled for many common workloads. LangChain still recommends turning the feature back on for long multi-step tasks, for less capable models that need structural reinforcement, and for user interfaces that display progress to end users. Claude Code keeps a todo list and re-injects the plan written in plan mode from disk after every compaction. Anthropic’s guide calls the general pattern “structured note-taking”: the agent writes a NOTES.md or TODO file outside the window and reloads it periodically. Its Claude Plays Pokémon example maintained tallies across thousands of game steps, read its own notes after each context reset, and resumed multi-hour sequences without losing the objective.
The pattern behind all of these implementations is that the goal exists as a mutable artifact — a file — rather than only as a message in conversation history. Messages age, get pushed toward the middle of the window, and eventually get summarized away. A file that is rewritten every few turns is always recent, always short, and survives any context reset. Whether that recitation cost — the tokens consumed by rewriting the todo list every cycle — is worth its per-turn tax depends on the model and the task length, which is exactly what the Deep Agents evals measured. For short tasks, the todo overhead reduces efficiency. For tasks that run past 30 tool calls, the insurance against goal drift becomes the dominant factor.
Mechanism 4: Memory Strategy Across Sessions
The last piece of the harness is what persists after the task ends and a new one begins. Claude Code re-injects the project-root CLAUDE.md and auto memory from disk after every compaction cycle. Amazon Bedrock AgentCore Memory stores events and runs configured extraction strategies in the background, so a coordinator agent can call a recall tool on the next run instead of re-researching the same information. AWS warns that without at least one extraction strategy configured, raw events are stored but nothing is extracted for retrieval — the memory fills up but remains effectively inaccessible. Anthropic’s file-based memory tool serves the same purpose on the Claude platform, allowing agents to write persistent notes that survive across separate invocations.
The limitation that every team must confront is that persistent context is not free. The ETH Zurich study covered in February 2026 found that repository context files like AGENTS.md do not generally improve task success while reliably raising inference cost. LLM-generated context files increased cost by 20 percent and 23 percent on the two benchmarks tested, and developer-committed files raised cost by up to 19 percent. Memory that reloads on every session is a standing tax on both the attention budget and the monetary budget. The Claude Code documentation gives matching advice: keep CLAUDE.md under 200 lines and move reference material into skills or path-scoped rules that load only when needed, not on every turn.
What is the most important thing for a practitioner to understand about memory strategy? The answer fits into a featured snippet: Persistent memory in agent systems should be treated as a cache with an eviction policy, not as an infinite store. The ETH Zurich study measured a 20 to 23 percent inference cost increase from LLM-generated context files with no corresponding task success gain. Effective memory strategy requires scoping what loads to what is needed for the current task, not what might be needed for any future task.
Testing Whether the Harness Actually Holds the Goal
Context management is only useful if the agent can still finish the task and recover details it no longer sees in its window. LangChain maintains targeted evaluations for exactly this purpose: tests that trigger summarization mid-task and then measure whether the agent continues toward its objective, and needle-in-a-haystack cases where a critical fact is summarized away and must be recovered through filesystem search. To generate enough events to compare prompt variants efficiently, the team triggers summarization at 10 to 20 percent of the window instead of the 85 percent default, and used a 25 percent trigger with Claude Sonnet 4.5 on terminal-bench-2 to study the effect of different compaction prompt designs.
The failure mode to watch for, in LangChain’s view, is goal drift: an agent that asks for clarification immediately after a summary, or wrongly declares the task complete when it has only completed a sub-step. Amazon Bedrock AgentCore Evaluations ships a goal success rate evaluator that can score the same traces against ground-truth task definitions. The practical implication is direct: if you run a harness and have not forced a compaction in a controlled test, you do not yet know what your summary prompt drops. Compaction is the only mechanism that permanently removes information from the active window, and it deserves the same testing rigor that teams apply to model selection or tool design.
The discipline of testing compaction behavior is still emerging across the industry. Most teams test on short tasks that never trigger compaction, then deploy to production tasks that do, and discover goal drift only through user complaints. The open-source tooling from LangChain and the evaluation infrastructure from Bedrock AgentCore represent the first generation of solutions to this testing gap. Expect this space to mature rapidly as more production deployments hit the wall of long-running agent tasks.
The four mechanisms — budgeting and offloading, compaction with intent preservation, todo-state recitation, and scoped memory strategy — form a coherent system for managing agent context overflow. No single mechanism is sufficient. Offloading defers the problem but does not eliminate it. Compaction solves the window limit but introduces goal loss risk. Todo recitation protects the goal but adds per-turn token cost. Memory strategy extends the agent beyond a single session but taxes every call. The harness that balances all four, with thresholds tuned to the specific model and task length, is what separates a shallow agent loop from a deep agent that can run for hours, execute hundreds of tool calls, and still finish the job it started.