Context Engineering for AI Agents
Brilliant at step one. By step forty, your agent forgot the task.
Why agents get worse the longer they run, and how to fix it
Chapters
- Intro
- Why context is the bottleneck
- How an agent fills its window
- Villain 1: context rot
- Villain 2: context poisoning
- The four levers
- Keep the cache warm
Who it's for
Engineers building AI agents (coding agents, research agents, support agents) that run for dozens of tool calls, and who see the agent start sharp and then drift, repeat itself, forget the goal or get expensive as a task goes on.
Context
An agent is a loop: the model reads its context, picks an action (usually a tool call), the tool's result is appended to the context, and the loop runs again. The model sees nothing but that context: the system prompt, the tool definitions, the conversation and every earlier tool result. "Context engineering" (a term popularised in mid-2025 by Andrej Karpathy, Shopify's Tobi Lütke and Anthropic) is the discipline of deciding what goes into that window at each step, the natural successor to prompt engineering once tasks span many turns.
Architecture
| Component | Role | Notes |
|---|---|---|
| System prompt + tool definitions | Stable instructions and the action space | Kept at the front, unchanged, so it stays in the KV / prompt cache |
| History | Earlier messages, actions and tool results | Grows every step; the main source of bloat |
| Retrieved context | Docs, code, records pulled in for this step | Should arrive just in time, not all up front |
| Context window | Everything the model reads on one call | A hard limit, and a soft one well before it (context rot) |
| LLM | Picks the next action | Quality depends on what the window holds |
| Notes / scratchpad (fix) | Plans, to-dos and facts written outside the window | e.g. NOTES.md or todo.md, re-read when needed |
| Compactor (fix) | Summarises old turns, clears stale tool output | Runs before the window fills |
| Sub-agents (fix) | Explore in their own clean windows | Return a short summary to the lead agent |
Request flow
- The user gives the agent a task.
- The agent builds the context: system prompt, tools, history, anything retrieved.
- The LLM reads it and returns an action, e.g.
search("..."). - The tool runs; its result (often thousands of tokens) is appended to the history.
- Steps 2–4 repeat. Every step the window is larger than the one before, until the task ends or the window is full.
The failure modes
Context rot. Models don't use a long context uniformly: reliability falls as the input grows, well before the advertised limit. Chroma's July 2025 study tested 18 models (including GPT-4.1, Claude 4, Gemini 2.5 and Qwen3) and all of them got less reliable as input length increased, even on simple retrieval and copying tasks; distractors and semantically similar text made it worse. Earlier, "Lost in the Middle" (Liu et al., 2023) showed accuracy is highest when the key fact is at the start or end of the context and drops when it is in the middle. In an agent, the task goal and the early decisions sit at the start while forty tool results pile up after them: the model drifts, repeats work or answers a question nobody asked.
Context poisoning. One wrong fact, a hallucination or a bad tool result, lands in the history. Because the history is re-read on every later step, the agent keeps building on it, and each new step reinforces it (Drew Breunig, "How Long Contexts Fail", June 2025, names this with its siblings: distraction, confusion and clash).
The quiet cost failure. Most providers cache the attention KV state of a prompt prefix. A single changed token early in the prompt (a timestamp in the system prompt, an edited earlier step, a tool list that changes) invalidates the cache from that point on, so every step pays full price. Manus reports that with Claude Sonnet cached input costs $0.30 per million tokens against $3.00 uncached, a 10x difference, and calls KV-cache hit rate the single most important metric for a production agent.
Why it happens
The window is treated as a transcript, a log that only grows, rather than as a working set the system curates. Everything the agent has ever seen competes for the same finite attention on every step.
The fix: four levers (write, select, compress, isolate)
- Write — keep plans, to-do lists and key facts in files outside the window (Anthropic calls it structured note-taking; Manus keeps a
todo.mdand treats the file system as memory). The agent re-reads the notes instead of scrolling through a transcript, and restating the goal at the end of the context keeps it in the model's recent attention. - Select — retrieve just in time: keep lightweight references (paths, ids, queries) and load the content only for the step that needs it, instead of stuffing everything up front.
- Compress — compaction: when the window nears its budget, summarise the conversation so far and restart from the summary; clear stale tool results that are no longer needed (Anthropic: tool-result clearing). Tune the compaction prompt for recall first, then precision.
- Isolate — sub-agents: hand broad searches to sub-agents that work in their own clean windows and return a short summary (often one or two thousand tokens) to the lead agent.
- And keep the cache warm — a stable prefix (no timestamps up front), an append-only history, tools masked rather than removed mid-task, deterministic serialisation. Manus also keeps failed actions in the context so the model can learn not to repeat them.
Trade-offs
- Compaction is lossy: a summary can drop the one detail that matters later. Measure recall on real traces.
- Notes and sub-agents add calls, latency and orchestration code; a sub-agent's summary can hide its mistakes.
- Just-in-time retrieval costs a tool call per lookup and depends on good references.
- Clearing or rewriting history breaks the prompt cache; do it rarely and in one go (compaction), not every step.
Numbers worth knowing
- 18 models tested in Chroma's "Context Rot" report (July 2025); all became less reliable as input length grew.
- $0.30 vs $3.00 per million input tokens, cached vs uncached, with Claude Sonnet as quoted by Manus (2025): 10x.
- Manus rebuilt its agent framework four times while learning how to shape context.
- The context-size chart in the video (≈3K tokens added per step) is illustrative.
One step, on a budget
def next_step(state):
ctx = [SYSTEM_PROMPT, TOOLS] # stable prefix
ctx += state.notes.read() # write
ctx += retrieve(state.goal, k=3) # select
ctx += state.history # append-only
if tokens(ctx) > 0.8 * BUDGET:
state.history = compact(state.history) # compress
ctx = rebuild(state)
if is_broad(state.goal):
summary = run_subagent(state.goal) # isolate
state.history.append(summary)
return llm(ctx)
Write, select, compress, isolate
Write
Plans, to-dos and key facts live in files outside the window
Select
Load documents just in time, only for the step that needs them
Compress
Summarise old turns and clear stale tool results
Isolate
Sub-agents explore in clean windows, return a summary
Recap: before and after
Cache-busting
- Timestamp at the top of the system prompt
- Earlier steps edited or deleted
- Tools added and removed mid-task
- Unstable JSON key order
Cache-friendly
- Stable prefix, dynamic data at the end
- Append-only history
- Tools masked, never removed
- Deterministic serialisation
Cached input: $0.30 vs $3.00 per million tokens (Claude Sonnet, via Manus)
Sources
Anthropic Engineering, "Effective context engineering for AI agents" (2025): context rot, compaction, structured note-taking, sub-agent architectures, just-in-time retrieval, tool-result clearing.
Chroma, "Context Rot: How Increasing Input Tokens Impacts LLM Performance" (July 2025), 18 models.
Liu et al., "Lost in the Middle: How Language Models Use Long Contexts" (2023, TACL 2024).
Drew Breunig, "How Long Contexts Fail" and "How to Fix Your Context" (June 2025): poisoning, distraction, confusion, clash.
Yichao "Peak" Ji, "Context Engineering for AI Agents: Lessons from Building Manus" (2025): KV-cache hit rate, stable prefix, append-only context, masking tools, file system as context, todo.md recitation, keeping errors.
LangChain, "Context Engineering for Agents" (2025): write, select, compress, isolate.
Coming next in the series: Prompt Injection in AI Agents