
Context Engineering - How to Manage Model Context So Your AI Agent Doesn't Lose the Thread
In June 2025 Andrej Karpathy suggested talking about "context engineering" instead of "prompt engineering". His definition caught on quickly: it's the delicate art and science of filling the context window with just the right information for the next step. The reason for the rename is practical. A prompt suggests a short instruction typed into a chat. In an AI agent that calls tools, reads files and collects results over dozens of steps, the instruction itself is a small part of what the model sees. The rest is system instructions, tool definitions, search results, conversation history and notes - and it's that rest that decides the quality of the answer.
A Chroma study showed that all 18 models tested lose quality as context length grows, and in the Manus agent every output token comes with about 100 input tokens. What context engineering is, why long context hurts, the four techniques (write, select, compress, isolate) and a context budget in practice.
The problem is that more context isn't better. A July 2025 Chroma study showed that all 18 models tested, including GPT-4.1, Claude Opus 4 and Gemini 2.5, perform worse as input length grows, even on simple tasks. An agent that "forgets" earlier decisions or starts making mistakes after an hour of work usually doesn't have too small a context window - it has a badly managed one. In this post I explain what context engineering is, why long context hurts, what the four core techniques are and how to apply them when building agents in a company.
/// WHY MORE CONTEXT ISN’T ALWAYS BETTER
What context engineering is
The context window is everything the model "sees" at a given moment: instructions, data, history and tool results. The model has no other working memory - if something isn't in the window, for the model it doesn't exist. Context engineering is designing what goes into the window, in what order and in what form, at every step of the agent's work.
The difference is clearest with an example. Prompt engineering answers the question: how do I phrase the instruction so the model understands it well. Context engineering answers the question: which of the thousands of available pieces of information should the model see now, and which are better kept outside the window and fetched only when needed. Writing good instructions still matters - I covered it in the post on prompt engineering for business - but with agents it's just one layer of several.
Why long context hurts
Model providers keep advertising larger context windows: 200,000 tokens, a million. It's tempting to conclude you can just put everything in. Research shows that's a mistake.
- Context rot. In the Chroma study, answer quality fell as input length grew in all 18 models and at every length increment tested. The drop was uneven and depended, among other things, on how many similar but irrelevant passages the text contained.
- Lost in the middle. Stanford's "Lost in the Middle" study showed that models make best use of information at the beginning and end of the context, and worst use of information in the middle.
- Cost and time. Every token in the window costs money on every call. In the Manus agent, the ratio of input to output tokens averages around 100 to 1 - almost the entire cost is reading context, not writing the answer.
The conclusion is simple: the context window is a resource with limited attention, not a warehouse. Every unnecessary passage lowers the odds that the model notices the important one.
Four techniques: write, select, compress, isolate
The clearest breakdown of techniques comes from the LangChain team. Each answers a different question.
/// FOUR CONTEXT ENGINEERING TECHNIQUES
* Framework from the LangChain team (write, select, compress, isolate).
1. Write outside the window. The agent keeps notes in files (e.g. NOTES.md or a task list) instead of holding everything in the conversation history. After dozens of steps it reads the notes and knows exactly where it left off. Manus describes the same pattern: the file system as practically unlimited memory, and a task list rewritten at the end of the context so the goal doesn't get lost in the middle of a long session.
2. Select only what's needed. Instead of loading all documents up front, the agent fetches them "just in time": it searches, reads a passage, checks. Anthropic calls this a just-in-time approach and recommends keeping lightweight pointers in context (file paths, identifiers, queries) rather than full contents. It's the same logic behind advanced RAG - good retrieval is half of context engineering.
3. Compress what's already there. When the window fills up, the older part of the conversation is summarized, and the agent continues with a fresh window and the summary. Anthropic recommends that the summarization prompt first maximize completeness and only then trim unnecessary detail. The simplest and safest form of compression is clearing old tool results that have already been used.
4. Isolate across agents. You split a complex task into subtasks for separate agents, each working in a clean window and returning only a short summary. In Anthropic's research system this setup scored 90.2% better than a single agent on internal tests, but used about 15 times more tokens than a regular chat. Isolating context costs money, so it pays off where the value of the result justifies the expense. When a multi-agent system makes sense and when it doesn't I covered in the post on multi-agent AI.
A context budget in practice
A well-designed agent has an explicit budget: you know how much space each type of information takes and what happens when the budget runs out. An example window layout for an agent handling customer queries:
[1] System instructions and rules fixed, at the start - they use the cache[2] Tool definitions only those needed for this task[3] Notes and memory short, updated by the agent[4] Search results and files passages, not whole documents[5] Conversation history recent steps in full, older ones summarized[6] Current task at the end of the window, right before the answer
The order isn't accidental. Fixed elements at the start let you use the cache, and the current task at the end lands where the model pays the most attention. Manus considers the cache hit rate the most important metric for an agent in production: with Claude Sonnet, cached tokens cost $0.30 per million versus $3 uncached, ten times more. A single changed character at the start of the instructions, such as a timestamp with seconds, is enough to invalidate the cache for everything after it. I cover lowering API costs, including caching, in more depth in the post on OpenAI API cost optimization.
Tools take up context too
It's easy to overlook that tool definitions go into the window on every call. A single MCP tool definition is typically a few hundred to over a thousand tokens, and GitHub's official MCP server with 94 tools takes about 17,600 tokens for definitions alone. A few connected servers can eat tens of thousands of tokens before the agent does anything.
There are two solutions. The first is tool search: the agent gets only a short index up front and loads a tool's full definition only when it wants to use it. The second is code execution instead of many calls: the agent writes a short script that calls the tools and processes the data itself, and only the result returns to the context. In an example described by Anthropic, this approach cut usage from 150,000 to about 2,000 tokens. I cover the protocol itself in the post on MCP, and how agents use tools in the post on AI agents.
In-session memory vs cross-session memory
Context engineering is about what happens within a single session of the agent's work. Memory across sessions - what the agent knows about a user after a week's break - is a separate problem I covered in the post on AI agent memory. The two layers connect: notes written during a session can become long-term memory, and long-term memory is one more source the agent selects from when deciding what to load into the window.
Model providers have started building these mechanisms into their platforms. Anthropic released automatic clearing of old tool results from the window and a file-based memory tool. In an internal search evaluation, combining the two improved results by 39% over the baseline, and in a 100-step search test, context clearing alone cut token usage by 84% and let the agent finish tasks that context overflow had previously cut short.
Patterns from real deployments
The same principles show up in projects I've built. In the Konkret system, the quote engine draws on 70 cost estimates totaling 3.2 GB - no context window can hold that, so the model gets only a few of the most similar estimates chosen by retrieval, not the whole archive. In GiftFinder, a model cascade splits the work so that a cheaper model processes large volumes of data, and a stronger one receives only selected, structured results for the final edit.
In both cases the rule is the same: the model at the end of the chain should see less, but better-chosen, information than the model at the start. On top of that comes a practice I use everywhere: tool results returned in a fixed structure rather than as raw text - I describe it in the post on structured outputs. A structured result takes less space and is easier to summarize or remove once it's no longer needed.
Common mistakes
- Throwing everything in "just in case". Whole documents instead of passages, all tools instead of the needed ones. Every excess element lowers quality.
- Variable elements at the start of instructions. A date, time or session ID in the first sentence invalidates the cache on every call.
- No plan for when the budget runs out. The agent runs until the window overflows and aborts or starts making mistakes, instead of summarizing and continuing.
- Summaries that lose details. Compression not tested on real runs can remove exactly the information the agent will need ten steps later.
- Multiple agents where one would do. Isolating context for tightly connected tasks increases cost and the number of hand-off errors.
A step-by-step implementation plan
- 1.Measure what's in the window. Log the full context of a few agent runs and check how much space instructions, tools, results and history take.
- 2.Put the fixed part at the start. Instructions and tool definitions without variable elements, so the cache works.
- 3.Limit tools to those needed for the task, or introduce tool search.
- 4.Fetch data on demand instead of loading everything up front; keep pointers in the window, not full contents.
- 5.Add notes outside the window for tasks longer than a dozen or so steps.
- 6.Introduce compression: clearing old tool results, and for long sessions summarizing history.
- 7.Test summaries on real runs, checking that the agent doesn't lose information it needs.
- 8.Consider sub-agents only for tasks that split cleanly, and do the cost math.
- 9.Monitor context length, cache hits and result quality over time.
---
I design and build AI agents that work reliably across many steps - from the context budget and on-demand retrieval, through compression and notes, to cost monitoring. I do this as part of AI app development, and the context engineering workshop is part of the AI Engineer course. Get in touch - I'll start by analyzing the full context of a few runs of your agent and listing what can be removed from it.
Worth reading next:
- AI Agent Memory - How to Make Your Chatbot and Agent Remember Users Across Sessions
- Prompt Engineering for Business - How to Talk to AI So It Actually Delivers
- Advanced RAG - Chunking Strategy, Hybrid Search and Reranking in Production
- Multi-Agent AI - When One Agent Isn't Enough and How to Build Agent Systems
/// RELATED_SERVICES
Need these concepts implemented? Explore the services related to this topic.
/// SOURCES
- 01Anthropic - Effective context engineering for AI agents
- 02Chroma - Context Rot: How Increasing Input Tokens Impacts LLM Performance
- 03Manus - Context Engineering for AI Agents: Lessons from Building Manus
- 04ACL Anthology - Lost in the Middle: How Language Models Use Long Contexts
- 05Anthropic - How we built our multi-agent research system
- 06Claude Blog - Managing context on the Claude Developer Platform
- 07Anthropic - Code execution with MCP: building more efficient AI agents
- 08LangChain - Context Engineering
/// RELATED_RECORDS
AI Video and Image Generation in Marketing - Tools, Costs and Copyright in 2026
OpenAI shut down Sora less than a year after the Sora 2 launch, and Veo 3.1 now costs $0.05 to $0.40 per second of video. Which AI video and image tools work in 2026, what an accepted asset really costs, who owns the copyright, how to label content since August 2026 and how to build a process with human review.
Social Media Automation with AI - One Piece of Content, Many Channels, Your Brand Voice Intact
LinkedIn's "Seems like AI slop" button was clicked over a million times in three weeks, and posts copy-pasted from a chatbot lose around 40% of their views. How to build an n8n pipeline that turns one piece into content for LinkedIn, a newsletter and X, keeps your brand voice and goes through human approval.
AI Contract Review - Risk Analysis, Version Comparison and a Pre-Signing Checklist
In the 2025 Vals benchmark, AI beat lawyers at data extraction and summaries but lost at marking up contracts (lawyers: 79.7%). How to build a contract review where AI finds the clauses and risks and a human decides - plus a clause playbook and protection against hidden instructions in documents.
Signal received?
Terminate
Silence
Initiate protocol. Establish connection. Let's build something loud.
