Context Engineering for Coding Agents: Beyond Prompt Engineering
Context engineering is designing everything an agent sees — instructions, tools, files, history — not just the prompt. Why context is a finite budget, context rot, just-in-time retrieval, progressive disclosure with skills, note-taking and compaction, subagents, and how to apply it to Claude Code.
"Prompt engineering" was about wording a single request well. With coding agents that run for hours, call dozens of tools and read hundreds of files, the wording of your request is a small part of what determines the outcome. What matters is everything in the model's context window at each step — and that's the subject of context engineering.
Anthropic's engineering team described it as finding the smallest set of high-signal tokens that maximises the likelihood of the outcome you want. This article turns that into practice for coding agents.
Context is a budget, not a bucket
Every model has a context window — today often a million tokens. It's tempting to treat that as "room for everything". Three reasons not to:
- Attention is finite. As context grows, models get worse at finding and using any particular piece of it. This degradation is often called context rot: the instruction you gave at the start competes with 400,000 tokens of logs, file contents and tool output.
- Every token costs on every turn. An agent resends its context with each step. Irrelevant material is paid for again and again. (How to keep Claude Code costs down.)
- Noise misleads. An outdated file version, a wrong guess from an hour ago, or a stale error message in context can steer later decisions.
So the goal isn't maximum context; it's maximum signal per token.
The components of an agent's context
| Component | Examples | Engineering question |
|---|---|---|
| System prompt & standing instructions | CLAUDE.md, rules |
Is every line earning its place on every turn? |
| Tool definitions | Built-in tools, MCP servers | Does the agent need all of these, right now? |
| Retrieved content | Files read, search results, docs | Was it fetched because it's needed, and only the relevant part? |
| Tool results | Test output, command output, logs | Is the agent reading 10,000 lines to find 3? |
| Conversation history | Past turns, decisions | Is old history still useful, or should it be compacted? |
| External memory | Plans, notes, specs on disk | What must survive beyond this context? |
Each is a lever.
Principle 1: Keep standing context small and precise
Standing instructions are loaded on every turn. Keep them to what's universally relevant: how to run, build and test; non-obvious conventions; hard rules. Anthropic suggests keeping CLAUDE.md under about 200 lines.
- Specific beats general. "Use
pnpm, nevernpm" beats "follow our package manager conventions". - Explain the why for non-obvious rules; models generalise better from reasons.
- Delete what the agent already does well without being told.
(How to write a CLAUDE.md that actually helps.)
Principle 2: Progressive disclosure
Instead of loading everything up front, give the agent pointers and let it load detail when relevant:
- Skills carry a short name and description in context; the full instructions load only when the skill is used. Move workflow-specific guidance (migrations, release process, PR review) out of
CLAUDE.mdinto skills. (Claude Code skills.) - MCP tool search — Claude Code defers MCP tool definitions so only names enter context until a tool is needed. Still, disable servers you don't use. (Connecting MCP servers.)
- Docs as files — a
docs/architecture.mdthe agent reads when touching architecture, rather than a summary pasted into every session.
Principle 3: Just-in-time retrieval
Agents work best when they explore like an engineer — list directories, grep, read the relevant function — rather than being handed a dump of the codebase. Lightweight references (file paths, function names, a query) are cheap; the agent loads content only when it decides it needs it.
Help it explore efficiently:
- a clear directory structure and descriptive file names,
- a short map in
CLAUDE.md("API routes insrc/server/routes/, one file per resource"), - typed code and code-intelligence tooling, so "go to definition" replaces reading five candidate files. (Using Claude Code on a large codebase.)
Principle 4: Make tool output high-signal
Tool results are often the biggest context consumer. A full test run, a verbose build, or a 5,000-line log can each swamp the window.
- Filter at the source. Run tests with reporters that show failures only; tail logs rather than cat them.
- Use hooks to preprocess. A hook can rewrite test commands to keep only failing output, or grep logs for errors before Claude sees them. (Claude Code hooks.)
- Design tools to return summaries with an option to fetch detail, rather than everything by default.
Principle 5: Externalise memory
The context window is working memory. Durable knowledge belongs on disk:
- Plans and specs before long tasks (spec-driven development),
- Progress notes / TODO files the agent updates as it goes,
- Decision records for choices that future sessions must respect.
With memory on disk, compaction and clearing become safe: summarise or start fresh, and point the agent back at the files. (/compact vs /clear.) This is also why the environment matters: notes, partial work and the running app need to persist between sessions for an agent to pick up where it left off. (Why your AI agent keeps forgetting.)
Principle 6: Isolate with subagents
Some work is reading-heavy but produces a small result: "find every place we call the payments API", "summarise why these 40 tests fail". Delegate it to a subagent with its own context window. The subagent burns through the reading; only its concise report returns to the main conversation. (Claude Code subagents.)
The trade-off: subagents cost their own tokens, and handing off loses nuance. Use them for well-defined, separable research — not for tightly coupled work.
Principle 7: Compact deliberately
Long-horizon tasks will exceed what's useful to keep verbatim. Compact at milestones, with explicit instructions about what to preserve (current task, decisions and reasons, exact failing tests, files changed), rather than letting the window fill. Set an earlier auto-compact threshold if sessions routinely run long.
Measure, don't guess
Context changes are hypotheses. Check them:
/contextshows what's occupying the window right now,/usageshows what's driving consumption over time,- a small set of representative tasks, re-run after changes, tells you whether quality actually improved. (Evals for coding agents.)
The summary
- Context engineering designs everything the agent sees, not just the prompt.
- Treat context as a budget: small standing instructions, progressive disclosure, just-in-time retrieval.
- Filter tool output, externalise memory to files, isolate heavy reading in subagents, compact at milestones.
- Measure with
/context,/usageand a small eval set.
EasySpawn gives Claude Code a persistent server, so the external memory context engineering depends on — plans, notes, specs, the codebase and the running app — is still there in the next session and from any device. See how it works or join the waitlist.
Related: What Is a Context Window? · Prompt Injection in Coding Agents · Why AI Coding Agents Need Persistent Workspaces · How to Write Good Prompts for AI Coding Tools · How Coding Agents Work
Keep reading
How to Resume a Claude Code Session (and What Resuming Can't Bring Back)
claude --continue and claude --resume reopen yesterday's conversation in seconds. But a resumed session restores the conversation, not the world it was working in. The commands, the habits that make resuming reliable, and the gap between conversation state and environment state.
How to Write a CLAUDE.md That Actually Changes What Claude Does
Most CLAUDE.md files are either empty or a wall of generic advice Claude would have followed anyway. What to put in one, what to leave out, how the files load, and how to tell whether yours is working.