How Coding Agents Work: The Loop, the Tools, and the Context Behind Them
Under the hood, Claude Code, Codex and similar agents are a model in a loop with tools and a carefully managed context. The agent loop, tools, file editing, search, context compaction, subagents, permissions and sandboxing — and what it means for you.
From the outside, a coding agent looks like magic: you describe a bug, and minutes later it's found, fixed and tested. Inside, the architecture is surprisingly small. Understanding it makes you much better at using agents — and at deciding where they should run.
The loop
Every coding agent is a variation on this:
messages = [system prompt, tool definitions, user task]
loop:
response = model(messages)
if response has no tool calls: break # done — final answer
for each tool call in response:
check permissions / hooks
result = execute(tool call)
messages.append(tool result)
The model never touches your files. It emits structured tool calls ("call Read with path src/cart.ts"); the harness — the agent program — executes them and feeds results back as new messages. The model then decides the next step with that new information. (What is function calling?, what is an AI agent?)
That's the entire control flow. Everything else is about which tools exist, what goes into the context, and what's allowed.
The tools
A typical coding agent's toolset:
| Tool | Does |
|---|---|
| Read | Read a file (often with line ranges) |
| Glob / list | Find files by pattern |
| Grep | Search file contents (usually ripgrep underneath) |
| Edit | Replace an exact string in a file |
| Write | Create or overwrite a file |
| Bash / shell | Run commands: tests, builds, git, package managers |
| Web fetch / search | Read docs and issues |
| Task / subagent | Delegate a sub-task to a fresh agent |
| MCP tools | Anything external: databases, GitHub, browsers, your APIs (What is MCP?) |
Each tool has a name, a description and a JSON schema for its input. Tool descriptions are effectively part of the prompt — models choose tools based on them.
Editing files
Rewriting whole files is slow, expensive and error-prone for large files. Most agents use exact string replacement: the model provides the old text (which must match exactly and uniquely) and the new text. If the match fails, the tool returns an error and the model re-reads the file and retries. Some systems use unified diffs or apply models instead. The common thread: edits are verified mechanically, and failures go back to the model as information.
Search instead of indexing
Many terminal agents don't pre-index your repository into embeddings. They search on demand with glob and grep — like a developer would — reading only what's relevant. This avoids stale indexes and works on any codebase immediately; the cost is more tool calls on very large repos. Some tools add semantic indexes on top. (Using Claude Code on a large codebase)
The shell is the universal tool
Bash is what turns a code editor into an engineer: run the tests, read the error, run the type-checker, start the dev server, use git, query the database. Feedback from real execution is the single biggest reason agents outperform one-shot code generation. It's also the most dangerous tool.
The context
The model sees only what's in its context window: the system prompt, tool definitions, project instructions (like CLAUDE.md), the conversation, and every tool call and result so far. (What is a context window?)
Long tasks fill it with file contents and command output. Harnesses manage this:
- Truncating huge tool outputs and telling the model how to get more.
- Compaction: when the context is nearly full, summarise the conversation so far and continue from the summary. Details can be lost — which is why instructions that must persist belong in files, not just chat. (/compact vs /clear)
- Prompt caching: the stable prefix (system prompt, tools, early conversation) is cached, so each turn of the loop doesn't re-pay for it. Agents make dozens to hundreds of model calls per task; without caching they'd be slow and expensive. (Prompt caching explained)
- Persistent memory in files: CLAUDE.md, AGENTS.md, notes and task lists that survive compaction and new sessions. (How to write a CLAUDE.md, why agents forget)
Good context engineering — giving the agent the right information, and keeping noise out — is most of what separates good agent setups from bad ones. (Context engineering for coding agents)
Subagents
A subagent is the same loop started with a fresh context, a focused prompt and often a restricted toolset. The parent delegates ("find every place we parse dates and report back"); the subagent burns through dozens of searches in its own context and returns a short summary. The parent's context stays clean. Subagents can also run in parallel. (Claude Code subagents)
Planning and todo lists
Many agents maintain an explicit task list as a tool ("TodoWrite"), updating it as they go. It's a cheap way to keep long tasks on track after compaction, and makes progress visible to you. Plan modes go further: research and propose without editing until you approve. (Claude Code plan mode)
Permissions, hooks and sandboxes
Because tools have real effects, the harness wraps execution in layers:
- Permission rules — allow, ask or deny per tool and pattern (
Bash(npm test)allowed,Bash(rm -rf *)denied). (Permission modes) - Hooks — your code runs before/after tool calls and can block them deterministically. (Claude Code hooks)
- Classifiers — some modes use a second model to review risky actions.
- Sandboxing — OS-level restrictions on filesystem and network access for shell commands.
- The environment itself — what the agent's user account, credentials and network can reach.
Layers 1–4 reduce mistakes. Layer 5 is the one that holds when everything else fails — including against prompt injection, where text the agent reads (a README, an issue, a web page) tries to redirect it. (Prompt injection in coding agents)
What this means in practice
- Give it feedback loops. Tests, type-checkers, linters and a runnable app turn guesses into verified changes.
- Write durable instructions into files, not just chat, because context gets compacted.
- Keep tasks scoped. Smaller tasks keep the context relevant.
- Narrow tools and permissions to what the task needs.
- Choose the environment deliberately. The agent can do whatever its environment allows. Running it where production credentials, your SSH keys and your personal files are reachable is a choice. (Sandbox: the wrong abstraction?, run AI-generated code safely)
- Persistence matters for long work. Agents that keep their workspace — files, installed dependencies, running services, git state — between sessions pick up where they left off. (Why agents need persistent workspaces)
The summary
- A coding agent is a model in a loop: tool call → harness executes → result → next step.
- Core tools: read, search, edit (exact replacement), shell, plus MCP for everything else.
- Context management — truncation, compaction, caching, file-based memory — makes long tasks possible.
- Subagents and todo lists keep big tasks focused.
- Permissions, hooks and sandboxes limit mistakes; the environment is the final boundary.
EasySpawn gives coding agents a persistent, isolated server of their own — repo, dependencies, running app and database — so the loop has real feedback and a safe boundary, and work continues after you close your laptop. See how it works or join the waitlist.
Related: What Is an AI Coding Agent? · The Claude Agent SDK · Context Engineering for Coding Agents · Evals for Coding Agents
Keep reading
Securing MCP Servers: Threats and Controls for Tool-Connected Agents
An MCP server turns a model's text into real actions against real systems. The threat model — tool poisoning, prompt injection via tool output, confused deputies, token passthrough, DNS rebinding on local servers, over-broad scopes — and the controls for building and deploying MCP servers safely.
The Sandbox Is the Wrong Abstraction for AI Coding Agents
The industry settled on ephemeral sandboxes for AI agents — isolated, disposable, destroyed after each task. That's exactly right for running untrusted code and exactly wrong for building software. Here's the distinction that matters.