How to Build an AI Agent: The Loop, the Tools, and the Guardrails
An AI agent is a model in a loop: it decides on an action, a tool runs it, the result goes back, repeat. Build one from scratch with the Claude API's tool use, then see when to reach for the Claude Agent SDK or other frameworks — plus stopping conditions, memory, cost control and safety.
Strip away the hype and an AI agent is a simple structure:
loop:
model looks at the goal + everything so far
model either answers, or asks to call a tool
your code runs the tool, appends the result
The model decides what to do next; your code does it and keeps the loop safe. (What is an AI agent?, How coding agents work)
Step 1: define tools
Tools are functions with a name, description and input schema. The model reads the descriptions to decide when to call them. (What is function calling?)
import Anthropic from '@anthropic-ai/sdk'
const client = new Anthropic()
const tools: Anthropic.Tool[] = [
{
name: 'search_orders',
description: 'Search a customer\'s orders by email. Returns up to 10 recent orders.',
input_schema: {
type: 'object',
properties: { email: { type: 'string' } },
required: ['email'],
},
},
{
name: 'issue_refund_request',
description: 'Create a refund REQUEST for a human to approve. Does not refund directly.',
input_schema: {
type: 'object',
properties: { orderId: { type: 'string' }, reason: { type: 'string' } },
required: ['orderId', 'reason'],
},
},
]
async function runTool(name: string, input: any) {
if (name === 'search_orders') return db.searchOrders(input.email)
if (name === 'issue_refund_request') return db.createRefundRequest(input.orderId, input.reason)
throw new Error(`Unknown tool ${name}`)
}
Step 2: the loop
async function agent(goal: string) {
const messages: Anthropic.MessageParam[] = [{ role: 'user', content: goal }]
for (let step = 0; step < 10; step++) { // hard cap on steps
const res = await client.messages.create({
model: 'claude-sonnet-5-5',
max_tokens: 2048,
system: 'You are a support agent. Use tools to look things up. Never promise refunds; request them.',
tools,
messages,
})
messages.push({ role: 'assistant', content: res.content })
if (res.stop_reason !== 'tool_use') {
return res.content.filter(b => b.type === 'text').map(b => b.text).join('\n')
}
const results: Anthropic.ToolResultBlockParam[] = []
for (const block of res.content) {
if (block.type !== 'tool_use') continue
try {
const output = await runTool(block.name, block.input)
results.push({ type: 'tool_result', tool_use_id: block.id, content: JSON.stringify(output) })
} catch (e) {
results.push({ type: 'tool_result', tool_use_id: block.id, content: String(e), is_error: true })
}
}
messages.push({ role: 'user', content: results })
}
throw new Error('Agent hit the step limit')
}
That's a working agent. Everything else is making it reliable.
Step 3: the parts that make it production-worthy
Stopping conditions
A step cap, a token or cost budget, and a timeout. Agents can loop. (Exponential backoff for API errors.)
Permissions and guardrails
- Tools enforce permissions in code — the model asking for another customer's orders must get nothing. (IDOR explained)
- Irreversible actions become requests a human approves, as above. (AI guardrails)
- Treat tool output from external sources as untrusted — it can contain instructions. (Prompt injection)
Context management
Long tasks overflow the context window. Trim old tool results, summarise, or store working notes in a file or database and read them back. (AI agent memory, Context engineering)
Structured final output
If code consumes the result, require structured output. (Structured output from LLMs)
Observability
Log every step: the model's choice, tool inputs, outputs, tokens, latency. Agent bugs are invisible without traces. (What is OpenTelemetry?)
Evaluation
Build a set of test tasks with expected outcomes and run them whenever you change the prompt, tools or model. (Evaluating LLM outputs)
Framework or from scratch?
| Option | Use when |
|---|---|
| From scratch (above) | Small agents, full control, learning |
| Claude Agent SDK | You want Claude Code's loop — file tools, bash, subagents, compaction, permissions, MCP — as a library (Claude Agent SDK) |
| Vercel AI SDK | TypeScript web apps, multi-provider, streaming UIs (Vercel AI SDK) |
| LangGraph and others | Complex graphs of steps and explicit state machines |
Start from scratch once to understand the loop, then adopt a framework for the parts you'd otherwise rebuild.
Where to run it
Agents that work for minutes or hours don't fit short serverless timeouts. Run them as background jobs on a server, with a queue and persisted state so a crash doesn't lose the work. (Background jobs, BullMQ tutorial)
EasySpawn gives agents somewhere to live: a persistent server where long-running jobs, their database and their logs stay put, with Claude Code to help build them. See how it works or join the waitlist.
Related: What Is an AI Agent? · The Claude Agent SDK · What Is Function Calling? · How to Build an MCP Server
Keep reading
MCP vs API: What's the Difference, and Do You Need Both?
An API is how programs talk to a service. MCP (Model Context Protocol) is a standard way for AI tools to discover and use those services. How they relate, why MCP servers usually wrap APIs, when to use each, and what MCP adds and costs.
Exponential Backoff and Jitter: Retrying Failed Requests Properly
When an API call fails with a timeout, 429 or 503, retrying immediately makes things worse. Exponential backoff waits longer after each failure, and jitter spreads retries out. Which errors to retry, Retry-After, a TypeScript implementation, idempotency, and retrying LLM API calls.