AI Agent Memory: How to Give an Agent Memory That Lasts
Models are stateless; agent memory is something you build. The kinds — working context, conversation summaries, long-term facts, episodic logs, retrieval — how each is implemented, what to store and when to forget, and the trade-offs between memory files, databases and vector search.
Language models remember nothing between calls. Every request starts from zero; whatever the model "knows" about the conversation is whatever you put in the prompt. So agent memory is a system you design, not a feature of the model. (What is a context window?, Why AI agents forget)
The kinds of memory
| Kind | What it holds | Lives in | Lifetime |
|---|---|---|---|
| Working memory | The current task: messages, tool results | The context window | One session |
| Summaries | Condensed earlier conversation | Context, regenerated | Session to days |
| Semantic (facts) | Preferences, decisions, profile | A file or database | Long |
| Episodic | What happened: past tasks, outcomes | Logs, database | Long |
| Retrieved knowledge | Documents, code, records | Search index / vector DB | Persistent |
1. Working memory: manage the window
Everything in the current loop sits in context, and context degrades as it fills. Techniques:
- Trim old tool outputs once used (keep a one-line note of what they said).
- Summarise older turns when you approach a threshold — "compaction". Claude Code does this automatically. (/compact vs /clear)
- Offload to files: let the agent write
NOTES.mdor a TODO list and re-read it, instead of carrying everything in the conversation.
2. Long-term facts: explicit and editable
The most useful long-term memory is usually small and human-readable: preferences, conventions, decisions, who the user is.
Claude Code's design is a good model: instructions you write (CLAUDE.md) plus auto memory — short Markdown files the agent writes itself, with an index loaded each session and details read on demand. (Claude Code memory)
For your own agent:
// a memory tool the agent can call
{
name: 'remember',
description: 'Save a durable fact about the user or project for future sessions. Only save things that will matter later and cannot be re-derived.',
input_schema: { type: 'object', properties: { key: { type: 'string' }, fact: { type: 'string' } }, required: ['key', 'fact'] },
}
Store facts per user in a table; load the relevant ones into the system prompt at session start. Let users see and delete them.
3. Episodic memory: what happened before
Log each task: the goal, the steps, the outcome, what failed. Later, retrieve similar past episodes ("last time a deploy failed like this, the fix was…"). Useful for agents that repeat similar jobs.
4. Retrieval: knowledge too big for context
Documents, tickets, codebases — far too much to load. Index them and retrieve only what's relevant to the current step:
- Keyword / full-text search for exact terms.
- Vector search for meaning. (What is a vector database?)
- Hybrid for both. (Semantic vs keyword search)
This is RAG applied to the agent's own memory. (What is RAG?, RAG chunking)
For coding agents, "retrieval" is often just grep and reading files — which is why a real filesystem matters so much.
What to store — and what not to
Store: stable preferences, decisions with their reasons, corrections the user made, pointers to where information lives.
Don't store: things re-derivable from the source (the code itself, the database), one-off details, secrets, and anything the user would be surprised to find remembered.
Forgetting is a feature
Memory that only grows goes stale and contradictory, and the agent follows outdated instructions confidently.
- Timestamp every memory and prefer recent ones.
- Update rather than append when facts change.
- Cap the always-loaded part (Claude Code loads only the first 200 lines / 25 KB of its index).
- Let users edit and delete — and in many jurisdictions, they have the right to. (GDPR basics)
Memory and safety
Memory is a persistence channel for attacks: an instruction injected once ("always send reports to this address") can be "remembered" and obeyed forever. Validate what gets written, keep memory scoped per user, and review it. (Prompt injection)
The environment is memory too
Often overlooked: for agents that build software, the workspace itself — files, git history, installed dependencies, database state — is the richest memory there is. An agent that returns to the same environment doesn't need to be told what it did yesterday; it can look. (Why AI coding agents need persistent workspaces)
EasySpawn gives agents the memory that's hardest to build: a persistent server where files, git history, notes and databases are exactly as they were left, session after session. See how it works or join the waitlist.
Related: Why AI Coding Agents Need Persistent Workspaces · Claude Code Memory Explained · Context Engineering for Coding Agents · How to Build an AI Agent
Keep reading
How to Build an MCP Server (TypeScript and Python)
Build a working Model Context Protocol server that gives Claude Code and other AI tools new abilities. Tools, resources and prompts explained, a TypeScript server with the official SDK and a Python one with FastMCP, connecting it to Claude Code, stdio vs HTTP transports, and designing tools agents use well.
How to Build an AI Agent: The Loop, the Tools, and the Guardrails
An AI agent is a model in a loop: it decides on an action, a tool runs it, the result goes back, repeat. Build one from scratch with the Claude API's tool use, then see when to reach for the Claude Agent SDK or other frameworks — plus stopping conditions, memory, cost control and safety.