AI agent memory: what it is, the kinds that exist, and how to choose
AI agent memory is any information an agent can use that was not in its prompt when the task started: facts it saved, notes someone wrote for it, or a record it shares with other agents. Everything else about the topic follows from three questions: what gets stored, where it lives, and who can read it.

What AI agent memory is
A language model has no memory of its own between calls. Every request starts from the text you send it and nothing else. An agent is a model in a loop that calls tools, and agent memory is the set of mechanisms that put the right information back into that text: what the agent learned last week, what the team decided last month, what another agent found an hour ago.
A useful definition, short enough to quote: agent memory is information that survives the end of a session and comes back into a later one, either because the agent retrieves it or because the system loads it for the agent. Everything in this guide is a variation on how that information is chosen, stored, found and retired.
The rest of this page walks through the kinds of memory that exist in practice, where each one lives, the problems every memory system has to solve, and how to choose between them. It links out to longer pieces where a topic deserves one.
The context window is working space
The context window is the text the model sees on this call. It is large now, often hundreds of thousands of tokens, and that size makes it tempting to treat it as memory. It behaves more like a desk than a filing cabinet. Everything on it is visible at once, it is cleared when the session ends, and piling more on it has costs.
Three costs show up in practice. Every token in the window is paid for on every turn, so a long pasted history makes each step slower and more expensive. Models attend less reliably to the middle of a very long context than to its start and end, so a fact buried on page forty of a transcript is easier to miss than the same fact placed near the top. And when the window fills, the harness has to drop or compress something, which is exactly the moment an agent forgets an instruction it was given an hour ago.
Memory is what lets the window stay small and relevant. Instead of carrying everything, the agent carries what this step needs and fetches the rest when it becomes relevant.
The kinds of agent memory
The literature borrows words from psychology (working, episodic, semantic, procedural), and they are useful up to a point. In running systems the more practical split is by mechanism. These are the seven you will meet.
1. Working memory: what is in the context right now
The current conversation, the files the agent has read this session, the tool results it has seen. It is complete and exact, and it disappears when the session ends. Every other kind of memory exists to put something back here.
2. Session history: the transcript, kept
Many systems store the full message history of past sessions and let the agent search or reload it. Letta, for example, persists every message in a database, so an agent can retrieve an old exchange even after it has been evicted from the context window. A transcript is the most faithful record of what happened. It is also the noisiest: the one sentence that mattered sits among hundreds that did not.
3. Instruction files: memory a person writes
A markdown file in the repository that the agent loads at the start of every session. Claude Code reads CLAUDE.md; Codex and Cursor read AGENTS.md. The file holds build commands, conventions and rules you would otherwise repeat. It is cheap, reviewable in a pull request, and portable to any tool that reads the file. It has no idea which of its lines are still true, and it loads whole whether or not a line is relevant to the task. We covered how Claude Code handles these files, and its own auto memory, in how CLAUDE.md and auto memory work.
4. Extracted memories: facts the system pulls out
A model reads the conversation and writes down what seems worth keeping: "the user prefers TypeScript", "the staging database is read-only on Fridays". Mem0 is the best-known example: its docs describe sending messages through a model that pulls out key facts, decisions and preferences. Its current storage is additive, so new memories accumulate and a correction is an explicit update call on the old memory. Extraction is automatic, which is its strength. What gets kept depends on what the extractor judged important, which is its weakness.
5. Semantic recall: search by meaning
Store text as embeddings in a vector index, and at query time return the passages closest in meaning to the question. This is the retrieval half of most memory systems, including many that describe themselves as graphs. It finds "we moved jobs to the queue table" when you search for "background processing", which keyword search would miss. It returns the most similar passages, which are not always the most important or the most current.
6. Knowledge graphs, including temporal ones
Store entities and the relationships between them, so recall can follow links: this decision affects that component, which has this open task. A temporal graph also records time. Graphiti, the open source engine behind Zep, describes explicit bi-temporal tracking and temporal edge invalidation: when a fact stops being true its edge is marked invalid, and the history is kept, so the graph can answer both what is true now and what was believed last March.
7. Shared project records: memory for a team of agents
A single store that several agents and several people read and write, organized around a project rather than a user. Entries are typed (a decision, a lesson from a failure, a risk, a task) and linked to each other and to the code they describe. Letta's shared memory blocks, attached to several agents at once, are one form of this. A committed instruction file is a simple form of it too, shared through git. The distinguishing question is whether what one agent learned reaches the next agent that needs it, whichever tool that agent runs in.
Where memory lives, and who can read it
The kind of memory tells you what is stored. Where it is stored tells you who benefits, and that decides more about day-to-day usefulness than any retrieval algorithm.
| Where it lives | Who can read it | Survives a new machine? | Typical examples |
|---|---|---|---|
| Files in the repository | Anyone who clones it, any tool that reads the file | Yes, through git | CLAUDE.md, AGENTS.md, Cursor rules |
| A folder or database on your machine | The agent on that machine | No | Auto memory, local memory plugins |
| A hosted memory service | Whatever holds the API key | Yes | Memory APIs used by AI products |
| A shared project record | Every agent and person on the project | Yes | Team memory servers reached over MCP |
Two consequences are easy to miss. Memory stored inside one tool is invisible to every other tool: switch from one coding agent to another for an afternoon and the second one starts from nothing. And memory stored on one machine is invisible to your teammates, so each of their agents relearns what yours already knows.
The hard problems every memory system has
Storing things is easy. Five problems decide whether stored memory helps an agent or quietly misleads it.
Relevance: getting the right thing back
An agent working on the checkout page needs the note about the payment provider's rate limit, and does not need forty notes about the marketing site. Similarity search gets close; ranking by recency, importance and links between entries gets closer. When retrieval misses, a system that loads a whole file compensates by loading everything, which brings back the context window's costs.
Staleness: knowing when a fact stopped being true
In March you write down that jobs go through the queue table. In June you remove the queue table. If nothing marks the March note as retired, an agent in July reads it, trusts it, and rebuilds against a table that no longer exists. Systems handle this differently: temporal graphs invalidate the old edge, extraction systems let you update or delete a memory, and some records retire a fact when a newer decision replaces it or when a date passes. A plain file relies on a person noticing. This is the problem that matters most as memory grows, and we wrote a whole piece on what happens when a fact stops being true.
Conflicts: two agents, two versions
Once several agents write to the same memory, they will disagree. One records that the API returns dates in UTC, another that it uses local time. A useful system keeps both visible and makes the conflict obvious; a naive one keeps whichever was written last. The same problem appears with work: two agents picking up the same task at the same time is a memory problem too, because neither knew the other had started.
Cost: every recalled token is paid for
Memory that is injected into every prompt costs tokens on every turn. A good system spends that budget on the few entries that matter for this step, and keeps the rest one retrieval away. Measure what your memory adds to each prompt before and after you adopt it; it is the easiest number to get and the one people most often skip.
Privacy: what leaves the machine
Hosted memory means project knowledge is stored on someone else's servers. Check what is sent (summaries, full transcripts, source code), who can read it, and whether you can export and delete it. Local memory avoids the question and gives up sharing in exchange.
One agent, or many
Most memory systems were designed for one agent serving one user: a support bot that remembers a customer, an assistant that remembers your preferences. The memory belongs to the relationship between that agent and that person.
Software projects increasingly have many agents: one planning, one implementing, one reviewing, sometimes in three different tools, alongside several people. Their memory problem is different in kind. The unit is the project. The facts that matter are decisions and their reasons, lessons from failures, and who is working on what. And the failure that hurts is one agent repeating a mistake another agent already made and recorded.
The useful test for any memory system: does what one agent learned reach the next agent that needs it, in a form it can trust?
For a single agent, per-user memory is often all you need. For a team of agents, look for memory that every tool can reach, that distinguishes a decision from a passing remark, and that has a way to retire what is no longer true.
How to tell whether memory is helping
Every memory product has a demo in which memory helps. Your project is the only benchmark that counts, and three checks take an afternoon.
First, read what gets injected. Most systems can show you what they added to a prompt. Count how much of it was relevant to the task. If most of it is noise, the agent pays for it on every turn and may act on the wrong entry.
Second, change a fact and wait a day. Tell the agent the project moved from one database to another, start a fresh session tomorrow, and ask about the database. A system that serves the old fact beside the new one without saying which is current will eventually mislead an agent at a bad moment.
Third, switch tools or machines. Do some work in one agent, then open the same project in another agent or on another computer and ask what was done. This separates memory that belongs to a tool from memory that belongs to the project.
Measuring this rigorously is harder than it sounds. We wrote up our attempts, including the ways a benchmark can run perfectly and measure nothing, in how do you measure whether agent memory actually helps.
Memory for AI coding agents
Coding agents are where most people meet agent memory first, and they have a specific shape. The codebase itself is a large, reliable memory: an agent can read the code to learn what the system does. What the code does not say is why. Why the retry limit is three, why the team stopped using the old queue, which approach was tried and abandoned, which file is about to be rewritten by someone else. That missing layer is what coding-agent memory is for.
A typical setup grows in stages. It starts with an instruction file:
# Build and test
- Run the test suite before every commit.
# Conventions
- Money is stored in integer cents.
# Decisions
- Background jobs go through the jobs table. We tried a hosted queue in
2025 and removed it; see docs/adr/007.That covers one person on one project well. As more people and more agents join, the file grows, lines go stale, and learned knowledge stays trapped in whichever tool learned it. Then teams add a memory tool. We compared eleven of them, with their prices and tradeoffs, in the best memory tools for AI coding agents.
How to choose
Answer these in order. Each one rules options in or out.
- Who needs to read the memory? One agent for one user points to per-user memory. Several agents or several people on one project points to shared memory.
- Which tools must reach it? If you use more than one agent, the memory has to live outside all of them, reachable through files every tool reads or through a protocol such as MCP.
- How will a fact stop being served? Look for invalidation, supersession, expiry or at least easy editing. If the answer is "someone will notice", plan for the day nobody does.
- Should writing be automatic or deliberate?Automatic extraction captures more with no effort. Deliberate, typed entries capture less and are easier to trust.
- Where may the data live? Local, self-hosted or a managed service, depending on what your project allows.
Where Stele fits
We build Stele, so weigh this paragraph accordingly. Stele is a shared project record for AI coding agents: typed decisions, lessons, risks and tasks, linked to each other and to the code, that Claude Code, Cursor, Codex and any MCP client read before they act and write back to as they learn. A fact can be superseded by a newer one, expire on a date, or retire when its task closes. It is a hosted service, it asks agents to write entries on purpose rather than capturing everything, and it is young. The free plan covers public projects and one private one; pricing has the rest, and getting started takes a few minutes.
Frequently asked questions
What is AI agent memory?
AI agent memory is information that survives the end of a session and comes back into a later one, either because the agent retrieves it or because the system loads it for the agent. It includes instruction files, saved session history, facts extracted from conversations, searchable embeddings, knowledge graphs and shared project records.
Is the context window the same as memory?
No. The context window is the text the model sees on the current call. It is cleared when the session ends, every token in it is paid for on every turn, and very long contexts are followed less reliably. Memory is what lets the window stay small by bringing back only what the current step needs.
What are the main types of AI agent memory?
In running systems: working memory (the current context), session history (stored transcripts), instruction files (such as CLAUDE.md or AGENTS.md), extracted memories (facts a model pulls from conversations), semantic recall (vector search by meaning), knowledge graphs including temporal graphs, and shared project records that several agents and people read and write.
Can multiple AI agents share the same memory?
Yes, if the memory lives outside any single agent. A committed instruction file is shared through git. Some frameworks attach one memory block to several agents. A memory server reached over MCP can be read and written by agents running in different tools. Memory stored inside one tool or on one machine is not shared.
How do you stop agent memory from going stale?
Give every fact a way to stop being served. Temporal graphs mark an old fact invalid while keeping its history, extraction systems let you update or delete a memory, and some project records retire a fact when a newer decision replaces it, when a date passes, or when the related task closes. A plain file needs a person to notice and edit it.