Skip to main content

AI Agent Memory: Short-Term and Long-Term

Models remember nothing between requests. How agents get short-term and long-term memory, the three kinds worth storing, and how memory goes wrong.

IntermediateVerdeshell Team · 5 min read · Last reviewed

A model remembers nothing between requests. Short-term memory is managing what stays in the context window during a task; long-term memory is storing what matters outside the model and loading it back when relevant.

Key takeaways

  • Language models are stateless: an agent’s memory is whatever the application puts back into the context window.
  • Short-term memory is the current task’s conversation and tool results, managed by trimming and summarising as the window fills.
  • Long-term memory is stored outside the model — commonly divided into episodic (what happened), semantic (facts) and procedural (how to do things).
  • Memory can be written by the agent through a memory tool, or extracted by the application after each session.
  • Stale, wrong or injected memories persist across sessions, so memory needs review, expiry, per-user separation and an edit button.
Agent memory: the context window and long-term memory storesThe context window holds what the model sees in this request: instructions, memories loaded for this task, the conversation so far, and tool results. Below it, long-term memory is stored outside the model in three kinds: procedural memory, how to do things, which supplies the instructions; semantic memory, facts about users, the business and documents; and episodic memory, what happened in past sessions. Relevant semantic and episodic memories are loaded into the window, and at the end of a session the conversation is summarised and saved as episodic memory.SHORT-TERM — THE CONTEXT WINDOW FOR THIS REQUESTInstructionsLoaded memoriesConversation so farTool resultsfreeLONG-TERM — STORED OUTSIDE THE MODELProceduralhow to do things: prompts, skillsSemanticfacts: users, business, documentsEpisodicwhat happened in past sessionsloadretrieve when relevantsummarise + saveThe model itself keeps nothing between requests. Memory is what the application stores and puts back.
Memory is whatever the application puts back into the context window. Long-term memory lives outside the model and is loaded when relevant.

Hover or tap the diagram to replay the animation.

Models do not remember

A language model keeps nothing between requests. Every call carries everything the model will consider — instructions, conversation, documents, tool results — inside its context window. When a chat assistant seems to remember your earlier message, the application sent that message again.

So “agent memory” is a design question for the application: what to keep, where to keep it, and what to put back in front of the model at each step.

Short-term memory: the current task

Within one task, memory is the running record of messages and tool results. In a long agent run, that record grows with every step until it crowds out what matters, or exceeds the window entirely.

The usual remedies: clear old tool results once they have been used, summarise earlier turns into a short account of progress, and keep the goal, constraints and what has already been ruled out pinned at the top. Model providers now offer some of this on their side — Anthropic’s API, for example, can clear old tool results and summarise earlier turns automatically.

Doing this well is a large part of loop engineering: a long-running agent that loses the thread keeps working but stops making progress.

Long-term memory: three kinds worth storing

Anything that should survive the end of a session has to be stored outside the model. A 2023 research framework for language agents, CoALA, divides long-term memory into three kinds, and the split is useful in practice.

Episodic memory — what happened before: summaries of past sessions, outcomes of earlier tasks, what was tried and failed.

Semantic memory — facts: a customer’s preferences, details about the business, knowledge the agent has gathered. A retrieval index over your documents is semantic memory in this sense.

Procedural memory — how to do things: the system prompt, skill files and the code of the tools themselves.

How memory is written and read

Writing happens in one of two ways. The agent can decide for itself, using a memory tool to save a note when it learns something worth keeping. Or the application can extract memories after each session — summarising it, or pulling out stated preferences — without the agent choosing.

Reading is the mirror image: load relevant memories into the context at the start of a task, or let the agent search memory through a tool when it needs to. The second scales better; the first is simpler and more predictable.

Both appear in current products. Anthropic’s API provides a memory tool in which the model reads and writes files in a memory directory that your application stores. ChatGPT combines memories a user can view and delete with references to past chats. The 2023 MemGPT research proposed treating the context window like an operating system’s fast memory, moving information in and out of slower storage as needed — the idea behind much of this design.

How memory goes wrong

Stale memories: a preference or fact that was true in March is applied in October. Memories need dates and expiry.

Wrong memories reinforced: a mistaken note, once saved, is loaded into every future session and treated as fact.

Memory poisoning: a prompt injection that gets saved as a memory keeps influencing the agent long after the malicious content is gone.

Leakage and privacy: memories stored without separation between users or clients can surface in the wrong conversation, and personal data kept in memory falls under the same retention and deletion obligations as any other.

Design rules

Store less. Most tasks need the current context and a good retrieval index, not a growing pile of notes.

Make memory visible and editable: users and operators should be able to see what is stored and delete it.

Record where each memory came from and when, and expire what goes stale.

Keep memory separate per user and per client, and never store credentials or secrets in it.

Test with memory on and off. If the agent is not measurably better with it, it is not earning its risk.

Next in the pathLoop Engineering: Design the Loop, Not Every Prompt

Want this built properly?

We design and build AI systems for clients. Tell us the problem and we will tell you honestly whether AI — and which kind — is the right fit for it.