Major

AI Agent Memory: Durable State Beyond the Context Window

An agent’s context window is not durable memory. See how databases, retrieval, and deterministic apps let work persist across runs, with governance built into execution.

Rahul Ramakrishnan
Diagram showing an AI agent using durable app data and logs across separate runs.

Key takeaways

  • A context window is working space for one task. It is not durable memory.
  • Retrieval returns candidates. An application record is the authority.
  • Memory design must define who may write, how long data lives, and what wins in a conflict.
  • Repeatable steps belong in deterministic apps, so the model reasons only where judgment is needed.
  • Persistence does not make data correct. Verify before any consequential action.

The actual definition of AI agent memory

AI agent memory is the information an agent can retain or retrieve across steps and runs to guide later decisions. It covers whatever sits in the prompt right now, whatever the agent can look up, and whatever the surrounding software has recorded about past work. It has no resemblance to human recall, and treating it as if it did is how teams end up debugging a personality instead of a data store.

IBM defines the term as an AI system's ability to store and recall past experiences. That is a fair starting point. It leaves open the question that matters in production: stored where, written by whom, and trusted how much?

We will use "memory" for the whole stack and "state" for the part of it that is authoritative. The distinction carries the rest of this article.

This piece covers memory architecture. It does not cover how to evaluate whether an agent's judgment is good, which is a separate and harder problem.

Why memory matters in production

A demo agent runs once and is judged on that run. A production agent gets interrupted. A deploy restarts the worker, an approval takes three days, a user returns on Thursday to a task they started Monday. If the only record of what the agent decided is a conversation that ended, the next run starts by rebuilding context from scratch, and it may rebuild it differently.

Three costs follow. The agent repeats setup work, which spends tokens on steps whose answers have not changed. It repeats questions to the user, which erodes trust fast. And it leaves no record you can audit, because the reasoning lived in a prompt that no longer exists.

The third cost is the one security teams notice. NIST's AI Risk Management Framework is organized around four functions, Govern, Map, Measure, and Manage, and each of them assumes you can say what a system knew and did. An agent whose memory is an unlogged context window gives you nothing to map or measure. For more on what persists after an agent request and why it matters, see what persists after an agent request.

Four pieces commonly called memory

People use "memory" for at least four different things. They have different lifetimes, different failure modes, and different owners. Conflating them is the root of most memory bugs.

Working context

This is the context window plus whatever task-local state the framework carries with it. LangChain's documentation describes short-term memory as the conversation retained within a single thread, saved by a checkpointer so the same thread ID can resume. It also notes that long histories can exceed the context window, and that even when they fit they can slow responses, raise cost, and distract the model. Teams respond by trimming or summarizing.

Both responses are lossy. A summary is the model's opinion of what mattered. Working context is the right place for the current task and the wrong place for anything you will need to defend later.

Retrieval and semantic memory

Semantic memory, in the common taxonomy, holds facts, concepts, and preferences. In practice it is an index over documents, notes, or extracted statements, queried by similarity. Products such as Mem0 package this as a memory layer with add and search calls.

Retrieval is good at finding relevant material in a large, loosely structured pile. It returns candidates ranked by closeness, not facts ranked by truth. Each result should carry provenance (where it came from and when) so the agent and the reviewer can weigh it.

Episodic history

Episodic memory is the record of what happened: this event, this action, this outcome. LangChain's memory post describes a loop of tracing agent behavior, analyzing it, and updating future instructions or skills. That loop produces learned behavior from history, and it depends on the history being captured faithfully in the first place.

An append-only event log is the cleanest form. Nothing is rewritten, so you can reconstruct what the agent saw and did at a given moment. This is where agent observability meets memory: the log that lets you inspect a run is the same log the next run can read.

Durable application state

This is the record of truth. The account's current health score, the invoice's reconciliation status, the ticket's owner. It lives in a database with a schema, permissions, and a defined set of operations that change it.

The key property is authority. When retrieval says one thing and the application record says another, the record wins, and the system knows why. The agent reads it, proposes changes to it through defined operations, and does not own it.

IBM and LangChain also use the terms short-term, long-term, episodic, semantic, and procedural. Taxonomies vary by author, and no single set is universal. Short-term and long-term describe lifetime. Episodic and semantic describe content. Procedural describes learned behavior. The four pieces above cut across them by asking a different question: who is the authority?

| Layer | Lifetime | Role | Authority | Governance | |---|---|---|---|---| | Context window | One task or thread | Working space for current reasoning | None; a scratchpad | Hard to inspect after the fact unless captured | | Retrieval store | Until re-indexed or deleted | Find relevant reference material | Advisory; ranked candidates | Needs provenance, access filters, deletion path | | Event history | Retention policy | Record what happened | Authoritative for what occurred | Append-only, attributed, exportable | | App database | Life of the record | Hold current truth | Authoritative for current state | Schema, role-based access, audit log |

A compact resume example

The following is a constructed illustration of an account health workflow. It is not a customer case.

Monday, an agent is asked to review a mid-sized account whose usage dropped. It queries the health app and gets the account record: owner, current risk level, last-updated timestamp. It retrieves two prior support threads from the index and notes their dates. It reasons that usage fell after a pricing change, which takes judgment, and it drafts outreach for the owner's approval. The approval is pending.

What is persisted? The draft and its status in the app database. An event appended to the log: agent reviewed account, retrieved these two threads, proposed this action. The risk level, unchanged until a person or a defined rule changes it.

Thursday, the owner approves. A new run starts with no conversation at all. The agent reads the record, sees the pending-then-approved status and the logged reasoning, sends the outreach through the app's send operation, and appends the result. It re-derives nothing, because nothing needs deriving.

The flow, in order: task, retrieve current state, reason where needed, execute app logic, append event.

How updates, deletion, and staleness should work

Memory without lifecycle rules becomes a liability. Four rules cover most of it.

Updates go through defined operations. The agent does not free-write to a record. It calls an operation that validates the change, applies permissions, and logs who made it.

Stale data carries a timestamp. Every retrieved item and every record shows when it was last confirmed. An agent that sees a three-month-old risk level should treat it as a hint and refresh it.

Retention is set per layer. Context vanishes with the task. Event history follows your audit retention policy. Records live as long as the business object does.

Deletion has to reach every layer. If a customer's data must be removed, it must go from the record, the event history where policy allows, and the retrieval index. A vector index that no one remembers to purge is how sensitive data survives a deletion request.

One caveat belongs in every design review. Retrieved data can be wrong or out of date. Before any consequential action, the agent should verify against the authoritative record, and a person should review the decisions where the stakes justify it.

Common misconceptions

A vector database is not a system of record. It is a good tool for semantic lookup, and nothing here argues otherwise. But it has no notion of current truth, no transaction semantics, and usually no per-record permissions that match your business roles. Teams that put authoritative facts in it discover the gap when two similar records disagree.

More context does not produce reliable state. A larger window lets the model see more at once. It does not persist anything after the run, and it does not tell you which of two conflicting statements is current. LangChain's own documentation lists distraction as a cost of long histories.

Persistence does not equal correctness. A wrong fact saved durably is a durable wrong fact. This is why authority, update rules, and verification matter more than the storage engine.

Memory does not make an agent remember the way a person does. The agent reads what the system gives it. Better memory means better-designed reads and writes.

The position we hold: memory design is the definition of authority, retention, and update rules, and retrieval quality is one input to it. Most memory products optimize the retrieval half. The harder half is deciding what the system is allowed to believe.

What we're building at Major in response

Memory should be an architectural property of the work. A prompt convention that says "remember to check the notes" fails the first time someone changes the prompt.

Major is the enterprise platform where agents build the software they run on. When an agent works out how to handle a repeatable part of a task, it builds an app for that part. The app has a managed database, storage, and logs, so the state the model used to carry in a context window lives in a place with a schema, permissions, and an audit trail. The agent reasons once to build it, then runs it. Reason once, run forever.

That gives the division of labor from the example above. The app executes the repeatable steps the same way every time and preserves the history. The agent reasons where judgment is needed: why did usage drop, what should the outreach say, does this case need an owner. Both live inside the governance layer, with scoped credentials and governance at the point of action. The system-level components are laid out in our piece on AI agent architecture.

This does not remove the model. It does not make stale data fresh, decide your retention policy, or replace human review of consequential calls. A workflow short enough to live in one prompt does not need any of this. The case for Major gets stronger as the workflow gets longer, crosses more systems of record, and has to be explained to an auditor months later.

If you want to see what an agent-built app with its own database and event log looks like in practice, see how Major keeps agent state in governed apps.

Related articles

Frequently asked questions

Do AI agents have memory?
Agents can use short-lived context inside a task and external persistence across tasks, such as databases, files, and retrieval indexes. A stand-alone model does not automatically retain a user's history between sessions. Whatever continuity an agent shows comes from the system around it, which decides what is stored and read back.
How do AI agents manage memory?
They capture events and outputs, store them in a context window, retrieval index, event log, or application database, then read back what the current task needs. Good designs add update rules, retention, and access policy. Retrieval supplies candidates, while the authoritative application record settles conflicts and gets verified before consequential action.
What are the four types of memory that every AI agent needs?
Taxonomies vary, and no single set is mandatory. A useful working split is working context for the current task, retrieved reference knowledge, episodic history of past events, and durable application state that holds authoritative records. Each has a different lifetime and a different owner.
Is a vector database enough for AI agent memory?
No. A vector database is strong at semantic lookup across loosely structured material. It returns ranked candidates, not current truth, and typically lacks transactions and role-based permissions matched to business records. Pair it with structured application state and an event log for authority and audit.