XavierFok
← all posts

Giving an AI agent memory that survives between runs

2026-08-15 · by Xavier Fok

# Giving an AI agent memory that survives between runs

A language model forgets everything the instant a run ends. You can have a long conversation where it seems to hold your entire situation in mind, and the next time you start it, blank slate. It remembers nothing about you, your last task, or anything it worked out an hour ago. For a casual chat this is harmless. For an agent whose job today builds on what it learned yesterday, the forgetting breaks everything. I run agents that need context across days, so I had to solve this properly, and the honest solution is far less magical than the word memory suggests.

Memory lives outside the model

The misconception to clear first, because it shapes everything after it. The model has no place inside it where your facts persist between runs. It can no more remember across runs than a calculator remembers your last sum. When people talk about giving an agent memory, what they are actually describing, whether they realize it or not, is storage wired up outside the model, with the right pieces fed back in at the start of each run.

The model stays forgetful. You build the remembering around it. Once that reframing lands, the mystique drains away and you are left with a normal engineering question: what to store, where to store it, and how to feed it back in.

A file, read first and written last

The version I actually use for most of my agents is a notes file, and it is exactly as humble as it sounds. At the start of a run, the agent reads the file, so the first thing it sees is the relevant context from before. At the end of the run, it writes back anything worth carrying forward. Read at the start, write at the end. That loop is the entire memory system.

No exotic database, no special infrastructure. The file is the memory, and the model is the thing that reads it and adds to it. An agent with this loop effectively remembers across days and weeks, and for the majority of what I build, this is the whole answer.

The instinct to keep everything, and why it fails

The obvious move once you have the loop is to remember everything. Why discard anything? Keep appending, and over time the agent accumulates a rich history of all it has ever done. I went down this road, and it fails for two reasons that compound.

The first is cost. Everything in the notes file gets fed to the model at the start of every run, and models charge by the amount of text they read. The file on disk costs nothing to keep. The file fed into every run is a tax, sized by how large the file has grown, paid again on each run. A small, tightly pruned file is a tiny tax. A file holding months of appended history is a large one, and if nothing ever prunes it, it grows without limit, which means every run costs more than the one before it, forever. A cost that rises monotonically with no ceiling is a genuinely bad shape for a cost to take. The moment you internalize that memory is read on every run, you stop wanting it big and start managing its size deliberately.

The second reason is worse than the money.

Stale facts poison the context

An accumulating memory fills with things that used to be true. Say the agent recorded a fact three weeks ago, correct at the time. The situation has since changed, and the old fact still sits in the file because nothing ever removes anything. Now every run begins with the agent reading that stale fact as if it were current, and acting on it with full confidence. As more history piles up, old facts contradict newer ones, and the agent has no reliable way to tell which version holds.

A memory full of outdated truths is worse than no memory at all, because a blank slate is at least honest about what it does not know. The stale file supplies confident, wrong context instead.

This lesson cost me a real failure. I let one agent's notes file grow freely, on the theory that more context could only help. One day it confidently acted on something that had stopped being true weeks earlier, because the dead fact was still in its memory, read in fresh that morning as if current. Nothing was broken. The model did exactly the right thing with what it was given. I had simply kept a fact past its expiry, and that class of failure is maddening to debug because every component is working correctly.

The skill is curation

Storing things is trivial. The actual skill in agent memory is deciding what deserves to be remembered, and just as much, what to forget or update. The notes file has to stay small and current. When something changes, the old version gets replaced rather than left sitting beside the new one. When something stops mattering, it gets removed. The file should read like a tidy summary of what is true and important right now, never like a transcript of everything that ever happened.

The comparison I hold in my head is a person's working notes. Nobody keeps every scrap of paper forever. You keep one clean page of what matters, and you cross things out when they change. The crossing out is the step people skip, and it is the step that makes memory useful. Good memory is mostly good forgetting.

Two mechanics make the curation practical. The first is structure. Give the memory defined slots: the current state of the task, the few durable facts that matter, the open items. When memory has a shape, an update means replacing a slot's value, and replacement is precisely what stops stale facts from accumulating in corners. The second is compression. Every so often, have the agent summarize its own notes, boil the sprawl down to what still matters, and discard the rest. Both mechanics attack the same enemy, unbounded growth.

There is a real tension underneath this, worth naming plainly. Remembered context genuinely makes an agent more capable, up to a point, since an agent that knows the relevant history beats one starting cold. Every remembered fact also costs something on every future run and carries a risk of going stale. So memory is a budget you spend deliberately. The question for each fact is whether remembering it will help future runs more than it costs them. Most facts fail that test. The few that pass are the ones worth keeping current.

The heavier version, and when you actually need it

At some scale a notes file stops being enough, and people ask about the serious machinery. If an agent must remember more than you would ever want to feed in at once, the answer is a searchable store. Everything gets kept in a form the agent can query, and each run pulls back only the few pieces relevant to the task at hand instead of reading all of memory every time.

That technique is real and useful at genuinely large memory sizes. I want to be equally honest that most agents never need it. For nearly everything I build, a small file kept current and read at the start is sufficient, and reaching for the searchable version early just adds a heavy thing to maintain for an agent that needed a text file. People jump to the complicated system because it sounds more serious. Match the machinery to the actual need, and graduate only when the simple file truly cannot hold what you require. The same discipline applies at either scale: keep what is current, drop what is stale, and treat every remembered fact as a recurring cost.

Disciplined note taking, nothing more

The pattern here connects to something I keep relearning across everything I build. The model arrives ready and capable, and the part you build is the careful system around it. For memory, that system is disciplined note taking. Read the right context in at the start. Do the work. Write back only what genuinely matters, updating and removing as things change. Done with discipline, the loop gives an agent something that behaves like memory across weeks. Done by hoarding, it gives you a slow, expensive agent confidently acting on facts that died a month ago.

So the next time someone describes long-term memory as a property of a model, you will know what actually sits under the hood. Storage someone wired up. A habit of reading it first and writing it last. The discipline to keep it small and current. The model remains as forgetful as ever, and the remembering is yours to build and yours to curate. The curating is the whole skill.

More breakdowns like this are on the [home page](/).

Get new guides and videos first — join the Telegram channel.