23 July 2026 / Applied AI

What I changed in agent memory

I now keep working context separate from durable records and searchable history. The change came from watching old material enter sessions where it was technically related but useless to the task in front of the agent.

Rows of library shelves filled with books beside a sunlit aisle.
Photo: Zoshua Colah (opens in a new tab)

I now keep working context separate from durable records and searchable history. The change came from watching old material enter sessions where it was technically related but useless to the task in front of the agent. Similarity is a weak reason to put something into the active context. A past conversation can mention the same project or tool and still describe an abandoned approach. A detailed result from last year may rank above a short current decision. When all stored material competes for the same prompt space, retrieval can make the agent less certain about the work it should actually do. Working context is the small set of information needed to continue the present task. It includes the objective, current state, constraints, relevant files and the next check. It changes quickly and should be easy to replace. I treat a handoff as the main entry point. The agent receives the current task record and follows links when it needs more detail. It does not begin by searching every conversation that mentions the project. This keeps incomplete ideas and old debugging output out of the initial prompt. Working context also needs an end. When the task finishes, some of it becomes an artefact or durable decision. The rest can remain in the session record without being promoted. Temporary command output, repeated explanations and a plan that was later discarded do not need to follow every future session.

The simplest check is whether a note would help someone perform the next action today. If it only explains how the work once unfolded, it probably belongs in searchable history. Durable memory holds information the agent should rely on across tasks: an approved preference, a current system fact, a standing constraint or a decision that still applies. These records need more structure than a transcript fragment. Each record should say what it is about, where it came from, when it became valid and how it can be corrected. A project decision should link to the issue, specification or approval that supports it. A user preference should retain its scope. "Use Australian English for this publication" should not silently become a rule for every person and project. I also keep the actor visible. A claim from a user, an observation from a tool and an inference made by an agent have different weight. Storing them all as plain facts makes later retrieval look more authoritative than the evidence allows. Corrections should supersede old records without erasing the history. The current record points to the newer decision, while the older one remains available for understanding past work. Active retrieval should prefer the current status unless the task specifically asks about the earlier period.

Session transcripts, old handoffs, logs and completed task notes still have value. I keep them searchable because they can answer questions about how a problem was investigated or why a change happened. They do not need automatic admission to the active prompt. Search results should show enough metadata to judge them before loading the full content. Date, project, record type, source and status are useful. A result marked as an abandoned plan should not look the same as an approved decision. A historical fact should keep the date that limits it. I prefer search that returns references and short excerpts. The agent can then open the relevant source rather than receiving a large bundle chosen by similarity alone. This adds a step, but it gives the task a chance to determine which history is worth reading. Access controls must follow the source. Moving text into a memory index should not make it visible to users or tools that could not read the original record. Deletion and retention rules need to reach the index as well as the primary store. The system can choose a memory tier from the request. Continuing a known task should load its current handoff. Applying a standing preference should query durable records for that user and scope. Investigating an old decision should search history and return source links.

This selection can be explicit in the agent's tools. Separate operations such as get_current_task, get_durable_records and search_history make the intent visible in the activity log. One general memory search is easier to wire up, but its results are harder to reason about. Ranking still matters inside each tier. Current status, source authority and scope may be more useful than semantic similarity. A current decision with modest text similarity can deserve priority over a long obsolete discussion that repeats the query terms. The agent should also be able to return no memory. Filling a context quota creates pressure to include marginal material. If the current task and source files are sufficient, an empty retrieval result is a valid outcome. Automatic capture is useful for preserving history, but durable memory should have a higher threshold. I promote information when it is expected to affect future work and has a source that can support it. For sensitive or consequential records, a person should confirm the wording and scope. A promotion process can ask:

  • Is this still true at the end of the task?
  • Which future work should use it?
  • What source supports it?
  • Who can correct or retire it?
  • Does it contain information that should remain restricted?

This review catches statements that sound permanent during a session but were only provisional. It also stops summaries from becoming the sole evidence when a better source exists. I evaluate retrieval with real task fixtures. For each fixture, I record which working context and durable records should appear, which historical items are tempting but unhelpful, and what source the agent should open if it needs more. The result is not just recall. I look at whether retrieved material changes the answer or action correctly. An irrelevant memory that the agent ignores still consumes context and review time. A stale record that changes a command is more serious. The next practical check is to take a task that has accumulated several sessions. Write one current handoff, identify the few durable decisions that still apply and move everything else behind search. Then run the task from a fresh session and inspect every memory item that enters before the first action.

Continue the thinking.

Comments are public and hosted in an open-source GitHub Discussions repository.

Loading comments connects your browser to GitHub. A GitHub account is required to post.

All blogs