Blog
Trying to remember everything made the agent worse
A total-recall experiment has made the problem clear: more retrieved history can crowd out the task in front of the agent. I am separating identity, durable decisions, project state and searchable archives so memory is selected by purpose instead of poured into every session.

The test stored more conversations and working notes, retrieved anything that looked semantically related, then added those results to the next prompt. Semantic similarity was treated as enough reason to include a record. In practice, the retrieved material often occupied the space needed to understand the current job. Old plans resurfaced after the plan had changed. A phrase shared by two unrelated projects could pull in the wrong context. Long transcripts brought back tentative ideas with the same weight as settled decisions. The model then had to infer which fragments were current, authoritative and relevant before it could start the actual task. A larger context window can hold more irrelevant material, so it does not solve poor selection. Every memory included in a session competes for attention and may change the answer, even if it was added with the harmless intention of being comprehensive. I am now treating memory as several types of record rather than one expanding transcript. The categories are practical, and their boundaries can change as the system develops. Identity covers the small amount of information that should remain stable across most sessions: who the agent is working with, broad preferences and standing constraints. These records should be concise and deliberately edited. A conversational guess should not quietly become part of identity.
Durable decisions record choices that are expected to affect later work. A useful decision record includes what was decided, when, within which scope and where the supporting source can be found. If the decision is superseded, the old record remains traceable but should no longer be presented as current guidance. Project state describes what is true now: current branch, open problem, active milestone, environment or next action. It changes often and needs an owner or source. Searchable archives hold transcripts, logs and documents that might matter later but do not deserve automatic inclusion. Separate record types make a chat remark less likely to travel through the system as a permanent instruction. A memory query needs more than semantic similarity. I want retrieval to consider the task, project, record type, status and time. A request to review code may need current project state and relevant decisions. A request to draft a personal note may need a small identity profile. Neither task needs a broad sample of every previous conversation. The retrieval sequence I am testing is:
- Identify the task and its project or subject.
- Load the minimum standing identity and constraints required for that task.
- Fetch current project state from explicit records.
- Retrieve a small number of related decisions, preferring current records.
- Search the archive only when the task or an unresolved reference requires it.
This makes archive search a deliberate operation. If a decision record points to a meeting note, the agent can retrieve that source when it needs detail. It does not need the whole meeting note in every session. Semantic search remains useful, but it should operate inside filters. A close vector match from an unrelated project is still unrelated. Metadata such as project identifier, record class and validity status provides boundaries that similarity alone cannot infer reliably. A remembered statement without a source is difficult to challenge. Was it copied from a specification, inferred from a conversation or written by the agent as a summary? Those origins affect how much authority it should have. For durable records, I am keeping a link to the source material and enough location information to inspect it. The source may be a file and section, an issue, a database record or a dated message. The memory can stay short because the detail remains available through that link. Corrections need to be explicit as well. Editing a memory in place can erase the reason an earlier session behaved differently. Record that one statement supersedes another, then make retrieval prefer the current record. This leaves a path for investigating unexpected behaviour without feeding obsolete text into normal work.
The same applies to uncertainty. A provisional interpretation should be labelled as such. It should expire, request confirmation or remain outside durable memory instead of hardening through repetition. I am setting a retrieval budget by both count and space. The exact numbers depend on the model and task, but the principle is to make inclusion scarce enough that the system must rank records. A retrieval result should explain why each item was selected. Useful fields include record type, project, status, source, retrieval score and token estimate. If the set exceeds the budget, the system can remove archive excerpts first, collapse related records or fetch a concise current summary backed by links. This also gives the agent a chance to notice poor retrieval. If half the budget comes from one old transcript, that is visible in the trace. Without the trace, a weak answer may look like a model failure when the prompt was already crowded with irrelevant history. Summaries can reduce space, but they introduce another layer that may become stale. I want summaries to state their coverage date and source set. A fresh source should take precedence when it conflicts with an older summary.
Evaluate whether a selected memory helps the agent complete a task without introducing an incorrect assumption. Recall on its own is a weak measure. Retrieving every related passage can score well while making the final work worse. I am building small fixtures from tasks where the needed context is known. Each fixture can specify records that should appear, records that must stay out and facts that require source lookup. Then the answer can be checked for task success, unsupported carry-over and use of superseded information. The next cleanup step is to reduce the default identity context, moving changing facts into project-state records and keeping transcripts searchable rather than automatic. For each memory that still loads by default, I want a plain answer to two questions: which tasks need this, and how will it be corrected when it changes?
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.