Blog
What I am testing in long-term AI memory
I am testing whether a graph can recover a decision through its relationships to a person, project and source. The immediate question is which records improve the next session and which simply make retrieval noisier.

I am testing whether a graph can recover a decision through its relationships to a person, project and source. The immediate question is which records improve the next session and which simply make retrieval noisier. Saving every conversation produces a large searchable archive quickly. The archive alone does not tell the agent which statement is current, whether it was a decision or a suggestion, who it applied to, or where it came from. The test starts with a narrow memory model and retrieval questions that a later session actually needs to answer. Most chat messages are coordination, intermediate reasoning or wording that was useful for a few minutes. They do not all need to become durable memory. Keeping them forever makes the search corpus larger without necessarily making the next answer better. The records I am interested in have a clear future use. A decision can prevent the same question being reopened. A constraint can stop a proposal that the project has already ruled out. A person's stated preference can shape how work is presented. A source reference can lead back to the evidence instead of preserving a loose paraphrase. Each stored record should therefore declare its type and scope. "Use Australian English" is a durable writing preference. "Try the second query" is probably session state. "The integration uses version 2 of this endpoint" may be project knowledge, but it also needs an effective date and source because the endpoint can change.
This classification can be wrong, so I am keeping the original source link and allowing records to be corrected or superseded rather than treating the first extraction as truth. A graph lets the memory store express relationships directly. A decision can be connected to the project it affects, the person who approved it, the source document that supports it and a later decision that replaces it. If a new session asks why a project uses a particular data boundary, retrieval can start at the project, follow active decisions of the relevant type and return the source. A text search might find every conversation containing "data boundary", including abandoned options and unrelated projects. I am keeping the initial relationship vocabulary small. Examples include APPLIES_TO, DECIDED_BY, SUPPORTED_BY and SUPERSEDES. Every edge needs a defined direction and meaning. Similar-looking relationships added casually will make queries harder to understand. Identifiers matter as well. A person's display name can change or collide with another name. A project may have several informal labels. Stable internal identifiers allow aliases to point to the same entity without merging two entities based on text similarity alone. Once the graph provides a bounded neighbourhood, semantic or keyword search can rank records inside that scope.
A retrieved sentence without provenance is difficult to trust. I want each durable record to point to the material from which it was created: a document version, approved note, issue or conversation turn. The memory record can store a concise statement for retrieval, but the source remains available for review. It also records when the statement was captured, who or what extracted it and whether a person confirmed it. Inferred records need to be labelled as inference. They should not quietly become an approved decision after being retrieved several times. Time needs explicit handling. created_at says when the memory record entered the system. effective_from says when the underlying decision began to apply. An optional effective_to or superseding edge tells the retriever when it stopped applying. Those dates answer different questions. Access control follows the source. A memory layer must not turn a restricted document into a broadly visible summary. Retrieval should check the requesting identity before returning the record or traversing to its source. A small evaluation set needs questions with known supporting records. Phrase them as a future session might ask them rather than as database queries. Examples include:
- What decision currently governs this project setting?
- Who approved it, and where is the supporting source?
- Has the earlier decision been replaced?
- Which stated preference applies to this output?
- Is there no durable memory for this question?
Include questions with no matching durable record. The correct result is no memory; returning the closest historical text would add weak context. For each question, record whether the correct memory appeared, its rank, whether stale or unrelated records appeared and how much context was returned. Inspect the relationship path as well. A correct answer reached through an accidental or overly broad path may fail as soon as more data is added. Run the same questions after adding distractor records from other projects and older decisions. This tests whether scoping and supersession continue to work as the graph becomes less tidy. Even relevant memory can be too much. The next session still needs room for its current instructions, source material and work product. Set limits by purpose. A writing task may need active voice preferences and one project brief. A technical decision may need the current constraint, the decision record and its source. Neither task needs the full history of every connected person and project. The retriever can rank candidates using relationship distance, record type, current status and source quality, then return a small bundle with identifiers and provenance. The agent can request a source or adjacent decision when needed. That is easier to inspect than injecting a long, flattened transcript before every prompt.
Summaries need care. A summary can reduce token use, but it can also merge an old view with a current one. I prefer summaries tied to a defined set of source records, with a version that changes when those records change. Long-term memory will contain mistakes. A person may correct a preference, a project may end or a source may be removed. The model needs a way to handle that without rewriting history silently. A correction can create a new record that supersedes the old one. Normal retrieval returns the active record, while an authorised audit can still see the sequence. A deletion request is different. It may require removing the record, derived summaries, embeddings and links to copied content according to the applicable retention rule. These operations belong in the test before more record types are added. Take one decision with a source, retrieve it, supersede it, and run the same question again. The result should show the new decision and retain a clear path to why it changed.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.