Blog
Why I’m testing a graph for long-term AI memory
A long-running AI collaborator needs to recover relationships between people, projects, decisions and evidence. I am testing a graph-backed memory layer because those connections are difficult to preserve in a pile of transcripts, and I want the stored reasoning to remain inspectable.

Chat history is useful when the next question follows directly from the last one. It becomes awkward when a collaborator needs to recover why a project decision was made weeks later, which source supported it, and whether that decision has since changed. A transcript records statements in sequence. Search may return a durable fact beside a discarded idea, or a current decision beside the one it replaced. Passing all of that history back into a model can crowd the immediate task with stale material. I want the memory layer to store selected records with a type, identity, date, status and source link. The conversation remains available when somebody needs to inspect the original context. The graph test asks whether explicit relationships make those records easier to retrieve and inspect. The model still receives selected context for each task; the graph is one way of selecting and tracing it. The first design task is deciding what deserves a durable record. A small initial model could include people, organisations, projects, tasks, decisions, claims and sources. Each record needs a stable identifier and a limited set of properties that can be checked. Relationships carry the context. A person works on a project. A decision applies to a component. A claim is supported by a source. A later decision supersedes an earlier one. A task implements a decision and produces an artefact. Those links let a query follow the path relevant to the current work.
I will create nodes only for records that answer a known retrieval question or maintain current project state. Turning every noun in a conversation into a node would add noise and false precision. I would begin with questions such as: What decisions currently apply to this project? Which sources support this claim? What changed after this task? Who owns the unresolved item? The data model can grow when a real question cannot be answered cleanly. A memory record without provenance is hard to trust. If the system states that a project uses a particular rule, a reviewer should be able to open the source and see when that rule was recorded. Each extracted claim or decision should link to a source reference. That may be a document and passage, a task record, a code change, or a specific part of a conversation. Store the source date and the date the memory was created. They are not necessarily the same. The relationship should also describe the nature of support. A source may state a fact directly, provide evidence for an inference, contradict an earlier claim, or merely mention the subject. Treating every link as “related to” weakens the graph until it becomes another pile of associations.
Where a source is private, retrieval must preserve its permissions. A relationship to restricted evidence does not make the evidence available to every user. The memory service needs to filter both records and traversed sources for the identity making the request. Long-term memory will contain mistakes. A person may correct a name, a project may change direction, or an extracted claim may have misunderstood its source. Updating the text in place would hide how later work came to use a different fact. For important records, keep correction as an event. Mark the earlier claim as superseded or disputed, link it to the replacement, and record the source and time of the change. Retrieval can prefer the current record while a reviewer can still inspect the path. Not every property needs full historical treatment. The test should focus on fields where change affects later reasoning. Project status, ownership, requirements and decisions are stronger candidates than harmless formatting details. Corrections also need an interface. If a user sees a wrong memory in context, they should be able to identify the record and propose or apply a change according to their permission. A vague “thumbs down” may signal a problem but does not repair the stored fact.
A graph makes many connections available. It does not decide which connections belong in a prompt. Retrieval still needs a bounded question and a budget. For a project planning task, start from the project node and retrieve current decisions, open tasks, owners and recent supporting sources. For a question about a person, permission may limit the result to their role on the current project. For a claim check, follow support and contradiction links rather than collecting every record that mentions the same words. Combine structural queries with text search where useful. A graph query can select the relevant records and relationships. Text or vector search can find passages inside the linked source material. The final context should label record type, date, status and source so the model can separate current facts from historical notes. Log which memory records were retrieved for a task. When an answer is wrong, that makes it possible to distinguish a bad stored fact, a retrieval miss and a reasoning error. The first evaluation should use a small, manually checked set of records. Create questions whose answers require one or two relationships, such as finding the source behind a decision or identifying what superseded it. Check whether the retrieved context is correct, current and permitted.
Then introduce difficult cases: two people with similar names, one project sharing a component with another, a decision with conflicting sources, and a claim corrected after it was first stored. The expected retrieval path should be written down before adjusting the query. Operational checks matter as well. Can the graph be backed up and restored? Can records be deleted when policy requires it? Can a reviewer find every durable record created from a source? Can the system rebuild derived links without duplicating them? I will start with one project, its decisions, the sources supporting them and one correction. For each test question, I will write down the expected records and relationships first, then compare them with what retrieval returns before adding more history.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.