Article chapter 08 of 08
Testing memory against the next task
From Long-term AI memory: records, relationships and retrieval
The way I'd judge a memory system is by the next task. Did it get the right context, leave out stale material and keep the source of each claim? Record counts and retrieval hit rates won't tell you that.
Build an evaluation set from cases shaped like real tasks. Each case should say what the task is, who's doing it, when, what scope they're authorised for, which records are expected, which should be excluded and what the right response is when memory is missing or disputed. Put in corrections, expired facts, entities with similar names, conflicting sources and queries that should return nothing.
Test the write path on its own. Given a source passage, does the system propose the right type, subject, value, scope, provenance and effective time? Does it refuse unsupported claims, and does it keep hypothetical language from turning into current fact?
Then test retrieval and context assembly. Checks I'd include:
- current records outrank superseded versions
- authority wins over recency that isn't backed by anything
- access filters keep restricted candidates out of ranking
- ambiguous entities stay unresolved
- graph traversal stops at the permitted depth
- the context package fits its budget without dropping the record that decides the answer
- source references reopen the exact supporting material
- a correction changes the next retrieval
Look at how the agent uses memory too. A correct record can still be applied outside its scope, so the expected output should tell the difference between using a preference for one project and treating it as a permanent rule.
Before release, I'd run a correction drill. Put in a fact with a source, retrieve it during a task, supersede it through the supported interface and run the same task again. Then go through the stored records, relationship validity, indexes, caches and the retrieved context. The second run should use the correction, and you should still be able to get to both pieces of evidence from it.
Think about that delivery date again: set, corrected, then talked about as a "what if". Long-term AI memory works when the next session gets the current date from a small record that points at its source and has replaced the old one, rather than having to guess it from a saved transcript. If you're building one of these, I'd start the evaluation set with that kind of correction, because it shows straight away whether your records, relationships and retrieval keep track of what's still true.