30 June 2025 / Applied AI / 8 chapters

Evaluate memory through the next task

From Long-term AI memory: records, relationships and retrieval

Evaluate memory against the next task: did it receive the right context, exclude stale material and preserve the source of each claim? Record counts and retrieval hits do not answer those questions.

Build an evaluation set from task-shaped cases. Each case should state the task, actor, time, authorised scope, expected records, excluded records and correct response when memory is absent or disputed. Include corrections, expired facts, similar entity names, conflicting sources and queries that should return nothing.

Evaluate the write path separately. Given a source passage, check whether the system proposes the correct type, subject, value, scope, provenance and effective time. Confirm that it refuses unsupported claims and does not turn hypothetical language into current fact.

Then test retrieval and context assembly. Useful checks include:

  • current records outrank superseded versions;
  • authority beats unsupported recency;
  • access filters prevent restricted candidates from entering ranking;
  • ambiguous entities stay unresolved;
  • graph traversal stops at the permitted depth;
  • the context package fits its budget without dropping the decisive record;
  • source references reopen the exact supporting material;
  • a correction changes the next retrieval.

Review the agent's use of memory as well. A correct record can still be applied outside its scope. The expected task output should distinguish using a preference for one project from treating it as a permanent rule.

Before release, run a correction drill. Insert a fact with a source, retrieve it in a task, supersede it through the supported interface, and repeat the same task. Inspect the stored records, relationship validity, indexes, caches and retrieved context. The new task should use the correction and retain a clear path to both pieces of evidence.

All articles