31 July 2024 / Applied AI / 8 chapters

Test retrieval and answers separately

From Building retrieval-augmented generation for real use

A wrong answer could come from missing ingestion, poor extraction, weak retrieval, unsupported generation or a source that's simply wrong. One end-to-end score won't tell you which stage needs fixing, so test retrieval and generation separately as well.

Build the evaluation set from representative question categories and source conditions: direct lookups, questions needing more than one passage, terminology differences, version-specific questions, permission boundaries, conflicting sources, missing information and questions the system should decline. Strip private values but keep whatever makes each case hard.

For each case, record the expected source passages (or an acceptable set of sources), the points the answer has to cover, claims it mustn't make and the expected decline behaviour. Plenty of questions have several valid wordings, so exact matching is usually too brittle, and a reviewer or structured rubric can judge whether the response is supported and complete.

Check retrieval first. Did the expected authorised source show up in the selected passages? Did an obsolete or restricted source outrank it? Did repeated chunks fill the result set? Look at results by category, because strong direct lookups can hide poor multi-source or no-answer behaviour in an overall score.

Then check generation against a fixed retrieved context where you can, so search changes don't muddy the result. Are the claims supported, do citations attach to the right passages, do qualifications survive summarisation and is conflict reported accurately?

Permission tests need explicit users and expected denials. Run the same question under roles with different access and check that restricted titles, excerpts and citation metadata don't turn up in the answer or any diagnostic view. Also test permission changes after ingestion and deleted documents after a sync.

Ask reviewers how long it takes to verify an answer and find the governing source. If checking means reading every retrieved document, the answer isn't saving anyone much time.

Store the configuration with each run (source snapshot, parser and chunk versions, embedding model, retrieval settings, generation model and prompt) so you can trace an improvement to a change and reproduce a regression. Keep failed cases as permanent fixtures when they represent a real category, but watch the balance. A set built only from recent failures can tune the system to odd wording while ordinary tasks get worse.

All articles