31 July 2024 / Applied AI / 8 chapters

Evaluate retrieval and answers separately

From Building retrieval-augmented generation for real use

A wrong answer may come from missing ingestion, poor extraction, weak retrieval, unsupported generation or an incorrect source. An end-to-end quality score does not identify which stage needs repair, so test retrieval and generation separately as well.

Build an evaluation set from representative question categories and source conditions. Include direct lookups, questions requiring more than one passage, terminology differences, version-specific questions, permission boundaries, conflicting sources, missing information and questions the collection should decline. Remove private values while preserving the structure that makes each case difficult.

For each case, record the expected source passages or acceptable source set, required answer points, prohibited claims and expected decline behaviour. Some questions have several valid wordings, so exact answer matching is often too brittle. A reviewer or structured rubric can judge whether the response is supported and complete.

Evaluate retrieval first. Did the expected authorised source appear in the selected passages? Did an obsolete or restricted source outrank it? Did repeated chunks consume the result set? Inspect performance by category rather than relying on one aggregate score, because strong direct lookup can hide poor multi-source or no-answer behaviour.

Then evaluate generation using a fixed retrieved context where practical. Check whether claims are supported, citations attach to the right passages, qualifications survive summarisation and conflict is reported accurately. This isolates prompt and model behaviour from search changes.

Permission tests need explicit users and expected denials. Run the same question under roles with different access. Confirm that restricted titles, excerpts and citation metadata do not appear in the answer or diagnostic interface. Also test permission changes after ingestion and deleted documents after synchronisation.

Ask reviewers how long it takes to verify an answer, identify the governing source and recognise uncertainty. A response that requires reading every retrieved document may be accurate but provide little practical help.

Store the configuration with each evaluation run: source snapshot, parser and chunk versions, embedding model, retrieval settings, generation model and prompt. Without that record, a better result cannot be traced to a change and a regression cannot be reproduced.

Use failed cases as permanent fixtures when they represent a real category. The evaluation set should grow from observed errors, but it still needs balance. A collection made only from recent failures can overfit the system to unusual wording while ordinary tasks degrade.

All articles