31 July 2024 / Applied AI / 8 chapters

Keep it fresh once people are using it

From Building retrieval-augmented generation for real use

Once the service is live, documents change, permissions move, connectors fail and users find gaps, and operators need to see each of those.

Monitor ingestion by collection and stage, showing the last successful discovery, fetch and index times plus outstanding failures. The application health check can be green while the index is stale, so data freshness needs its own status. If a published procedure should be searchable within a set period, measure which source version is actually in the index, and do the same for revocation and deletion.

Keep a feedback route that captures the question, the answer, the authorised source references, the configuration version and what the user's problem was, without asking them to paste private content into some unrelated ticket. Categories could include wrong source, outdated source, missing source, unsupported statement, unclear answer and access problem. Source problems go to content owners and system problems to the delivery team.

Plan for partial failure. If the embedding service is down, new documents can wait while existing search keeps working. If the permission service is down, the safe response might be to deny retrieval instead of using cached access past its approved window. If citation links fail, the interface shouldn't pretend the answer is still fully verifiable.

Collections also pick up abandoned drafts, duplicates and documents nobody owns. Remove or quarantine that material through an approved process and reconcile the index, because more documents can make answers worse when nobody can say which ones apply.

Before release, I'd check that:

  • Each collection has a business and technical owner.
  • Current-version and deletion rules are defined.
  • Extraction quality has been sampled for each format.
  • Access is enforced during retrieval.
  • Chunks keep their source and version metadata.
  • Ingestion can restart and reconcile.
  • Evaluation includes unsupported, conflicting and restricted questions.
  • Citations open the exact permitted source location.
  • Operational views show stale and failed items.
  • Feedback can be traced back to the configuration and evidence used.

Run that check again whenever you add a source collection, change the embedding or generation model, change chunking or open it up to a wider audience. Most of what gets RAG from a demo on a tidy folder into real use is the work around the model: keeping each user to the current sources they're allowed to see, and being able to trace every answer back to them. So for a first implementation, I'd pick one maintained collection, write its answer and access contract, and build the evaluation set before indexing anything else.

All articles