31 July 2024 / Applied AI / 8 chapters

Make ingestion something you can rerun

From Building retrieval-augmented generation for real use

Documents change, parsers improve, chunking rules get revised and embedding models get swapped out, so ingestion has to be rerunnable. A one-off script that loads a clean folder leaves you no way to update or correct anything.

Give each source item a stable identifier from the source system instead of relying on its filename. Store the source version, retrieval time, content hash, parser version, access metadata and processing status. With those fields the pipeline can skip what's unchanged, process revisions and remove content that no longer exists.

Break the processing into stages so failures stay visible:

  1. Discover the source items and record the expected set.
  2. Fetch each item with its metadata and permissions.
  3. Extract structured text and keep the extraction result.
  4. Normalise repeated noise without wiping out layout that means something.
  5. Split the content into retrievable units.
  6. Create embeddings and index entries.
  7. Reconcile the finished index against the expected source set.

Each stage should be restartable for an individual item. If one big file fails, mark it and carry on according to a policy you've written down, because a silent skip leaves search looking complete when part of the collection is missing.

Each chunk should carry the stable document identifier, version, section or page location, title, owner and access attributes, with the original source references kept outside the chunks. That's what lets you do citations, filtering, deletion and investigation, and reprocess without treating the vector database as the source of truth.

Extraction needs quality checks too. Quarantine empty output, suspiciously short output from a large file, repeated page furniture and obvious encoding failures. Optical character recognition (OCR) might suit scanned material, but mark that text so evaluation can account for recognition errors. Important tables might need a format-specific representation instead of being flattened into lines.

The update path has to delete as well. When a source item is replaced, retire the old chunks in the same controlled operation, and when access is revoked, make sure stale entries can't be returned. A nightly refresh might be fine for one collection while permission changes need a faster event-driven path.

Finish every run with a reconciliation of expected, processed, unchanged, failed and removed documents, and show failures to an operator with the source identifier and stage. An index can accept new records without complaint and still be wrong if nobody checks the old or missing ones.

All articles