Article chapter 03 of 08
Build ingestion as a repeatable data pipeline
Ingestion should be rerunnable. Documents change, parsers improve, chunking rules evolve and embedding models may be replaced. A one-off script that loads a clean folder leaves no reliable path for updates or correction.
Assign each source item a stable identifier derived from the source system rather than its filename alone. Store the source version, retrieval time, content hash, parser version, access metadata and processing status. These fields let the pipeline distinguish unchanged material, process revisions and remove content that no longer exists.
Use staged processing so failures remain visible:
- Discover source items and record the expected set.
- Fetch each item with its metadata and permissions.
- Extract structured text and retain the extraction result.
- Normalise repeated noise without erasing meaningful layout.
- Split the content into retrievable units.
- Create embeddings and index entries.
- Reconcile the completed index against the expected source set.
Make each stage restartable for an individual item. If one large file fails, mark it and continue according to a stated policy. A silent skip leaves the search interface looking complete even though part of the collection is missing.
Keep original or approved source references outside the generated chunks. A chunk should carry the stable document identifier, version, section or page location, title, owner and access attributes. This metadata supports citations, filtering, deletion and investigation. It also makes reprocessing possible without treating the vector database as the source of truth.
Extraction needs quality checks. Reject or quarantine empty output, implausibly short output from a large file, repeated page furniture and obvious encoding failures. For scanned material, optical character recognition may be appropriate, but the resulting text should be marked so evaluation can include recognition errors. Important tables may need a format-specific representation rather than flattened lines.
An update path also needs to delete. When a source item is replaced, remove or retire the earlier chunks in the same controlled operation. When access is revoked, make sure stale index entries cannot be returned. A nightly refresh may suit one collection, while policy updates or permission changes may require a faster event-driven path.
Finish each run with reconciliation: expected documents, processed documents, unchanged documents, failed documents and removed documents. Surface failures to an operator with the source identifier and stage. An index can accept new records successfully while remaining wrong because old or missing records were never checked.