31 July 2024 / Applied AI / 8 chapters

Know who owns each source and who can see it

From Building retrieval-augmented generation for real use

Before a document goes near the index, you need to know who owns it, which version applies, who can see it and how changes will get through later. So before any bulk ingestion, build a source register. For each collection, record the system of record, business owner, technical access method, document types, permission model, update pattern, retention rule and known quality issues, plus how quickly deleted or revoked material has to disappear.

Mixed formats need their own handling. Word documents, PDFs, spreadsheets, web pages and scanned images don't give you equivalent text. A PDF might have a usable text layer, a broken reading order or no text at all, tables can lose the link between headings and values, and repeated headers and footers crowd the retrieval results. The pipeline should note the format and extraction method for each item and keep a warning when extraction was only partial.

Sort out versions at the source where you can. A shared drive might hold a published policy, a working draft and an old copy in an archive folder, and all three can look relevant to a question. Identify the current version from source metadata or an explicit publication state instead of guessing from filenames.

Permissions have to carry through into the retrieval layer. If access is only checked at ingestion, a later permission change never reaches the index. Store the source identifier and the access attributes you need to evaluate the current user at query time, or keep a synchronised permission index with a clear freshness rule. Once a restricted passage is in the model's context, a prompt telling the model to hide it won't protect it, so the filtering has to happen before then.

One file might have a public introduction and restricted appendices. If the source system can't express that, splitting the file during ingestion won't create real permission boundaries, because the access model has to come from an authorised source.

Every collection also needs an owner who can say whether the material should still be available and which version is current, since a technical team can't work that out from document similarity. Before ingestion starts, sample each format and permission class and check the text is readable, current status can be worked out, and a test user only gets what their role allows.

All articles