Article chapter 02 of 08
Inventory sources, owners and permissions
A document is not ready for retrieval simply because the application can read it. The team needs to know who owns it, which version applies, who may see it and how changes will reach the index.
Build a source register before bulk ingestion. For each collection, record its system of record, business owner, technical access method, document types, permission model, update pattern, retention rule and known quality issues. Include how quickly deleted or revoked material must disappear. Use the register when configuring ingestion and during later reviews.
Mixed formats need explicit handling. Word-processing files, PDFs, spreadsheets, web pages and scanned images do not produce equivalent text. A PDF may have a usable text layer, a broken reading order or no text at all. Tables can lose the relationship between headings and values. Headers and footers may repeat on every page and crowd retrieval results. The pipeline should identify the format and extraction method used for each item, then retain warnings when extraction is partial.
Resolve document versions at the source where possible. A shared drive may contain a published policy, a working draft and an old copy in an archive folder. Similar text can make all three look relevant to a question. Decide how the current version is identified, using source metadata or an explicit publication state rather than guesses based on filenames.
Carry permissions into the retrieval layer. If access is checked only when a file is ingested, a later permission change will not reach the index. Store the source identifier and access attributes required to evaluate the current user at query time, or maintain a synchronised permission index with a clear freshness rule. Once a restricted passage reaches the model, an instruction to hide it is not an access control.
Record document-level and section-level differences. One file may contain a public introduction and restricted appendices. If the source system cannot express that distinction, splitting the file during ingestion will not create legitimate permission boundaries. The access model needs to come from an authorised source.
Give every collection an owner who can answer two operational questions: should this material still be available, and which version should the system treat as current? A technical team cannot infer those answers safely from document similarity.
Before ingestion begins, sample each format and permission class. Confirm that the extracted text is readable, current status can be determined, and a test user receives only the material their role permits.