Article chapter 05 of 08
Retrieve only what this user is allowed to see
I'd treat retrieval as an application request like any other: an authenticated user, an authorised collection, a query representation, ranking rules and a limit on what goes into the model's context.
Filter for access before any passage gets near the model, so the retrieval service only returns chunks the current user can read. Filtering a mixed result set afterwards also gives poor answers, because restricted passages take the top spots and leave too little permitted evidence. If your search technology supports it, put the permission attributes in the query filter.
Semantic vector search helps when the question and the source use different wording, and keyword search is still better for exact identifiers, product names, policy numbers and unusual terms. A combined approach can pull candidates from both and rank them with consistent rules. The right balance depends on your questions and sources, so it's a configuration you evaluate.
Keep query preparation restrained. Spelling normalisation, acronym expansion and known terminology mappings help, but automatic query rewriting can change what the user meant. Keep the original query, log the rewritten one, and test cases where date, region or account type change the answer.
Ranking can use metadata as well as text relevance, so current published material can rank above drafts and a source for the user's region can outrank a general copy. Recency only helps when newer actually means applicable, though. An older contract or policy might still govern a historical question, so date filtering should follow what the user is trying to do.
When repeated chunks from one document push out other evidence, set diversity rules. The model might need a procedure and its exception from different sections. Retrieve enough candidates to rank well, pass a smaller set you can inspect into generation, and log the selected source identifiers, versions and scores.
Decide what happens when retrieval is weak. A low score doesn't map neatly to "no answer" across every collection, so calibrate thresholds and evidence checks against your evaluation cases. If the passages don't support what's being asked, return what you did find or say the sources don't answer it. Searching more widely should be a deliberate product decision, especially when the wider collections have different permissions.