Article chapter 01 of 08
Work out the job before you build the index
Retrieval-augmented generation (RAG) hands a language model selected source material along with the user's question, so it answers from that material instead of only what it learned in training. A folder of tidy documents, a vector index and a prompt can answer a handful of prepared questions, which is why the demos come together so quickly.
Getting it into real use needs a much narrower idea of what "working" means. I'd start with who's going to use it, what they're trying to get done and what they'll do with the answer. A support assistant that helps staff find a procedure, a public answer service and a tool that drafts regulated advice could all use the same retrieval technique, but the acceptable sources, response time, review and what happens when it gets something wrong will be different for each.
Collect the kinds of questions people actually ask, without copying private content into a test set nobody's protecting. Think finding a stated rule, combining facts from two sections, finding the current form, explaining a process, comparing versions and recognising that the material doesn't contain an answer. For each one, write down what evidence a good response has to show, so someone checking it can find the source and the spot in it.
Then write the boundaries down in plain language:
- Which source collections are included.
- Which users or roles can query each collection.
- What kinds of answer the system is allowed to produce.
- Whether the answer is information, a draft or something that triggers an action.
- Which questions need to go to a person or another system.
- What the interface does when the evidence is weak or missing.
That stops "chat with company knowledge" from becoming the requirement. Account access, public answers and internal drafting might share retrieval components, but they carry different risks.
For a first slice, pick a collection where you know who owns the sources and can check the answers. One maintained procedure library is much easier to look after than a crawl across shared drives, inboxes and collaboration tools. Before choosing a model, I'd also write an answer contract: what the user can expect, which evidence they'll see and when the system will say it can't answer. That then guides the ingestion, retrieval and interface decisions.