31 July 2024 / Applied AI / 8 chapters

Chunk content around meaning and source location

From Building retrieval-augmented generation for real use

Chunking determines what the retriever can return to the model. Fixed character windows are easy to implement, but they can separate a heading from its rule, a table row from its columns or an exception from the instruction it qualifies.

Start with the structure available in the source. Headings, paragraphs, list items, table boundaries and page references can help produce units that remain understandable on their own. Keep the title and heading path with each chunk so a passage from "Cancellation" is not presented without the policy or product it belongs to.

Evaluate chunk size against the actual questions. A very small unit can match a term precisely while omitting conditions in nearby text. A large unit carries context but may dilute the matching signal and consume the model's context window with irrelevant material. Overlap can preserve boundary information, although too much overlap creates near-duplicate results that crowd out other sources.

Do not apply one rule blindly to every format. A short procedure step may need its prerequisites and warning. A long policy section may split at subheadings. A table might be indexed by logical rows with column names repeated. A question-and-answer page may work as one pair per chunk. Store the chunking strategy and version so results can be reproduced during evaluation.

Preserve source location in a form the user can follow. Page number is useful for stable PDFs. A heading path and anchor may be better for a web page. Spreadsheet references need sheet and range. Do not generate a polished citation label that cannot lead back to the evidence.

Handle repeated and boilerplate text. Navigation, confidentiality notices, standard footers and copied introductions can dominate similarity search. Remove known layout noise during normalisation, but retain content that changes the meaning or status of the document. De-duplication should identify exact and near-exact copies for review rather than automatically deciding which version is authoritative.

Test chunks directly before tuning retrieval. Read a sample without the surrounding document and ask whether the passage identifies its subject, keeps relevant conditions and can support a citation. Then inspect the chunks created around headings, page breaks, lists and tables. Many apparent model failures begin with source text that was split or extracted badly.

All articles