azyware
Technology

Chunking strategies for RAG: document type decides the window

EZ
Eazyware
· 7 min read
Quick answer

Which RAG chunking strategies work best for different document types?

Chunk by structure, sections, tables, clauses, rather than fixed windows; contracts, tickets and manuals each need different strategies. The right chunk is the smallest unit that still answers a question on its own, and it carries metadata so filters and citations work. Here are the strategies per document type.

RAG chunking strategies decide whether the passage that answers a question can be retrieved at all. A fixed 500-token window cuts a contract clause in half, separates a table from its header, and glues three unrelated Slack messages together. The model then receives fragments that no longer mean what the document meant, and no amount of prompt work recovers that. The rule we apply: chunk by the document's own structure, choose the window per document type, and attach metadata to every chunk. This article sets out how that plays out for contracts, manuals, tickets, spreadsheets, wikis and chat, and how to know when you have got it wrong.

Chunking is one stage of the retrieval pipeline we build under the Retrieval & Knowledge Engineering service, and it is the stage most often responsible for a category of questions failing on the golden set.

Why chunk size in RAG is the wrong question

Teams ask "what chunk size should we use?" as if there were one number. There is not, because the right unit is defined by the content, not the token count. A clause in a services agreement is a unit: it has a number, a heading and a meaning that survives on its own. A step in a repair manual is a unit, but only with its parent procedure's title attached. A ticket comment is not a unit; the ticket's problem statement plus its resolution is. The question to ask is: what is the smallest piece of this document that still answers a question without the surrounding text? That is the chunk. Token count follows.

Two consequences. First, chunks in one corpus will vary in length, and that is fine; embedding models handle a range. Second, a chunk must carry enough metadata (document, section path, date, owner, access control) that retrieval can filter on it and a citation can point to it. A chunk without metadata is a passage you cannot find again.

Chunking strategy by document type

Document typeUnit to chunk onAttach as metadataCommon failure with fixed windows
Contracts and policiesNumbered clause or sub-clauseClause number, section heading, defined terms, effective date, versionClause split mid-sentence; definitions separated from the clause that uses them
Manuals and SOPsSection or procedure step groupHeading path (chapter > section), product model, revisionStep 4 retrieved without the procedure it belongs to
Support ticketsProblem statement, resolution, internal notes as separate chunksTicket ID, product, status, customer tier, dateChit-chat comments merged with the fix; PII embedded
Spreadsheets and tablesRow group with header repeatedSheet name, column headers, source fileHeader lost; numbers with no column meaning
Wiki and Confluence pagesSection under an H2 or H3Page title, breadcrumb path, last edited, spaceStale and current sections mixed; navigation text indexed
Slack and Teams threadsThread with parent messageChannel, participants, date, permalinkReplies without the question; three threads in one chunk
Scanned PDFs and formsLayout region (block, table, field group)Page number, OCR confidence, form typeColumns interleaved; tables flattened to prose

Structural chunking: parse first, then split

Structural chunking needs a parser that preserves structure. Plain text extraction throws away the headings, table boundaries and clause numbers that the chunker needs. Layout-aware parsing keeps them: headings become a path, tables stay tables, footnotes attach to the right paragraph. Open tools such as Docling handle much of this for PDFs and office documents, and the parsing stage should be tested per document type with a handful of real files before any chunking rule is written.

Once structure is available, the splitter walks it: for a contract, emit one chunk per clause, with sub-clauses either merged into the parent or emitted separately depending on length; for a manual, emit one chunk per section but never split a numbered procedure; for a wiki page, emit one chunk per heading block. Where a unit is too long for the embedding model, split it at paragraph boundaries and give each piece the same metadata plus an index so the pieces can be reassembled at answer time.

Overlap, parent context and small-to-big

Overlap between fixed windows is a patch for a problem structural chunking mostly removes, but two related techniques remain useful. Parent context prepends the heading path ("Warranty > Exclusions > Consumables") to the chunk text so the embedding carries the context a reader would have. Small-to-big retrieval indexes small chunks for precise matching but returns the enclosing section to the model so the answer has room to be accurate. Both are cheap and both improve groundedness on long documents.

Semantic chunking: when structure is missing

Semantic chunking splits text where the meaning shifts, usually by measuring the similarity between adjacent sentences and cutting at the low points. It is the right tool when a document has no usable structure: transcripts, long emails, free-text notes, exported chat without thread markers. It is the wrong tool for a contract, because a contract already tells you where the units are and a similarity threshold will disagree with the drafter. Use semantic chunking as a fallback for unstructured sources, not as a default.

Metadata is not optional

Every chunk carries at minimum: source document ID, section path, version or last-modified date, and the access-control list of the source. Add domain fields where they change retrieval: product model for manuals, customer tier for tickets, jurisdiction for contracts. This metadata does three jobs. It lets the retriever filter before ranking, which is how permission-aware retrieval is enforced. It lets citations point to a specific clause or page rather than a file. And it lets the refresh pipeline find and replace the chunks of a changed document without re-indexing everything.

How to know your chunking is wrong

Build a golden set of one to three hundred real questions with the passage that answers each, grouped by document type. Run retrieval and measure recall per group. A group with low recall while others are fine is almost always a chunking problem: the passage exists but was split so that no chunk matches the question. Inspect the failures by eye; the pattern is usually obvious within ten examples. Fix the rule for that type, re-index that type only, and re-run. The method is described in how to measure RAG quality, and it is the only reliable way to tune chunking, because the effect of a change is invisible without a measurement.

A worked example

A facilities-management company had a RAG assistant over equipment manuals and their maintenance ticket history. Technicians complained it answered procedure questions with the wrong step and could not find past fixes for a known fault code. The corpus had been chunked with a single fixed window and overlap.

The review grouped the golden set by type. Manual questions failed because procedures were split across chunks and the heading was lost, so a retrieved step did not say which procedure it belonged to. Ticket questions failed because the fault code appeared in a comment that had been merged with two unrelated comments, and because vector search alone did not match the code exactly. The fix was structural: manuals were re-parsed with headings preserved and chunked by procedure with the heading path prepended; tickets were split into problem, resolution and notes with the fault code as metadata; keyword search was added alongside vectors for the codes. Recall for both groups rose on the golden set, and the technicians' complaints stopped. The same discipline underpins the KYC document-intelligence work, where a form field split from its label is a compliance problem rather than an inconvenience.

Team and timeline

Chunking is decided in the first two weeks of a retrieval build, alongside the corpus audit and the golden set, and revisited whenever the eval shows a document type failing. It is typically one retrieval-focused AI engineer and a data engineer for the parsing pipelines. Across a full build of four to ten weeks, chunking and parsing take roughly a fortnight; the rest is retrieval tuning, permissions, refresh and the answering layer. The Retrieval & Knowledge Engineering service starts at $14,000 / ₹8.8L, and a three-week ProofRun at $6,250–10,500 will tell you which of your document types need which strategy before you commit to a full build. Current figures are on the pricing page; the LLM application practices for evals apply throughout.

Before you start: a checklist

  • Inventory document types and pick three real files of each
  • Test the parser on those files; check headings, tables and clause numbers survive
  • Define the unit per type: clause, procedure, thread, row group, section
  • Decide the metadata fields per type, including access control and version
  • Collect 100–300 real questions grouped by document type
  • Use semantic chunking only where structure is genuinely absent
  • Prepend heading paths and consider small-to-big for long documents
  • Measure recall per document type before and after every chunking change

Glossary

  • Structural chunking: splitting on the document's own units (clauses, sections, threads) rather than a token count.
  • Semantic chunking: splitting where adjacent sentences stop being similar; a fallback for unstructured text.
  • Parent context: prepending the heading path to a chunk so its embedding carries the surrounding meaning.
  • Small-to-big: indexing small chunks for matching but returning the enclosing section to the model.
  • Recall: the share of golden-set questions for which the correct passage appears in the retrieved set.

See why basic RAG fails in production for the six failures chunking sits among, hybrid search for the exact-match problem chunking cannot solve alone, and document intelligence for the parsing side of scanned files.

Let the document tell you where its units are, attach the metadata that lets you find them again, and measure recall per type so the next change is a fix rather than a guess.

Frequently asked questions

What is the best chunk size for RAG?

▾

There is no single number. The right chunk is the smallest unit of the document that still answers a question on its own: a clause, a procedure, a thread. Length varies across a corpus and embedding models handle that; measure recall per document type instead of tuning a token count.

Is semantic chunking better than fixed-size chunking?

▾

For unstructured text such as transcripts or notes, yes. For documents with their own structure, contracts, manuals, wiki pages, structural chunking beats both, because the drafter has already marked the units and a similarity threshold will disagree with them.

Do chunks need metadata?

▾

Yes. Source ID, section path, version and access control at minimum. Metadata is how the retriever filters before ranking, how citations point to a clause rather than a file, and how refresh replaces only the chunks of a changed document. Our retrieval builds treat it as mandatory.