ARR/Field Notes/Corporate
Field notes · Building AI apps

Chunking breaks meaning.

Your RAG pipeline found the right document, cited it correctly, and still answered half the question. The model isn't the problem. The 512-token knife you cut your documents with is.

For engineers building RAG & AI apps · 7 min read
A project note split into four fixed-size chunks. The cut falls mid-sentence and even mid-word. Only chunks one and three are retrieved, so the cause of the delay and the decision never reach the model, and it answers with half the story.
Click to open full size
The story

Right document. Half the answer.

You've built a retrieval pipeline over your team's notes. Someone asks: "Why was the Bangkok launch delayed, and what did we decide?" The answer is all in one short paragraph of project_notes.md. Your splitter cut it every 512 tokens:

#1 On 3 March the team agreed to move the Bangkok launch from April to June. The delay was caused ────────────────────────────────────────── cut #2 by the supplier missing the certification deadline for the dosing pumps, which meant the units could not legally be installed. ────────────────────────────────────────── cut #3 To protect the launch, Nok will run a private preview for the f ────────────────────────────────────────── cut #4 irst ten waitlist clients in May, using two pre-certified demo units.

Chunks #1 and #3 share the most words with the question, so they win the similarity search. The cause lives in #2. The decision lives in #4. Neither is retrieved.

Model answer
"The launch moved to June because of a certification issue."
grounded · cited · half the answer

Every word of that answer is supported by the retrieved context. Your groundedness check passes. And the user still doesn't know whose certification failed or what the team decided.

Why it happens

Boundaries set by a counter, not by meaning.

Fixed-size chunking is blind to structure. It splits wherever the token count lands: mid-sentence, between a heading and its content, between a claim and its caveat, occasionally mid-word. A chunk torn out of its document also loses the context that made it findable, such as the pronoun's referent, the acronym defined three paragraphs up, or the section title that scoped everything below it.

Then retrieval scores each fragment on its own. A fragment that echoes the question's wording beats the fragment that actually answers it, so the model is handed the half of the idea that sounds relevant and none of the half that is.

01 · index

Split by length

Every document is cut every N tokens, wherever the count lands.

02 · retrieve · where it breaks

Fragments compete alone

Top-k returns the pieces that look like the question. The cause and the decision sit in the pieces that don't.

03 · generate

A faithful half-answer

The model summarises exactly what it was given. Grounded, cited, incomplete.

Five signs, five fixes

Your evals pass. Your users still get half answers.

01
Stitched fragments

The answer spans a chunk boundary

The model gets one side of the cut and fills the gap, or confidently answers with the half it has.

✓ With ARR

Boundaries are placed where one idea ends and the next begins, so each piece returned is a complete thought, whatever its length.

02
Overlap echo

Top-k full of near-duplicates

Sliding-window overlap and copied documents fill your top-k with the same text, crowding out the passage you actually needed.

✓ With ARR

Each distinct passage comes back once, so your k slots go to different ideas, not repeats.

03
Top-k whack-a-mole

Raise k, pay for it

To recover lost context you raise k or add more calls. Cost and latency climb, and the extra chunks bring noise with them.

✓ With ARR

Because every piece is a complete idea, fewer calls are needed to get the full picture.

04
The stale index

The doc changed; the vectors didn't

A note was updated after indexing. Until the next re-embed, your pipeline serves the old text as current.

✓ With ARR

ARR tells the AI when something it read has changed since, instead of serving the old version as current.

05
Drifting citations

Chunk IDs that point at the wrong text

Re-chunk or edit a document and your stored offsets now cite different text than the model actually saw.

✓ With ARR

Every reference carries a fingerprint of the exact text it points to, so anyone downstream can confirm it's unchanged (in shared setups).

Why better chunking isn't enough

There's no right chunk size. Only trade-offs.

The standard advice is to tune: smaller chunks for precision, larger chunks for context, overlap to catch sentences that straddle a boundary, semantic or late chunking, contextual headers, multi-scale indexes. All of them help, and all of them trade one failure for another. Small chunks keep detail and lose context; large chunks keep context and blur detail. Even the right size turns out to depend on the query (AI21; see also Weaviate and Redis on the trade-offs).

ARR takes a different route: it moves reasoning into the retrieval step, drawing each boundary around a complete idea and linking related ideas across documents, so the model receives connected context instead of fragments. It plugs in under the AI you already use; your model doesn't change.

A fair caveat

We've measured ARR in our own benchmark on code and text, against the standard tools agents use, including embedding search. We haven't yet published a head-to-head against specific chunking strategies such as late chunking or contextual retrieval. Early access is how we close that gap, on your data.

Try this tomorrow

The split-answer test

  1. Find a document where the answer to a real question spans two paragraphs: a cause and a decision, or a rule and its exception.
  2. Ask your pipeline that question and log the retrieved chunks.
  3. Check whether both halves arrived, and whether any chunk starts or ends mid-sentence.

If one half is missing, your model didn't fail. Your boundaries did. That's the part ARR replaces.

Early access

Plug ARR into your pipeline before anyone else.

We're opening ARR to a small group of builders first. Tell us about your stack and we'll be in touch when your place is ready. No payment required.

More in ARUKAS Field Notes:
Your AI is answering from an old document
When legal AI gets the clause wrong
Why AI coding agents break things three files away
Running parallel coding agents without them overwriting each other
Which version of the policy did your AI just apply?
AI citation errors start before the AI writes anything
Duplicate files, duplicate totals: where AI reconciliation goes wrong
What AI prior-art search doesn't tell you it missed
How ARR works →