You've built a retrieval pipeline over your team's notes. Someone asks: "Why was the Bangkok launch delayed, and what did we decide?" The answer is all in one short paragraph of project_notes.md. Your splitter cut it every 512 tokens:
Chunks #1 and #3 share the most words with the question, so they win the similarity search. The cause lives in #2. The decision lives in #4. Neither is retrieved.
"The launch moved to June because of a certification issue."
Every word of that answer is supported by the retrieved context. Your groundedness check passes. And the user still doesn't know whose certification failed or what the team decided.
Fixed-size chunking is blind to structure. It splits wherever the token count lands: mid-sentence, between a heading and its content, between a claim and its caveat, occasionally mid-word. A chunk torn out of its document also loses the context that made it findable, such as the pronoun's referent, the acronym defined three paragraphs up, or the section title that scoped everything below it.
Then retrieval scores each fragment on its own. A fragment that echoes the question's wording beats the fragment that actually answers it, so the model is handed the half of the idea that sounds relevant and none of the half that is.
Every document is cut every N tokens, wherever the count lands.
Top-k returns the pieces that look like the question. The cause and the decision sit in the pieces that don't.
The model summarises exactly what it was given. Grounded, cited, incomplete.
The model gets one side of the cut and fills the gap, or confidently answers with the half it has.
Boundaries are placed where one idea ends and the next begins, so each piece returned is a complete thought, whatever its length.
Sliding-window overlap and copied documents fill your top-k with the same text, crowding out the passage you actually needed.
Each distinct passage comes back once, so your k slots go to different ideas, not repeats.
To recover lost context you raise k or add more calls. Cost and latency climb, and the extra chunks bring noise with them.
Because every piece is a complete idea, fewer calls are needed to get the full picture.
A note was updated after indexing. Until the next re-embed, your pipeline serves the old text as current.
ARR tells the AI when something it read has changed since, instead of serving the old version as current.
Re-chunk or edit a document and your stored offsets now cite different text than the model actually saw.
Every reference carries a fingerprint of the exact text it points to, so anyone downstream can confirm it's unchanged (in shared setups).
The standard advice is to tune: smaller chunks for precision, larger chunks for context, overlap to catch sentences that straddle a boundary, semantic or late chunking, contextual headers, multi-scale indexes. All of them help, and all of them trade one failure for another. Small chunks keep detail and lose context; large chunks keep context and blur detail. Even the right size turns out to depend on the query (AI21; see also Weaviate and Redis on the trade-offs).
ARR takes a different route: it moves reasoning into the retrieval step, drawing each boundary around a complete idea and linking related ideas across documents, so the model receives connected context instead of fragments. It plugs in under the AI you already use; your model doesn't change.
We've measured ARR in our own benchmark on code and text, against the standard tools agents use, including embedding search. We haven't yet published a head-to-head against specific chunking strategies such as late chunking or contextual retrieval. Early access is how we close that gap, on your data.
If one half is missing, your model didn't fail. Your boundaries did. That's the part ARR replaces.
We're opening ARR to a small group of builders first. Tell us about your stack and we'll be in touch when your place is ready. No payment required.
More in ARUKAS Field Notes:
Your AI is answering from an old document
When legal AI gets the clause wrong
Why AI coding agents break things three files away
Running parallel coding agents without them overwriting each other
Which version of the policy did your AI just apply?
AI citation errors start before the AI writes anything
Duplicate files, duplicate totals: where AI reconciliation goes wrong
What AI prior-art search doesn't tell you it missed
How ARR works →