Your agent changed a function and fixed every caller it could find. The tests it ran passed. The break surfaced three files away, through a wrapper it never followed, on the first invoice in production.

You ask your coding agent to make format_price() return a Money object instead of a string. A reasonable change. The agent searches for format_price(, finds four call sites, and updates each one. One of them is money() in utils/money.py, a thin wrapper that just returns the result, so the agent decides it needs no change.
build_line_items() never mentions format_price, so no search for that name will ever find it. It calls money() and glues the result onto a string. Two files further on, render_invoice() is where it finally blows up, the first time a real invoice is generated.
"Updated format_price and all 4 call sites. Tests pass."
Most coding agents navigate by searching text: find the name, open the file, edit the match. That finds direct callers well. It can't follow a value once it passes through a wrapper under a different name, so everything one hop beyond the wrapper is invisible.
A bigger context window doesn't rescue it. Loading the whole repository spreads the model's attention thinner rather than sharper, and agents still miss cross-file dependencies with everything in view (DEV, Augment Code). Language servers help with direct references, but they need to be running for each language and they stop at the first hop too.
The agent finds every place format_price is written.
The value flows on through money() under a new name. Text search can't follow it.
"All call sites updated." True for the ones it could see.
This is the one field where we've measured ARR directly. We ran the tools coding agents rely on today against the same search tasks on real code. Here are the six that matter most for "will this change break something?"
| Task | Text search | File reads | git | Language server | Embeddings | ARR |
|---|---|---|---|---|---|---|
| What breaks, several levels deep | ✕ | ~ | ✕ | ~ | ✕ | ✓ |
| What does this function call? | ✕ | ~ | ✕ | ✓ | ✕ | ✓ |
| Name defined in many places (warn me) | ~ | ✕ | ✕ | ~ | ✕ | ✓ |
| Files needed but missing | ~ | ✕ | ~ | ✕ | ✕ | ✓ |
| Is my earlier reading still valid? | ✕ | ~ | ~ | ~ | ✕ | ✓ |
| Which functions changed (no git, no backup) | ✕ | ✕ | ✕ | ✕ | ✕ | ✓ |
✓ handles it · ~ partially · ✕ can't. Internal benchmark, not independently verified. A language server does answer "what does this function call?". We've kept that row in because it's the honest picture.
The agent fixes direct callers. Callers of the wrapper, and their callers, break later in tests or in production.
The dependency chain is followed several levels deep before the change is written, so every affected caller is in the same change.
The agent jumps to the first config or utils it finds and edits that one. The one actually imported stays untouched.
You're warned when a name has more than one definition, with every one of them listed.
The file changed since the agent read it. The edit lands on the wrong lines or quietly reintroduces deleted code.
The agent is told whether its earlier read is still valid before it writes.
A near-identical function with renamed variables doesn't match a name search, so the fix lands in one copy and the bug survives in the other.
Functions are matched by their shape, not their names, so every clone is returned and fixed together.
The agent invents a stub for an import that really is missing, or tries to recreate system headers and generated code it simply can't see.
Genuinely missing files are listed, and kept apart from system, built-in and generated ones.
Good test coverage will eventually catch the invoice bug. But by then the agent has finished, the change is in review or merged, and someone has to reconstruct a chain the agent never saw. Fixing it is cheapest before the agent writes, when the whole chain is known.
ARR sits under the agent you already use. It doesn't replace your model, your editor or your tests. It gives the agent what a careful engineer looks up before a risky change: what depends on this, and how far that goes.
The table above is from our own internal benchmark on a code corpus. It isn't independently verified, and real repositories vary. Some tasks also look different on larger setups, which is why early access runs on your code, not ours.
Everything on that list is a dependency your agent couldn't see. That's the part ARR traces.
We're opening ARR to a small group of engineering teams first. Tell us how you ship with agents and we'll be in touch when your place is ready. No payment required.
More in ARUKAS Field Notes:
Your AI is answering from an old document
When legal AI gets the clause wrong
Chunking breaks meaning
Running parallel coding agents without them overwriting each other
Which version of the policy did your AI just apply?
AI citation errors start before the AI writes anything
Duplicate files, duplicate totals: where AI reconciliation goes wrong
What AI prior-art search doesn't tell you it missed
How ARR works →