An outage at a key supplier costs your business a week of sales. Someone asks the contract-review AI a simple question: is the supplier liable?
The AI finds clause 14.2 of the master supply agreement: "The Supplier shall be liable for all losses arising from its negligence," and answers immediately.
"Yes. Under §14.2 the supplier is liable for all losses."
The clause didn't end there. On the next line it continues: "except for indirect or consequential losses, including loss of profit, which are excluded in full." Lost sales are exactly the kind of loss that carve-out is written to exclude. The AI never saw it.
The cases that make headlines are fabrications: lawyers fined for citing cases that never existed. Even purpose-built legal research tools have been found to get answers wrong in more than one in six queries, and general chatbots far more often (a Stanford study summarised by BriefCatch).
But in contract review the more common failure is quieter. Before your AI answers, the system splits your agreements into pieces and hands it the few that match the question. Those pieces are cut by length, not by meaning. A clause that runs onto the next line can be cut in two, and the exception lands in a piece the AI never receives.
The quote is real. The citation is correct. A spot-check confirms the words are in the contract. That's what makes it dangerous.
"Is the supplier liable if the outage cost us sales?"
The obligation lands in one piece, the carve-out in the next. Only the first piece reaches the AI.
"Yes, all losses." Accurately quoted. Legally wrong.
Your AI receives the first half and answers confidently on incomplete terms.
Your AI receives the whole clause, however it's laid out on the page.
Your AI quotes the obligation and never sees the "except where" that changes everything.
When two terms sit together, your AI is handed both, so the obligation and its exception arrive as one.
Your AI answers from whichever draft it found first, with no sign the signed version says something different.
The versions are paired and you see exactly what differs before your AI answers.
Your AI quotes one agreement and never mentions the same wording appears in another, with a different meaning there.
You're told outright whether the wording is unique or appears in several places.
Schedule 3 was never attached. Your AI summarises the deal as if it were complete.
Every document that's referenced but missing is listed, so your AI can say the deal isn't complete.
The usual safeguard is to verify the AI's citations. That catches invented cases. It doesn't catch this, because the citation is real and the quoted words really are in the contract. To catch a missing carve-out you'd have to reread the whole clause yourself, every time, and then the AI has saved you nothing.
ARR doesn't replace your model or your judgment. It changes what the model is handed: complete clauses, their exceptions, the right version, and a warning when something is missing.
ARR won't stop a model inventing a case that never existed; that's a different failure. What it addresses is the AI being handed the wrong or partial source. We've measured these abilities in our own benchmark on code and text, not yet on contracts, so treat this as how ARR is built to work here, not a measured result. That's what early access is for.
If it doesn't, the AI didn't misread the contract. It was never given the whole clause. That's the part ARR fixes.
We're opening ARR to a small group first. Tell us how your team works with contracts and we'll be in touch when your place is ready. No payment required.
More in ARUKAS Field Notes:
Your AI is answering from an old document
Next: Chunking breaks meaning
Why AI coding agents break things three files away
Running parallel coding agents without them overwriting each other
Which version of the policy did your AI just apply?
AI citation errors start before the AI writes anything
Duplicate files, duplicate totals: where AI reconciliation goes wrong
What AI prior-art search doesn't tell you it missed
How ARR works →