The paper is real. The DOI resolves. The number is in the results section. And the sentence right after it, the one that says the effect didn't replicate, never made it into the answer. That's the citation error your reference checker won't catch.

Your lab is scoping a new project and asks its AI research assistant a simple question: does Compound X reduce inflammation markers in mice? The assistant searches your team's library and finds the paper you'd expect. It answers in seconds, with a citation.
"Compound X reduces IL-6 by 40% in mice [Results 3.2]."
Section 3.2 does say that. It says it about the higher dose, in the first cohort of eight animals. The very next sentence reads: "The effect was not seen at the lower dose, and did not replicate in cohort 2."
Nothing was fabricated. Every element of the summary can be traced to the paper. It's still the wrong conclusion, and it's exactly the kind that ends up in a grant application, a slide deck or a literature review.
The headline problem is real. In one multi-model study of AI-generated academic references, only about a quarter were entirely correct and nearly 40% were wrong or invented (Enago), and a Lancet analysis reported a sharp rise in fabricated references in published papers (STAT).
But research assistants that read your own library fail differently. Before the AI writes a word, the system splits papers into pieces and hands it the few that best match your question. A striking result matches well. The limitation stated in the next sentence often lands in a different piece, one that never gets picked. The AI then summarises what it was given, faithfully and incompletely.
"Does Compound X reduce inflammation markers in mice?"
The headline result is retrieved. The caveat in the next sentence isn't.
Every word traceable to the paper. The conclusion still wrong.
Your AI quotes the headline result and leaves out the dose, cohort or replication caveat stated right after it.
When a finding and its caveat sit together, your AI is handed both, along with the complete section around the result, not a fragment.
The preprint reported 40%; the published version, after review, reported 22%. Your AI blends them, or quotes the one you'd least want.
The versions are paired and you see exactly what changed between them.
A collaborator re-ran the analysis and updated the data. Your AI keeps reporting the numbers it read last month.
Your AI is told when data it read has changed since, so it doesn't report stale numbers as current.
The same result appears in a preprint, the published paper and a conference abstract. Your AI counts it as three pieces of evidence.
Each distinct passage comes back once, so repeated text isn't mistaken for independent support.
A gene, compound or author written slightly differently, and your AI reports that it isn't mentioned in your library.
Your AI still finds it when the spelling is off.
Reference checkers do valuable work: they catch DOIs that don't resolve and papers that were never written. They can't catch this, because the reference is genuine and the quoted result really is in the paper. To catch a dropped caveat you'd have to reread the section yourself, every time, and then the assistant hasn't saved you much.
ARR doesn't replace your model or your scientific judgment. It changes what the model is handed: complete results with their limitations, a clear signal when versions of a paper or dataset disagree, and no inflated count from repeated text.
ARR won't stop a model inventing a reference that doesn't exist; that's a different failure, and reference checkers remain the right tool for it. Finding related work purely by topic is also still an area where embedding search does better. We've measured ARR's abilities in our own benchmark on code and text, not yet on scientific literature, so treat this as how ARR is built to work here. That's what early access is for.
If it doesn't, the AI didn't misread the paper. It was never handed the caveat. That's the part ARR fixes.
We're opening ARR to a small group first. Tell us how your team works with the literature and we'll be in touch when your place is ready. No payment required.
More in ARUKAS Field Notes:
Your AI is answering from an old document
When legal AI gets the clause wrong
Chunking breaks meaning
Why AI coding agents break things three files away
Running parallel coding agents without them overwriting each other
Which version of the policy did your AI just apply?
Duplicate files, duplicate totals: where AI reconciliation goes wrong
What AI prior-art search doesn't tell you it missed
How ARR works →