ARR/Field Notes/Research
Field notes · Scientific research

AI citation errors start before the AI writes anything.

The paper is real. The DOI resolves. The number is in the results section. And the sentence right after it, the one that says the effect didn't replicate, never made it into the answer. That's the citation error your reference checker won't catch.

For researchers, lab leads & systematic reviewers · 6 min read
An illustrative results section. The AI received only the first sentence, that Compound X reduced IL-6 by 40% in mice at the higher dose. The next sentence, saying the effect was not seen at the lower dose and did not replicate in a second cohort, never reached it, so its summary reports the 40% reduction without the caveat.
Click to open full size
The story · an illustrative example

A real finding. Half of it.

Your lab is scoping a new project and asks its AI research assistant a simple question: does Compound X reduce inflammation markers in mice? The assistant searches your team's library and finds the paper you'd expect. It answers in seconds, with a citation.

AI summary
"Compound X reduces IL-6 by 40% in mice [Results 3.2]."
real paper · real number · missing caveat

Section 3.2 does say that. It says it about the higher dose, in the first cohort of eight animals. The very next sentence reads: "The effect was not seen at the lower dose, and did not replicate in cohort 2."

Nothing was fabricated. Every element of the summary can be traced to the paper. It's still the wrong conclusion, and it's exactly the kind that ends up in a grant application, a slide deck or a literature review.

Why it happens

Fabrication is the loud failure. Fragments are the quiet one.

The headline problem is real. In one multi-model study of AI-generated academic references, only about a quarter were entirely correct and nearly 40% were wrong or invented (Enago), and a Lancet analysis reported a sharp rise in fabricated references in published papers (STAT).

But research assistants that read your own library fail differently. Before the AI writes a word, the system splits papers into pieces and hands it the few that best match your question. A striking result matches well. The limitation stated in the next sentence often lands in a different piece, one that never gets picked. The AI then summarises what it was given, faithfully and incompletely.

01

Someone asks

"Does Compound X reduce inflammation markers in mice?"

02 · where it breaks

The finding travels alone

The headline result is retrieved. The caveat in the next sentence isn't.

03

A faithful half-summary

Every word traceable to the paper. The conclusion still wrong.

Five signs, five fixes

Every citation checks out. The science doesn't.

01
The dropped caveat

The limitation in the next sentence

Your AI quotes the headline result and leaves out the dose, cohort or replication caveat stated right after it.

✓ With ARR

When a finding and its caveat sit together, your AI is handed both, along with the complete section around the result, not a fragment.

02
Preprint vs published

Two versions, two numbers

The preprint reported 40%; the published version, after review, reported 22%. Your AI blends them, or quotes the one you'd least want.

✓ With ARR

The versions are paired and you see exactly what changed between them.

03
The updated dataset

The results file changed

A collaborator re-ran the analysis and updated the data. Your AI keeps reporting the numbers it read last month.

✓ With ARR

Your AI is told when data it read has changed since, so it doesn't report stale numbers as current.

04
The double count

One finding, three appearances

The same result appears in a preprint, the published paper and a conference abstract. Your AI counts it as three pieces of evidence.

✓ With ARR

Each distinct passage comes back once, so repeated text isn't mistaken for independent support.

05
The spelling variant

IL-6 vs IL6, Müller vs Mueller

A gene, compound or author written slightly differently, and your AI reports that it isn't mentioned in your library.

✓ With ARR

Your AI still finds it when the spelling is off.

Why citation checkers aren't enough

A checker confirms the paper exists. Not that the conclusion survives it.

Reference checkers do valuable work: they catch DOIs that don't resolve and papers that were never written. They can't catch this, because the reference is genuine and the quoted result really is in the paper. To catch a dropped caveat you'd have to reread the section yourself, every time, and then the assistant hasn't saved you much.

ARR doesn't replace your model or your scientific judgment. It changes what the model is handed: complete results with their limitations, a clear signal when versions of a paper or dataset disagree, and no inflated count from repeated text.

A fair caveat

ARR won't stop a model inventing a reference that doesn't exist; that's a different failure, and reference checkers remain the right tool for it. Finding related work purely by topic is also still an area where embedding search does better. We've measured ARR's abilities in our own benchmark on code and text, not yet on scientific literature, so treat this as how ARR is built to work here. That's what early access is for.

Try this tomorrow

The caveat test

  1. Pick a paper in your library with an important limitation stated right after its main result.
  2. Ask your AI assistant for that paper's main finding.
  3. Check whether the limitation appears anywhere in its answer.

If it doesn't, the AI didn't misread the paper. It was never handed the caveat. That's the part ARR fixes.

Early access

Be one of the first research teams to test ARR on your own library.

We're opening ARR to a small group first. Tell us how your team works with the literature and we'll be in touch when your place is ready. No payment required.

More in ARUKAS Field Notes:
Your AI is answering from an old document
When legal AI gets the clause wrong
Chunking breaks meaning
Why AI coding agents break things three files away
Running parallel coding agents without them overwriting each other
Which version of the policy did your AI just apply?
Duplicate files, duplicate totals: where AI reconciliation goes wrong
What AI prior-art search doesn't tell you it missed
How ARR works →