Scientific research › Tian et al., 2026

Citation Failures: The Study Proving You Should Diagnose Before You Optimize

Tian, Chen, Tang, Liu & Jia (arXiv:2603.09296) built the first taxonomy of citation failure modes — the distinct reasons a page never gets cited by an AI engine — and an agentic system (AgentGEO) that diagnoses the failure, applies a targeted repair, and iterates. Result: over 40% relative improvement in citation rates while modifying only 5% of the content, versus 25% modified by baseline methods. The lesson is blunt: blind optimization rewrites eight times more of your content for worse results. Diagnosis first, surgery second.

Why “not cited” is not one problem

A page can fail to be cited at different stages of the pipeline: never retrieved, retrieved but not fetched, fetched but unparseable, parsed but unattributed, read but unused in the final answer. Each failure mode has a different cause and a different fix. Treating them all with the same “make the content better” pass is why so much GEO effort produces nothing measurable. The paper also distinguishes contribution (your content influenced the answer) from citation (you got the visible credit and the traffic) — only the second sends you buyers.

The results that matter

Finding Number Practical translation
Targeted repair beats blind rewriting +40% relative citation improvement, touching only 5% of content Small, diagnosed edits outperform full rewrites
Baselines modify far more for less ~25% of content modified by baseline methods Un-targeted optimization is expensive and risky
Generic optimization can harm long-tail content Niche pages need diagnosis even more than head pages
Some failures resist content fixes entirely If the blocker is authority or entity recognition, no rewrite will save you — the fix is off-page

How this fits the research map

Aggarwal et al. (KDD 2024) proved content changes can move AI visibility. SAGEO Arena (Kim et al., 2026) warned that naive changes often backfire. Tian et al. close the loop: the way out is a diagnose → repair → verify cycle, stage by stage, page by page. That is the operating model modern GEO converges on — and it is precisely the loop CapstonAI runs: measure your mentions per engine, identify the failure mode behind each gap, fix it with minimal, reviewed edits, and re-measure.

What this means for your business

  • Stop rewriting whole pages on instinct. Find why each page is skipped, and fix that.
  • Protect your long-tail pages: generic AI-optimization passes can make them less citable, not more.
  • Accept that some gaps are off-page problems — entity signals, third-party validation — and budget for them separately.
  • Verify every fix with per-engine measurement. Uncited fixes are unfinished fixes.

Frequently asked questions

What is a citation failure?

Any point in the AI search pipeline where your page drops out: not retrieved, not fetched, not parsed, not attributed, or not used in the final answer. Each stage is a distinct failure mode with its own fix.

What is AgentGEO?

The agentic system built by Tian et al. (arXiv:2603.09296): it diagnoses a page’s failure mode using their taxonomy, selects a targeted repair from a tool library, and iterates until citation is achieved.

Why does modifying less content work better?

Because most of a failing page is not the problem. Diagnosed repairs touch the specific blocker; blanket rewrites disturb everything else — including what already worked.

Can every page be repaired into a citation?

No. The study found some documents face challenges optimization alone cannot fix — typically authority and entity-level gaps. Those need off-page work, not another rewrite.

Get your citation failure diagnosis — free AI visibility audit →

Related: All GEO scientific research · Engine citation behavior · AI citation tracking · Citation selection vs absorption


Veille IA · AI Search Watch
Chaque semaine : AEO, GEO et modèles IA — l’impact sur le référencement des entreprises de toutes tailles. / Weekly AEO & GEO briefing.