GEO Scientific Research 2026: Peer-Reviewed Evidence Map for AI Citation Optimization

GEO Scientific Research 2026: Peer-Reviewed Evidence Map for AI Citation Optimization

Most published GEO content in 2026 is vendor marketing dressed as research. Only a handful of peer-reviewed empirical studies actually quantify how ChatGPT, Perplexity, Claude, and Google AI Overview select and use citations. This silo page maps the academic GEO literature that practitioners can cite with confidence: the University of Toronto comparative study of AI Search vs. Google (Chen et al., 2025), the citation-selection-vs-absorption measurement framework from the geo-citation-lab dataset (Zhang, He & Yao, 2026), and the original GEO formalization by Aggarwal et al. Below: the empirical findings practitioners must internalize, the methodologies behind them, and a citation-ready reading list organized by topic.

TL;DR: Two academic papers anchor evidence-based GEO in 2026: (1) Chen et al. (arXiv:2509.08919) document an overwhelming earned-media bias in AI Search, low cross-language domain stability, and a structural big-brand bias; (2) Zhang, He & Yao (arXiv:2604.25707) separate citation selection from citation absorption, show that ChatGPT cites fewer but deeper sources, and reveal that Q&A formatting alone does not improve absorption. Use these as the empirical foundation, not vendor claims.

Every study in the lab

Each analysis below gets a dedicated deep-dive page in English and French. Newest first. Peer-reviewed and preprint work is listed alongside our own first-party measurement, which is labelled as such — the distinction matters and we make it visible rather than blurring it.

August 2026 · First-party measurement Not peer-reviewed · method disclosed

CapstonAI direct-measurement study — entry vs. rank

990 queries, 4,755 pages, 922 domains, 79 scored attributes. Google surfaces 966 domains across the query set; generative engines surface 162 to 348. Entity clarity governs entry into the answer; site-wide consistency governs rank.

2026 · Critical survey, 45 studies arXiv:2607.14035

Martinez — Optimizing visibility in generative engines

Proven GEO effects are conditional on context; no reviewed technique shows a stable, cross-platform effect on organic discoverability.

2026 · Large-scale baseline, 102k answers arXiv:2606.20065

Kumar — Brand visibility across AI search engines

Only 2.9% of citations point at the brand’s own domain; first-run visibility sorts by brand stature into 72.9% / 43.6% / 11.4% tiers.

2026 · Targeted repair experiment arXiv:2603.09296

Tian et al. — Citation failure repair

+40% relative citation gain while modifying only 5% of the content. Repair beats rewriting.

2026 · SAGEO Arena (Yonsei) arXiv:2602.12187

Kim et al. — Structural information GEO

Naive GEO often degrades retrieval (−4.54 on body-only edits); structural optimisation adds +22%, and +35% combined with sourced statistics.

2026 · 21,143 citations analysed arXiv:2604.25707

Zhang, He & Yao — Citation selection vs. absorption

ChatGPT cites fewer sources (6.88 per answer) but uses each far more deeply. Q&A formatting alone does not improve absorption (−5.74%).

2025 · University of Toronto arXiv:2509.08919

Chen et al. — Earned-media bias in AI Search

AI Search overwhelmingly cites earned media over brand-owned content. Cross-engine overlap is low (Jaccard 0.10–0.25); translating a query changes the entire evidence ecology.

2024 · KDD’24 arXiv:2311.09735

Aggarwal et al. — The GEO benchmark

Up to +40% in-testbed visibility. Verbatim quotation (+41%), added statistics (+34%) and cited sources (+29%) beat an authoritative tone alone (+13%).

These findings are what our free tools are built on, and each week’s changes are tracked in AI Search Watch.

Free CapstonAI scan →    Pricing

The academic GEO landscape in 2026

Three streams of academic work shape evidence-based GEO. The first stream formalized the field: Aggarwal, Murthy, Sheth, Bose & Krishna (2024) introduced the GEO framework, the GEO-bench benchmark, and showed that black-box content interventions can lift visibility by up to 40% in generative engines. The second stream measures observed engine behavior: Chen, Wang, Chen & Koudas (2025) at the University of Toronto ran large-scale controlled experiments across verticals, languages, and paraphrases to compare Google with ChatGPT, Claude, Perplexity, and Gemini. The third stream studies citation mechanics: Zhang, He & Yao (2026) built the geo-citation-lab dataset (602 prompts, 21 143 valid citations, 23 745 citation-level feature records) to separate which sources get selected from which sources actually shape the generated answer.

The combined picture is more useful than any single paper. Aggarwal et al. proved GEO works as an intervention. Chen et al. mapped which sources AI engines prefer (overwhelmingly earned media, with sharp engine-by-engine differences). Zhang et al. quantified the depth-vs-breadth tradeoff and identified the page-level features that drive answer-level influence.

What the Toronto study (Chen et al., 2025) established

The Toronto group ran controlled experiments comparing Google’s top-10 results with web-enabled responses from Claude (3.5 Sonnet), ChatGPT (4o search-preview), Perplexity (sonar-pro), and Gemini (2.5 Flash with Google Search grounding). Each cited URL was classified as Brand (official manufacturer/retailer), Earned (independent reviews, media, government), or Social (community platforms like Reddit, YouTube, Quora).

  • Earned-media dominance. Across automotive, consumer electronics, software, and other verticals, AI Search overwhelmingly cited earned sources. Claude and ChatGPT were the most earned-heavy (above 80% in most verticals); Gemini sat in the middle with more brand content; Perplexity included more social sources (notably YouTube).
  • Low cross-engine overlap. Jaccard similarity between engines on the same vertical was typically 0.10 to 0.25. Each engine is sampling a different evidence pool.
  • Cross-language instability. Domain overlap across languages was generally near zero for GPT (it swaps site ecosystems by language), while Claude maintained much higher cross-language stability by reusing the same authority domains.
  • Big brand bias. Unbranded prompts in the cola vertical defaulted to market leaders (Coca-Cola, Pepsi, Dr Pepper), with niche brands appearing at much lower frequencies.
  • Paraphrase sensitivity is smaller than language sensitivity. Reformulating a query changes some citations; translating it changes the entire evidence ecology.

Deep-dive: Earned media bias in AI Search →

What the citation-absorption study (Zhang, He & Yao, 2026) added

The Zhang group analyzed 602 controlled prompts across ChatGPT, Google AI Overview/Gemini, and Perplexity using the public geo-citation-lab dataset. They proposed and tested a two-stage measurement framework: citation selection (which sources a platform chooses) and citation absorption (how deeply a cited page shapes the generated answer, measured by a composite influence score combining reference count, position, paragraph coverage, TF-IDF similarity, and n-gram overlap).

  • Citation breadth and depth diverge. Mean citations per prompt: ChatGPT 6.88, Google 12.06, Perplexity 16.35. Mean fetched-page influence: ChatGPT 0.2713, Google 0.0584, Perplexity 0.0646. ChatGPT cites fewer sources but uses each one more intensively.
  • Q&A format alone is weak. Q&A pages averaged 0.0947 influence versus 0.1005 for non-Q&A pages, a relative difference of -5.74%. FAQ packaging without underlying evidence density does not earn deeper absorption.
  • Evidence genres matter. Pages containing code (+76.88%), numbers/statistics (+61.55%), definition markers (+57.33%), comparison content (+55.28%), and how-to content (+41.20%) showed substantially higher mean influence.
  • Domain type ranks differently in selection vs. absorption. News appears frequently in candidate pools but news_media pages averaged only 0.0726 influence, while encyclopedia pages averaged 0.2144. Selection probability and absorption intensity are separate outcomes.
  • Top selected domains in the dataset: youtube.com (560), en.wikipedia.org (352), reddit.com (315), reuters.com (287), linkedin.com (187), nytimes.com (174), pmc.ncbi.nlm.nih.gov (167), facebook.com (151), forbes.com (146), finance.yahoo.com (146).

Deep-dive: Citation selection vs. absorption →

Combined empirical map: what practitioners should believe in 2026

ClaimEvidenceSourceConfidence
AI engines overwhelmingly prefer earned media over brand-owned contentMultiple verticals, multiple engines, Brand/Earned/Social classificationChen et al. 2025High (replicated across verticals)
ChatGPT cites fewer sources but uses them more deeply than Perplexity or Google21 143 citations, mean influence 0.2713 vs. 0.0646 vs. 0.0584Zhang, He & Yao 2026Medium-high (snapshot, single dataset)
Q&A formatting alone does not improve absorption23 745 citation features, -5.74% relative differenceZhang, He & Yao 2026Medium (descriptive only, no controlled intervention)
Cross-language domain overlap is engine-dependent and generally low for GPTCross-language Jaccard heatmaps across 6 languagesChen et al. 2025High (multiple language pairs)
Unbranded prompts default to market leaders (big brand bias)Cola vertical: Coca-Cola, Pepsi dominate ChatGPT and Perplexity outputsChen et al. 2025Medium-high (one vertical extensively tested)
Evidence genres (definitions, stats, comparisons) drive deeper absorption than format alone+57% to +77% mean influence upliftZhang, He & Yao 2026Medium (observational, not experimental)
GEO interventions can lift visibility up to 40% in benchmark conditionsGEO-bench experiments on Perplexity.ai and synthetic benchmarksAggarwal et al. 2024Medium (benchmark, not field)

First-party measurement: what makes a page cited, and what makes it rank

Everything else on this page is third-party work we did not run. This section is ours, and we flag it for what it is: an internal measurement study completed in August 2026, not a peer-reviewed publication. The dataset is not public. We publish it because the method — including the controls that could have invalidated it — is disclosed below, and because the finding on index compression has no equivalent in the literature.

What was measured

330 search intentions, each written in 3 formulations, produced 990 queries. Those queries were run against five sources — ChatGPT with and without web access, Gemini with and without grounding, and Google as a control. Every cited page was captured with a headless browser (text, images and screenshots), producing 4,755 pages across 922 domains. Each page was then scored on a battery of 79 attributes in 7 categories on 0–4 scales, for roughly 300,000 attribute-by-page evaluations, using an open-weights model for text and a vision model for the captures.

Finding 1 — the citation index is five times smaller

On the same 990 queries, Google surfaced 966 distinct domains. The generative engines surfaced between 162 and 348. This is the single number practitioners should carry: the web did not merely change format, the number of sites that exist at all from the answer layer’s point of view shrank by a factor of three to six. Competition for citation is not a harder version of competition for ranking; it is a different-sized board.

Finding 2 — being cited and being ranked are two different problems

The study deliberately separated two questions, and they did not have the same answer.

  • Being cited at all is a page-level property, driven by entity clarity — whether a machine can tell who you are, where you are, and what makes you distinct. The two engines rank these attributes almost identically, which is unusual and makes the lever unusually robust: consistency of structured identity data scores 0.86 on Gemini and 0.87 on ChatGPT, and canonical identification of the establishment 0.85 on both.
  • Ranking well once cited is a site-level property, driven by consistency — nine site-wide consistency attributes dominate every ranking on all five sources (0.55 to 0.60 in head-to-head duels, +0.27 to +0.31 in correlation). Terminological consistency, brand hierarchy, internal entity linking and page architecture belong to this group. No single page can fix them.

Finding 3 — web search widens coverage, it does not level the field

Pages cited only thanks to web search scored lower than pages cited from the model’s own memory. But when a page was cited on both sides, search pushed it up: +0.12 on ChatGPT (p = 0.006) and +0.13 on Gemini (p = 0.004). On ChatGPT the main effect is coverage — 300 of the 990 queries returned nothing from memory, and browsing answered 221 of them. Being findable does not substitute for being known; it adds a second, narrower door.

How we tried to break it

This is the part we would want to see in anyone else’s study, so it is the part we publish in most detail. Four independent analytical angles were used, including a random forest evaluated on domains never seen during training, which neutralises both attribute overlap and memorisation of well-known sites. Two controls were run on data whose answer was known in advance:

  • Negative control — entirely random scores. The classical binomial test declared this pure noise significant at q ≈ 0.0001. The test we retained declared nothing. That result is worth pausing on: a standard, widely used statistical method reported a confident finding on data containing no signal whatsoever.
  • Positive control — an attribute fabricated to follow rank exactly. It was detected at the top (79% of duels, +0.75), with no spurious findings elsewhere.

What this study does not say

The design is observational. It measures associations; it does not demonstrate causation, and no page here should be read as a promise that changing an attribute will move a number. Effect sizes are also uneven: around +0.25 on average, but as low as +0.01 on some sources. The study tells you what to work on first, not what each point is worth. Treat it as a prioritisation instrument, not a forecast.

Test these findings on your own site →

The full GEO research reading list

Vague 21 — Toronto empirical study deep-dives

Vague 22 — Citation absorption framework deep-dives (FR)

Wave 23 — Extended reading list (verified on arXiv)

How to cite this research properly

Both primary studies are available on arXiv with permanent identifiers. Cite them as follows in any GEO whitepaper, blog post, or client report:

  • Chen, M., Wang, X., Chen, K., & Koudas, N. (2025). Generative Engine Optimization: How to Dominate AI Search. University of Toronto. arXiv:2509.08919. https://arxiv.org/abs/2509.08919
  • Zhang, K., He, X., & Yao, J. (2026). From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms. arXiv:2604.25707. https://arxiv.org/abs/2604.25707
  • Aggarwal, P., Murthy, V., Sheth, V., Bose, A., & Krishna, S. (2024). GEO: Generative Engine Optimization. arXiv:2311.09735.

FAQ — GEO scientific research

Are these papers peer-reviewed?

Chen et al. (2025) is an arXiv preprint from the University of Toronto research group; Zhang, He & Yao (2026) is also an arXiv preprint with public dataset and reproducible pipeline. Aggarwal et al. (2024) appeared at KDD 2024. arXiv preprints are common in this fast-moving field because journal cycles are too slow for the pace of AI search evolution. Treat findings as empirical descriptions of observed behavior at the moment of publication, not as permanent laws.

Why focus on academic studies when vendors publish their own data?

Vendor data is valuable but suffers from selection bias (vendors report what makes their tools look good). Academic studies use open prompt sets, multiple platforms, transparent methodology, and reproducible pipelines. They are the only sources that can be cited as independent evidence in client proposals and board reports. That objection applies to us too, which is why our own first-party study above is labelled separately, why its limits are stated in the same section as its findings, and why we publish the negative control that could have invalidated it. Judge it on its method, not on our logo.

How fast does this research age?

Quickly. The Toronto group explicitly warns that engine behavior is dynamic and that exact percentages should be treated as illustrative of relative trends, not permanent facts. Re-baseline every 6 months by running the same prompt panels and comparing engine outputs. The conceptual frameworks (earned-media bias, evidence-container hypothesis, selection vs. absorption) are more durable than the specific numbers.

Are there other peer-reviewed GEO sources worth tracking?

Yes. The Aggarwal et al. (2024) original GEO paper introduced GEO-bench. Citation-repair, agentic GEO, and structural GEO work has expanded the field. Studies of source attribution in answer engines (mentioned in the related-work sections of both Chen et al. and Zhang et al.) document hallucination, inaccurate citation, and evidence-claim mismatches. Track the arXiv cs.IR (Information Retrieval) and cs.CL (Computation and Language) categories monthly for new contributions.

Methodology cheat-sheet (cite these in your own work)

  • Chen et al. methodology: ranking-style prompts (e.g., “Top 10 X”), web-enabled API calls to each engine, domain extraction via tldextract, Brand/Earned/Social classification with rule-based and AI-assisted labeling, Jaccard overlap for cross-engine and cross-language stability, paraphrase templates (justification, source, quote, confidence, ranked, imperative, keyword-only).
  • Zhang, He & Yao methodology: 602 prompts across 4 layers (main A=432, style B=60, language C=60, scenarios D=50), influence_score = 0.20·ref_count + 0.15·(1-first_position_ratio) + 0.20·paragraph_coverage + 0.25·TF-IDF cosine + 0.20·(bigram+trigram overlap)/2, 72 feature dimensions per citation, fractional logit or beta regression recommended for absorption modeling.
  • Aggarwal et al. methodology: GEO-bench benchmark with synthetic and real-world queries, content intervention strategies (citations, quotations, statistics, formal tone), position-adjusted word count and subjective impression scores as visibility metrics.

Tools and related reading

🇫🇷 Version française : Recherche Scientifique GEO 2026 — Toutes les analyses Zhang 2026 disponibles en français, avec les 5 pages V22 de CapstonAI.

Ready to apply GEO research to your brand?

Free CapstonAI scan →

Last updated: May 2026. Primary sources: Chen, M., Wang, X., Chen, K., & Koudas, N. (2025). Generative Engine Optimization: How to Dominate AI Search. University of Toronto. arXiv:2509.08919. Zhang, K., He, X., & Yao, J. (2026). From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms. arXiv:2604.25707. Aggarwal, P., Murthy, V., Sheth, V., Bose, A., & Krishna, S. (2024). GEO: Generative Engine Optimization. arXiv:2311.09735.

🇫🇷 Version française : Recherche Scientifique GEO 2026 — Toutes les analyses Zhang 2026 disponibles en français.

Free GEO Tools & Templates

Apply this research to your business — download the free calculators, audits and playbooks.

GEO ROI Calculator for CFOs

15-minute business case based on Zhang et al. 2026 + 86-customer cohort benchmarks.

GEO Metrics Defensibility Audit

25-point checklist to make your AI visibility metrics survive a board-level audit.

Multi-Engine GEO Scorecard

Score your brand across ChatGPT, Perplexity and Google AI Overview. 75% of cited sources differ between engines.

Test these findings on your own site

Research tells you how AI engines behave in general. It cannot tell you whether they cite you. The free CapstonAI audit turns these findings into your numbers: mentions, citations and gaps per engine, for your market’s real buyer prompts.

Run the free AI visibility audit →
Or check a single page’s structure with the free Chrome extension