Scientific research › Martinez, 2026

The GEO Evidence Audit: What 45 Studies Actually Prove

Martinez (arXiv:2607.14035, July 2026) published the first critical survey of generative engine optimization, reviewing 45 studies from November 2023 to July 2026. Its conclusion is uncomfortable and useful: GEO techniques can measurably change how an already-retrieved page is cited, but “no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior.” In plain terms — the research shows you how to be quoted better once the engine has your page. It does not yet show anyone how to make the engine find you in the first place. Those are two different problems, and most of the market sells them as one.

Where the famous “+40%” actually comes from

Almost every GEO pitch deck quotes “up to 40% more visibility.” The survey traces that number to its source and shows what it measures. In Aggarwal et al. (KDD 2024), the metric is position-adjusted word count — the share of the generated answer attributed to your source, weighted by where it appears. For the quotation-addition tactic it rose from 19.3 to 27.2, roughly +41% in relative terms.

The survey is explicit that this “does not mean that 40% more readers will click, nor that a page will gain 40% in retrieval probability. It means that, in this testbed, a source already provided to the generator receives a larger position-weighted share of attributed text.” The generalised claim that “GEO increases visibility by 40%” is listed among the claims the review rejects as unsupported. The effect is real; the interpretation sold around it is not.

The visibility pipeline, not the ranking

The paper’s central reframing: visibility in generative engines is “not a single ranking task but a stochastic, partially observable pipeline.” It separates seven outcome variables rather than one score — activation, crawling and indexing, retrieval, reranking and context allocation, generation and citation, absorption and fidelity, and finally attention, clicks and conversion. A tactic that improves one stage can be irrelevant, or harmful, at another. This is why a single “AI visibility percentage” is a misleading KPI: it collapses seven different failure points into one number that cannot tell you which one is broken.

The results that matter

Finding Number Practical translation
Scope of the audit 45 studies, Nov 2023 – Jul 2026 The first systematic check of what the GEO field has actually established
Origin of the “+40%” claim 19.3 → 27.2 position-adjusted word count (≈ +41% relative) A share-of-answer-text gain in a fixed testbed — not clicks, not retrieval
Most reproducible levers Topical relevance and position in context What survives replication: match the query, and get placed early in the model’s context
Answer instability between days Jaccard 0.34 – 0.42 across repeats The same query returns substantially different sources day to day — single measurements are noise
ChatGPT repetitions that never searched the web 57.8% Much of what people call “AI visibility” never touches your site at all
Google AI Overview citations outside the organic top 10 53% of cited domains Ranking first and being cited are two different games
Proven effect on organic discoverability None found across 45 studies No technique is yet shown to make you findable — only more citable once retrieved

The methodological warnings worth knowing

The survey catalogues why so many GEO results overstate themselves, and each item is a question you can ask any vendor:

  • Fixed-context testing. Most experiments inject content into documents the engine was already given, so they cannot measure discoverability at all.
  • Underestimated randomness. With day-to-day Jaccard overlap of 0.34–0.42, a before/after screenshot proves nothing without repeated sampling.
  • LLM-judge circularity. The risk peaks “when the same model generates the rewrite, the response, and the score.”
  • Missing denominators. Citation rates computed only over answers that already contained citations build in selection bias.
  • No user validation. Position-weighted metrics assume earlier citations get more attention, without eye-tracking or behavioural evidence.

The review assigns high confidence only to effects that are conditional on context, and low to very low confidence to claims about organic discovery or conversion. Its summary of the commercial discourse is direct: “claims about GEO return on investment clearly outstrip the academic evidence.”

How this fits the research map

This survey does not contradict the studies we have covered — it locates them. Aggarwal et al. (2024) showed content changes can move attributed text. SAGEO Arena (Kim et al., 2026) showed naive rewriting often degrades retrieval and that structure limits the damage. Tian et al. (2026) showed diagnosed, targeted repair beats blanket optimization. Martinez adds the frame around all three: these are context-stage results, they are real, and they should be sold as what they are. That is the standard CapstonAI measures against — per-engine, repeated over time, with the denominator shown.

What this means for your business

  • Treat any flat “+X% AI visibility” promise as a question, not a fact: at which pipeline stage, measured how, over how many repetitions?
  • Fix retrieval before rhetoric. If the page is not crawlable, indexed and topically tight on the query, no amount of rewriting is being tested by the research.
  • Measure repeatedly. With answer overlap between 0.34 and 0.42 day to day, one check tells you almost nothing.
  • Do not assume your Google ranking protects you: 53% of domains cited in AI Overviews are not in the organic top 10.
  • Track citation share per engine over time — the one metric the evidence base actually supports.

Frequently asked questions

Does this survey say GEO does not work?

No. It says GEO has proven effects at one stage of the pipeline — changing how an already-retrieved page is cited and quoted — and that no reviewed technique has yet demonstrated a stable, cross-platform effect on being discovered organically. The work is real; the marketing around it overreaches.

Is the “+40% visibility” figure false?

The measurement is genuine but routinely misdescribed. It refers to position-adjusted word count rising from 19.3 to 27.2 for quotation addition in Aggarwal et al.’s testbed — a larger share of the answer text for a source already given to the model. It is not 40% more clicks and not a 40% better chance of being retrieved.

What should I measure instead of a single visibility score?

Citation share per engine, sampled repeatedly over time, with the denominator stated. The survey’s seven-stage pipeline is the reason: one number cannot tell you whether your problem is crawling, retrieval, context position or citation.

What does the evidence support doing?

Topical relevance to the query and favourable position in the model’s context are described as the most reproducible levers. Practically: be genuinely on-topic, be structurally clean so you survive retrieval and parsing, and carry extractable evidence — definitions, figures, comparisons, steps.

See your real citation share, engine by engine — free AI visibility audit →

Related: All GEO scientific research · Citation failure repair · SAGEO Arena & structural GEO · AI Search Watch