Scientific research › Kumar, 2026

AI Brand Visibility: What 100,000 AI Answers Actually Show

Kumar (arXiv:2606.20065, June 2026) published the first large-scale baseline of brand visibility across AI search engines: 102,025 prompt responses and 149,912 citations, covering 102 brands on ChatGPT, Gemini, Perplexity, Claude and Grok between March and May 2026. The headline is not the ladder of brand fame — it is where the citations go. Only 2.9% of them point at the brand’s own domain. Roughly three quarters point at corporate and third-party pages about the brand, and the single most-cited content format is the ranked “best-of” listicle. If your AI visibility plan consists entirely of editing your own website, this study says you are working on 2.9% of the surface.

The brand-stature ladder

On first measurement, against unbranded prompts, visibility sorts into three clean tiers. Global household names (the paper cites Stripe and Nike) appear in 72.9% of relevant answers. Established mid-market and regional brands (Olipop, Klaviyo) appear in 43.6%. Niche and small brands appear in 11.4%. Each step down costs roughly 30 percentage points, and the separation is statistically solid (Kruskal–Wallis H = 38.32, p = 4.78×10−9).

This matters because it sets an honest starting line. A small brand appearing in about one relevant answer in nine is not broken — it is normal, and it is the number to improve against. The study is explicit that the hard case is “everyone outside the already-authoritative top brands”: SMEs, D2C brands, creators and early-stage startups. Most published GEO advice is validated on exactly the brands that do not need it.

The results that matter

Finding Number Practical translation
Scope of the baseline 102,025 responses, 149,912 citations, 102 brands, 5 engines, Mar–May 2026 The first measurement large enough to argue with
Brand-stature ladder 72.9% / 43.6% / 11.4% Global vs mid-market vs niche visibility on first run — about 30 points per step
Citations pointing at the brand’s own domain 2.9% Your website is a small minority of what engines actually cite about you
Citations to corporate and third-party brand pages 75.2% The surface that decides your visibility is mostly owned by other people
Most-cited content format Ranked “best-of” listicles — 21.0% of all citations One page type carries a fifth of everything; being absent from it is expensive
Sentiment vs mention volatility 45.5% of runs flip sentiment vs 6.8% for mention (≈ 6.7×) Whether you appear is fairly stable; how you are framed is not
Mention determinism 77.5% of brand–prompt–engine cells always or never mentioned Most visibility is structural, not luck — which is why it can be worked on
Baseline drift with no intervention ChatGPT −1.34%/run, Perplexity −0.76%/run; others flat Doing nothing is not neutral on the two engines that move

Where engines actually get their sources

Across all 149,912 citations, the distribution is lopsided in a way that should redirect most budgets. Corporate and third-party brand pages take 75.2%. The brand’s own domain takes 2.9%. Among the remaining non-corporate sources, video leads: YouTube 4.2%, ahead of editorial tech and business media at 3.8%, Reddit and community forums at 3.3%, Wikipedia and reference at 2.6%, and software review sites such as G2 and Capterra at 1.1%.

Within content pages, the format breakdown is just as pointed. Ranked listicles account for 35.7% of cited content pages and 21.0% of all citations. Generic articles follow at 31.0% of content pages, how-to guides at 9.7%, comparison pages at 4.5%. In plain terms: the “best CRM for small business” type of page is the single highest-leverage object in AI search, and it is almost always somebody else’s page.

Per-engine differences

Day-1 recognition on unbranded prompts varies by engine: Perplexity 23.9%, ChatGPT 22.1%, Gemini 18.7%. The paper also reports Claude at 51.5% and Grok at 12.0%, both on a smaller agency-tier sample — treat those two as indicative rather than settled. The practical reading is the one CapstonAI measures against: there is no such thing as a single “AI visibility score.” A brand can be well established on one engine and invisible on another, and a blended number hides exactly the gap you would want to act on.

What the study does not claim — and who wrote it

Two caveats belong on the record, because they change how the numbers should be used.

  • It is a baseline, not a causal test. The paper measures what visibility looks like; it does not demonstrate that any specific tactic improves it. The author is direct about this and closes by proposing seven protocols (v1.1) designed to test causality — including a randomised closed-loop trial of implemented recommendations, a web-search-on/off natural experiment to separate training-data visibility from retrieval visibility, and a regression of visibility on schema markup. Those are stated expectations and study designs, not results.
  • The data comes from a commercial tracker. The brands analysed are those tracked on Ranqo, the author’s own product. That is a self-selected sample of companies already paying attention to AI visibility, which plausibly skews the tiers and the trajectories. It does not invalidate the distributions — 149,912 citations is a lot of evidence — but it is the kind of disclosure that should accompany the numbers wherever they are quoted, including here. CapstonAI works in the same market.

How this fits the research map

This is the empirical counterpart to Martinez’s critical survey, which found no reviewed technique with a proven effect on organic discoverability. Kumar does not contradict that — he measures the terrain the survey said was under-measured. Read together with Tian et al. on targeted repair and SAGEO Arena on structural GEO, the picture is consistent: work on being retrievable and correctly represented, measure per engine, repeat over time, and be suspicious of single scores.

What this means for your business

  • Audit the third-party surface first. Find the ranked listicles in your category and check whether you are on them, and how you are described. That surface carries 21% of citations; your own site carries 2.9%.
  • Set expectations by tier. If you are a small or niche brand, roughly 11% first-run visibility is the baseline — judge progress against that, not against a household name.
  • Track sentiment separately from mention. Framing flips in 45.5% of runs against 6.8% for mention, so a single unflattering answer is noise. Trend it; do not react to it.
  • Measure engine by engine. Day-1 recognition ranges from 18.7% to 23.9% across the three large-sample engines — a blended score would hide that entirely.
  • Do not assume stability. Baseline visibility drifts down on ChatGPT and Perplexity without intervention, so a flat trend line is already an outcome.

Frequently asked questions

Does this mean my website does not matter for AI visibility?

No. It means your website is not where most citations land. Only 2.9% of the 149,912 citations pointed at the brand’s own domain, while 75.2% pointed at corporate and third-party pages about the brand. Your site still has to be crawlable, accurate and on-topic — it is the source others draw from — but editing it alone addresses a small share of the cited surface.

What is the single highest-leverage page type?

The ranked “best-of” listicle. It is 35.7% of cited content pages and 21.0% of all citations in this dataset. Being present, accurately described and favourably placed on the listicles in your category is the most concentrated opportunity the data shows.

My brand appears in far fewer answers than a competitor. Is something wrong?

Probably not, if the competitor is larger. Visibility sorts by brand stature into roughly 72.9%, 43.6% and 11.4% tiers, about 30 percentage points apart. Compare yourself to brands of similar stature, and track your own trend over time rather than the absolute gap.

Why does the same brand look positive one week and negative the next?

Because sentiment is the unstable signal. It flipped in 45.5% of runs, against 6.8% for whether the brand was mentioned at all — roughly 6.7 times noisier. Sentiment needs repeated sampling before it means anything; a one-off screenshot does not.

Does this study prove that GEO tactics work?

No, and it does not claim to. It is a descriptive baseline of what visibility looks like across engines and brand tiers. The author proposes seven follow-up protocols specifically to test whether recommendations causally improve visibility. Until those run, treat the numbers as a map of the terrain, not as proof that a given tactic moves it.

See where you actually get cited, engine by engine — free AI visibility audit →

Related: All GEO scientific research · The GEO evidence audit · Citation failure repair · AI Search Watch