Which AI Search Statistics Come From Datasets Nobody Can Inspect?
· Published August 4, 2026 · 4 min read
Claims last verified August 19, 2026
The figures as commonly quoted: three findings that get stated as industry facts, in decks and in strategy documents, with no qualifier attached.
All three come from vendor datasets that exist nowhere outside the single article reporting them. Search for the numbers and you get the article and its requoters.
The Consensus Gap
As quoted: 2.37% of cited URLs appear in all three of ChatGPT, Perplexity and Google AI Overviews for the same prompt, and 91.07% appear in only one.
What was measured: 3.7 million URL citations, Q3 2025 to Q1 2026, from data supplied by Omnia, a visibility tracking vendor. Reported by Kevin Indig on May 11, 2026. Omnia publishes no report, no blog post and no dataset containing these figures.
Where they diverge: Indig’s own disclosure is better than the requoting deserves. He publishes the caveats in full. The pool is skewed toward Omnia’s customer base, weighted toward Spain plus the UK, the Nordics and other EU markets, with prompts run in each country’s primary language, and his stated instruction is to treat the findings as directional for European AI search. That geographic limit is dropped almost everywhere the number appears. Our own attribution gap piece quoted 91% with no market attached until 2026-08-19, when the European scope was added to both places it appears there.
The direction is independently corroborated. Writesonic analysed 161,286 prompts across four engines in May and June 2026, of which 70,879 returned citations from all four, and found 3.8% of sources agreed by all four and 72% to 73% of cited domains appearing on exactly one engine. That is a domain-level measurement against Omnia’s URL-level one, and URL-level overlap should be lower, which is the direction the two differ in. The Consensus Gap’s direction survives independent checking. Its decimals do not.
The ski ramp
As quoted: 44.2% of citations come from the first 30% of a page.
What was measured: 18,012 verified citations, supplied by Gauge, an AI visibility analytics vendor, reported by Indig on February 16, 2026. Gauge publishes no standalone research report on citation position.
Where they diverge: less than in the other two cases. The method disclosure inside the article is unusually complete for vendor data. The sentence transformer is named, the vector dimensionality given, the matching cosine similarity and the confidence threshold stated. That is more than most vendor research offers. It still cannot be checked by anyone outside the two parties, because the data is not published.
The 74% shortlist figure
As quoted: three quarters of buyers pick the top AI recommendation.
What was measured: 48 US participants and 185 task observations, run by Citation Labs and Clickstream Solutions, reported April 7, 2026. 74% of participants chose the item ranked first in an AI Mode response. Citation Labs does not publish the study.
Where they diverge: Indig states explicitly that the findings are directional and not population-level estimates. By the time the figure reaches a slide it is a fact about buyers rather than about 48 people doing 185 tasks. A qualitative study of 48 people is a legitimate and useful thing. It is not a population estimate, and 74% of 48 is 36 people.
The pattern worth naming
None of this means the analyses are bad. It means one thing worth saying out loud: each of these vendors sells a product the finding makes more valuable. Omnia sells multi-engine tracking, and a finding that engines are largely disjoint argues for multi-engine tracking. Gauge sells content optimization, and a finding that models read the top of the page hardest argues for content optimization.
That is not an accusation of bad faith. It describes which analyses get funded and published, and the visible consequence is that this literature contains almost no null results.
We also cannot prove a negative here. The evidence is an absence: no report on the vendor’s site, and a search that returns only the original article and its requoters. One published methodology page from Omnia, Gauge or Citation Labs would settle it, and we would correct this page the same day.
The honest version of the claim
Cite all three as “reported by Kevin Indig from data supplied by [vendor]”, never as an industry finding. Keep the European scope attached to the Consensus Gap. Keep the 48 participants attached to the 74%.
This page is part of Where the AI Search Statistics You Keep Quoting Actually Came From, an audit of seven of the most repeated figures in AI search optimization, with the checks we run on each one.