Four Numbers We Published and Had to Correct
· Published August 4, 2026 · 4 min read
Claims last verified August 19, 2026
We audit other people’s statistics, so our own belong on the record first. Every error below is ours, most of them from the same week we started checking everyone else’s.
We have made every category of error in this piece, some of them this week.
We reported a correlation from a compressed range. An eight-company pilot gave a Spearman correlation of 0.783 between our Retrieval Readiness score and AI citations. Every company in that pilot scored between 84 and 90, and a correlation computed across a range that narrow carries almost no information. The full run across 43 companies on August 3, 2026 returned a rank correlation of -0.068 between readiness and citation rate, and organic traffic to citation volume over the same 43 companies came back at 0.878. Brand size dominates; our readiness score does not predict citation rate at all in that cohort. We published the null.
Our scoring rubric rewarded pages with nothing on them. Version 1 of our index was purely additive, so structural hygiene earned points regardless of whether the page had content. The homepage of bybit.com returned 92,941 bytes of HTML containing 2 words of extractable text and scored 54 out of 100. That is the failure mode of any additive AEO checklist with no gate: a page that cannot be read at all collects points for schema, crawler access and clean markup. Version 3 replaces the relevant checks with gates that cap the score.
Measuring less made scores look better, and we quantified it. Where a site blocks us, some dimensions cannot be measured, and our scorer removes those weights from the denominator rather than scoring them zero. That is the right call, and it has a direction. Across the 29 companies in our cohort measured on all eight dimensions and capped by no gate, scoring only the four dimensions we can measure everywhere produces a score 11.6 points higher on average than scoring all eight, computed from our own cohort file on August 4, 2026. Any partially measured domain in our index is therefore flattered relative to a fully measured one, and the index has to say so wherever the two appear in the same table.
We published a study and retracted it three times in one day. On August 4, 2026 we published a citation study and pulled it. The headline said only 15% of AI citations in buying answers reach a vendor site. That number came from defining “vendor” as one of the 55 companies in our own index, so every exchange outside our sample counted as a non-vendor. Classified by what a domain actually is, vendor and brand sites take 81.5% of buy-intent citations and 78.7% of comparison-intent citations. The reversal we described did not exist. The prompt design was also confounded: the comparison set asked about product categories and the buy set asked about assets, so subject and intent varied together and neither could be attributed. And a 31% citation rate had been presented as an industry constant when the distribution ran from 0% to 100%, the interquartile range was 18 to 41, and one company held 41,629 of the 59,160 mentions in the aggregate. Companies with fewer than 20 mentions averaged 42.5% against 27.0% for those over 100, so the flat line was an artifact of averaging a stable large-brand figure against unstable small-brand ratios.
The common cause was writing first and stress-testing only when challenged. We now require a research brief before any study and a content brief before any piece, both of which force the denominator, the expected range and every possible outcome to be written down before anyone looks at the data. This piece has one.
We have also quoted a drifted figure ourselves. Our attribution gap article from July 2026 carries the 91% Consensus Gap number with no mention of its European skew and the 75% shortlist figure with no mention of the 48 participants behind it. Both have had the scope added on 2026-08-19.
What changed as a result
The companion piece, where the AI search statistics you keep quoting actually came from, applies the same treatment to seven figures the industry repeats.