← Signal

ChatGPT Asks the Exchange. Everyone Else Asks Reddit.

Four engines answering the same safety question. One reaches back to the exchange's own website, the other three reach for community and review sites.

Ask ChatGPT whether a crypto exchange is safe and it answers you, about half the time, out of that exchange’s own website. Ask Gemini, Perplexity or Google’s AI Overviews the same question and they almost never do.

That is the whole finding. Everything below is the evidence for it.

TL;DR

  • Across 13 exchanges, ChatGPT’s answer to “is X safe” drew 52.1% of its cited sources from X’s own domain.
  • Google AI Overviews drew 10.2%. Gemini 8.2%. Perplexity 8.2%.
  • We ran ChatGPT twice, seven days apart, identical model and design. 52.6% then 52.1%. The behaviour is stable, not a bad day.
  • The third-party half splits by kind, not just by volume. ChatGPT reaches for regulators. Every other engine reaches for Reddit and Trustpilot.
  • 1,300 engine calls in total. Every one answered.
SHARE OF CITED SOURCES THAT ARE THE BRAND'S OWN SITE ChatGPT 52.1% AI Overviews 10.2% Gemini 8.2% Perplexity 8.2%
13 exchanges, 4 safety prompts, 5 runs each, per engine. Denominator is every classified cited source across all answered runs.

Why this question

“Is this exchange safe” is the highest stakes thing a crypto buyer types. The answer decides whether money moves. It runs at roughly 9,600 searches a month in this category before you count the “is X legit” and “is X a scam” variants.

We expected the engines to answer it out of Reddit, Trustpilot and the review sites, and for the brand’s own trust and proof of reserves pages to be nearly absent. That is the thesis we have argued elsewhere: you do not own your own safety verdict, so no amount of on site trust content changes who the engine reaches for.

On three engines out of four, that held. On ChatGPT it did not, and we are publishing the miss.

What we did

Four prompts per exchange, held fixed: is X safe, is X legit, is X a scam, is X trustworthy. Thirteen exchanges. Five runs of every prompt on every engine, because one run is a coin flip and we have written about that mistake before.

That is 260 calls per engine and 1,300 in total. All 1,300 answered. Nothing was retried, nothing was dropped, and a run that produced no usable answer would have contributed no sources rather than scoring zero. None did.

Brand-own means the cited source is the exchange’s own registrable domain or a subdomain of it. That definition was fixed in the study brief before any data was collected, along with the rule that a per exchange share is only reported above 30 classified sources. Every exchange cleared that floor in every arm.

The replication

The first ChatGPT run was one day’s behaviour on one model. A 52.6% result that inverts your own thesis is exactly the result you should distrust, so we did not publish it. We re-ran the identical design against the identical model seven days later.

CHATGPT BRAND-OWN SHARE, SEVEN DAYS APART 52.6% 14 Aug 52.1% 21 Aug shift 0.5 points
Identical prompt set, identical model, identical run count. 260 calls each time.

Half a point. Per exchange the moves are mostly under four points. The largest single move was one exchange falling 11.8 points, which is worth knowing and is not large enough to move the headline.

The half that is not the brand

The remaining citations are where it gets more interesting, because the engines do not just cite third parties at different rates. They reach for different kinds of third party.

MOST CITED THIRD PARTY SOURCES ChatGPT justice.gov7.8% sec.gov7.8% trustpilot.com3.5% bbb.org2.5% cftc.gov1.9% REGULATORS AND ENFORCEMENT Every other engine reddit.com11.1% youtube.com8.2% trustpilot.com7.9% coinbureau.com6.0% coinledger.io5.3% COMMUNITY AND REVIEW
Right hand column shows Google AI Overviews, the highest of the three non-ChatGPT engines on each source. Gemini and Perplexity rank the same domains in the same order at slightly lower shares.

ChatGPT’s top two third party sources on a safety question are the Department of Justice and the SEC. Every other engine’s top source is Reddit.

That is a different theory of what “safe” means. One engine treats it as a regulatory question and goes to the enforcement record. The others treat it as a reputation question and go to what people said.

Every exchange, every engine

ExchangeChatGPT 14 AugChatGPT 21 AugGeminiPerplexityAI Overviews
Bitso75%77%12%12%22%
Luno79%67%8%8%12%
Coinbase63%63%9%10%8%
Bitvavo65%62%14%12%29%
Kraken62%59%8%11%4%
Crypto.com60%54%7%8%3%
OKX54%53%9%6%11%
Binance39%44%15%10%15%
Bybit47%43%4%3%6%
Upbit39%42%3%5%8%
KuCoin37%41%8%7%3%
Bitget36%37%6%5%8%
Bitstamp30%36%9%9%13%

Brand-own share, per exchange, per engine. Every cell is above the 30 source reporting floor.

Two things worth reading off it. The ChatGPT column is not uniform. Bitso sits at 77% and Bitstamp at 36%, so this is not a flat rule the engine applies to everyone, and something about the sites themselves moves it. The other three columns are uniform, and they are uniformly low. Whatever varies between these exchanges, it barely moves any engine other than ChatGPT.

What this changes

If you are optimising for ChatGPT, your own security, compliance and proof of reserves pages are load bearing. They are cited about half the time. That is the one place where the conventional advice, write good trust content on your own site, is measurably correct.

If you are optimising for anything else, those same pages are close to irrelevant on this question. Roughly nine in ten citations go somewhere you do not control, and the fix is not another page. It is the enforcement record, the review profile, and the community thread.

And if a vendor shows you one blended AI visibility number, this study is the reason to distrust it. A 52% and an 8% averaged together produce a figure that describes no engine anyone actually uses.

We were wrong about ChatGPT, on the query where being wrong matters most. The thesis survives on three engines out of four, which is not the same as surviving.

Method, limits and how to check us

  • Cohort: the 13 exchanges named in the table above, chosen for volume and held constant across every arm.
  • Prompt set: four prompts per exchange, fixed before collection and never changed between arms. See the fixed set method for why the set is frozen.
  • Runs: five per prompt per engine. 260 calls per arm, 1,300 in total, all answered.
  • Engines and models: ChatGPT gpt-5.6-terra, Gemini gemini-3.5-flash, Perplexity sonar-pro, and Google AI Overviews. Model names are recorded because a model change invalidates comparison.
  • Dates: Perplexity and AI Overviews 11 August 2026, Gemini and ChatGPT 14 August, ChatGPT replication 21 August. Arms collected on different days are compared with that caveat; the two ChatGPT arms are the only within engine comparison and they were designed for it.
  • Brand-own definition: the exchange’s own registrable domain or a subdomain of it, matched against an explicit domain list rather than by name. A substring rule would have counted bitsonline.com as Bitso citing itself.
  • Reporting floor: per exchange shares are reported only above 30 classified cited sources, fixed in the brief before collection. Every exchange cleared it in every arm.
  • Denominator: every classified cited source across all answered runs. A run with no usable answer contributes no sources and is excluded, never scored zero.
  • What this is not: a ranking of which exchange is safe. We measured which sources the engines cite, not whether any answer was correct.
  • Cost: $39.53 in engine calls across all five arms.
  • Raw data: the per engine brand-own results.