// AI answer monitoring · Answer Monitor

The AI visibility tool that tells you what AI says about you

Your buyers' real questions, run through six engines on a fixed prompt set, every week. Every answer checked against a fact sheet you sign.

// What it measures

Four numbers. Three of them exist today

Every measurement runs against a prompt set that is fixed on purpose. AI visibility tracking is only a benchmark if the questions hold still, because if the questions change, the figure cannot be compared.

Shipping

AI share of voice, per engine

How often your domain is the source an engine reaches for on your prompt set, and which domains take the slot when it is not you. Reported per engine, never blended, because the engines disagree with each other constantly.

Shipping

Brand mention share

How often your name appears in the answer text itself, and whether you are named ahead of the alternatives. Kept as its own number rather than averaged into AI share of voice.

Shipping

Ghost citation rate

Kevin Indig calls it a ghost citation: the engine links your page as a source and never says your name. In his study with Semrush, across 3,981 domain appearances, 61.7% of citations were exactly that, and only 13.2% got both the link and the name. Being cited and being named are different outcomes needing different work, so we report them as two numbers.

Shipping

Answer accuracy

Every claim the engines make about you, checked against a fact sheet you signed off on, with the wrong ones separated from the ones that appeared once and did not repeat.

In build

Drift alerts

A model update can change what an engine says about you overnight. Every scan is stored, so the baseline starts the day you run the first one. The automatic comparison and the alerting are being built, not shipped.

Every run goes through the provider APIs, logged out. That is deliberate: a signed-in answer carries the user's memory, their custom instructions and whatever the provider is testing on their account that week, so two people asking the same question get two different answers and neither one can be compared with itself next quarter. Logged-out API measurement is the only version of this number that means the same thing in October as it did in July. The full method, including how many times each prompt is run and why, is written up on how we measure.

// What a report row looks like

A sentence, and the page that taught it

Below is real output from our own scan of Uniswap, a public protocol we use as a test subject. It is not a client and no client data appears anywhere on this page.

zion@monitor ~ answer_accuracy

$ zion monitor --subject uniswap --engines perplexity,chatgpt,gemini --runs 2

  • fact sheet: uniswap.yaml v1, verified 2026-08-04
  • a real run from 2026-08-04, NOT the product spec: a live monitor runs six engines and at least five runs per question
  • 4 prompts x 2 runs x 3 engines = 24 runs, 16 answered
  • 143 claims checked: 70 match, 10 contradict, 63 not covered
gemini gemini-3.5-flash governance REPEATED 2 of 2

asked "Who is behind Uniswap?"

gemini stated "The Uniswap Foundation is an independent nonprofit organization dedicated to growing and securing the decentralized Uniswap ecosystem."

verdict Outdated rather than invented. True until December 2025, when most Foundation operations moved to Uniswap Labs.

cited alongside it

  • uniswapfoundation.org
perplexity sonar governance HELD FOR REVIEW

asked "Who is behind Uniswap?"

perplexity stated "The Uniswap Protocol is stewarded by the Uniswap Foundation."

verdict The fact sheet says UNI holders govern the protocol. Both sentences can be read as true at once, so this goes to a person before it goes in a report.

cited alongside it

  • about.uniswap.org
  • blog.uniswap.org

A domain listed here was cited by the engine in the same breath as the claim. That makes it the first page to check, not proof of causation. Where an engine ties no source to the sentence, the report says "not traceable" in those words rather than blaming whichever link happened to be listed first.

Read the first row

The engine is repeating something that was true until December 2025, when most Foundation operations moved to Uniswap Labs, and the page it cites was never corrected to say so. That is a content problem with an address, and it is fixable.

Read the second row

The automated pass flagged it. A person held it, because the fact sheet's sentence and the engine's sentence can both be true at once. Nothing reaches a client unreviewed, which is why we show you the queue instead of pretending it does not exist.

// Then what

A retainer fixes the wrong answers Answer Monitor finds

Nobody can edit a model, and anyone who says otherwise is selling something they cannot deliver. What can be done is what the second row above points at: find the pages the engine cited when it got you wrong, work out which ones you control, and rebuild the signal there. In our Uniswap scan, wrong claims traced back to about.uniswap.org and blog.uniswap.org, pages the brand owns and could edit the same afternoon.

That correction work is a service, and it is the one that turns a measurement into a change.

See how we fix it →

// Citations and mentions move independently

Most AI citations never say your name

Being cited means the answer used your page as a source. Being mentioned means the answer said your brand name. Buyers only ever see the second one. Most AI visibility platforms average the two into a single figure, which hides the entire gap between them.

62% of citations are "ghost citations": your page is the source, your name is never said
13% of appearances earn both a citation and a mention
−0.229 correlation between citations and mentions across 1,094 categories, so they move independently

Figures from Kevin Indig's Growth Memo, produced with Semrush, rounded above and exact here at 61.7% and 13.2%: 3,981 domain appearances across 115 prompts, 14 countries and 4 engines for the ghost citation split, and 1,094 US categories on ChatGPT for the correlation. Read them with the caveats. Semrush sells a tool that tracks both metrics, so the finding argues for its own product, and its co-published version is not an independent replication. The dataset records whether a brand appeared, not the sentiment around it, so a mention is not necessarily a good mention. We report the split because it holds across two datasets, and because the alternative is provably an average of two things that do not move together.

Answer Monitor FAQ

What is an AI visibility tool?

An AI visibility tool measures whether AI answer engines name your brand, cite your site, and describe you correctly when someone asks a question in your category. Classic rank trackers cannot see any of it, because there is no results page to read. A useful one reports three separate things: how often you are cited as a source, how often you are named in the answer text, and whether what is said about you is true.

What is the best AI visibility tool?

It depends on whether you want a number or an explanation. Most tools sample prompts, hand you a score and leave you to work out what caused it, and the prompt set often changes between runs, so the score moves for reasons that have nothing to do with you. What matters more than the vendor is the method: a prompt set that is frozen between measurements, several runs per prompt because engines are not deterministic, every engine reported separately rather than blended, and every wrong answer traced back to the page that taught it.

How is Answer Monitor different from an AI visibility dashboard?

A dashboard hands you the raw output of an automated pass and calls it insight. That output is a triage queue, not a finding list: in our own testing at least one flagged contradiction in eleven did not survive a human reading it, because two statements that sound opposed turned out to both be true. So a person reads every finding before you do, and the ones that do not hold never reach you. That review is most of what you are paying for, which is also why this is a service with software underneath rather than a seat licence.

Do you give our brand a trust score or an accuracy score?

No, and we will not build one. A single number derived from AI output has no defensible method behind it, and the moment it appears in a board deck somebody treats it as fact and makes a decision on it. What you get instead is auditable: what a named engine said, on a named date, in response to a named prompt, with the sources it cited alongside it. You can reproduce every line of that. Nobody can reproduce a score.

Can you show us whether our answers are getting better over time?

Not automatically yet. Every scan is stored, so the baseline starts the day you run the first one, and the raw material for a before and after comparison is already on disk. The scan to scan diff and the alerting on top of it are being built rather than shipped. Anyone selling you drift alerting today is selling a snapshot on a monthly invoice, and you will not find out which until month two.

Two ways in

Answer Monitor runs against your prompt set and your verified facts. Either the Diagnostic builds them properly, with every wrong answer traced to the page that caused it, or the monitor builds a smaller set of fifteen in its first month. Both are real. Neither is a shortcut: a monitor pointed at questions nobody vetted reports noise very precisely, which is why the fifteen are still built with you and still checked against a fact sheet you sign.

The Diagnostic: $7,500, six weeks. Roadmap at week three, corrections shipped and re-measured by week six. Monitor alone: $1,500 a month, month to month, no minimum.