Conductor fired 14,000 API calls at four AI engines, 50 runs per cell, the same prompt strings every time. On purchase-intent questions, the brand overlap between any two runs came out at 40%. Their sentence, not mine: "Out of every 10 brands that appeared across any two runs, only four showed up in both."
Conductor sells an AEO platform. So that is an interested party publishing a number that makes its own category look shaky, which is exactly why I would take it seriously.
Now hold it next to the thing on your dashboard. One number between 0 and 100, refreshed weekly, labeled AI visibility, produced by a tool you cannot see inside. When I published the manual method for measuring share of AI voice on this site, I wrote a three-runs-minimum rule into it, because the answer moves between runs and one check tells you close to nothing. That is the gap this post is about: what you can actually audit when someone senior asks you to defend that number out loud.
The score looks precise. It isn't.
Semrush's 2026 AI Visibility Index analyzed 126 million US AI search prompts from January through April 2026, across ChatGPT, Gemini, Google AI Mode and AI Overviews. Sitting inside it is a worked example of the format I am arguing against: Patagonia "maintained an AI visibility score of approximately 79-80 throughout the study period."
Read the 79 as a reader rather than as a marketer. What is in it? It blends how often the brand got named with how often its domain got used as a source, over a prompt set Semrush chose, at run counts Semrush chose, on the models Semrush sampled, combined by a weighting Semrush has not published. Nobody outside that company can take the 79 apart, reweight it, or rebuild it with a different tool. Semrush sells the product that produces the score. That is the normal state of this market, not a scandal.
The demand behind it is real, and that is the uncomfortable part. Semrush's own survey found 45% of marketing leaders cannot accurately measure their brand's visibility inside AI-generated answers, and only 9% have tools that track every relevant metric across platforms. That is a room full of people who need a number by Thursday. A tidy 0 to 100 walks straight into the gap. Two significant figures read like measurement precision. Here they are presentation precision, sitting on top of a sampling process nobody outside the vendor gets to inspect.
What's the difference between an AI citation and an AI mention?
A mention is your brand named in the text of an AI answer, link or no link. A citation is a domain the engine used as a source for what it said, usually shown as a numbered reference or a link chip. Two different events, with different denominators, and they do not move together.
Semrush puts the distinction plainly in the same index: "Mentions show how often a company appears in an answer, while citations show which domains and pages AI platforms use as evidence." Their data also sizes the gap. "On Gemini, the overlap between mentioned brands and cited domains can be as low as 30%."
Engine behavior widens it further. Semrush measured ChatGPT citing an average of 15 sources per response against Gemini's average of 3. Same brand, same week, five times as many slots in one engine's bibliography as the other, for reasons that have nothing to do with the brand. Any composite that folds mentions and citations into a single figure has to pick a weighting between two events that behave differently on every surface. That choice is where the score gets made, and it is the part you never see.
Mention and citation are two separate systems, not one funnel
Seer Interactive, an agency that sells GEO services, ran the largest test of this I can find: 541,213 LLM responses across 20 brands, published 24 March 2026 by John Lovett. The finding that matters: when a brand is mentioned in a response, its citation rate is 53.1%. When that same brand is not mentioned, its citation rate is 10.6%.
Five times, running the wrong direction for anyone who assumes a citation drags a mention along behind it. Seer's line is better than mine: "The citations are the bibliography, not the brainstorm."
Citation rate for the same brand, mentioned vs not mentioned
Seer calls the gap a ghost citation, and describes it like this: "your content cleared the retrieval threshold. Your brand did not clear the mention threshold. Those are two separate systems." Your page can be good enough to source while your brand is not known enough to name. A blended score averages those two facts into one flat number that tells you to do nothing in particular.
Two limits on that data, both of which Seer states themselves. They do not have access to LLM token generation logs, so this is strongly supported behavioral evidence and not proven architecture. And the analysis leaves out Claude and Meta entirely, because neither platform returns citation URLs. Whatever happens on those two surfaces is not inside the 541,213.
The citation event is real. It's also not what you think it captures.
The unit I am arguing for is not a metaphor. It is a structured object the engines already hand back.
OpenAI's
Responses API,
with the web search tool on, returns a url_citation annotation carrying the URL,
the title, and the character span in the answer where that source was used.
Perplexity's Sonar API
returns a citations array of the URLs used to generate the response, plus
search_results.
Gemini's grounding
returns groundingMetadata: the queries it issued, groundingChunks
with uri, title and domain, and groundingSupports that map spans of the answer
back to specific sources.
That is an event with a date, an engine, a prompt, and a source URL attached to it. You can log it, diff it week over week, and hand the raw rows to anyone who wants to check your work. A 79 does none of that.
Here is the scope limit, and it is load-bearing. Those payloads describe your own API call. They are not a feed of what a person saw inside the ChatGPT app or the Gemini app on their phone, and no such feed exists, at any price, for anyone. You are sampling the engine, not watching your buyers.
Why two vendor scores can't agree, even in good faith
I went looking for a clean head-to-head, two named tools scoring the same brand in the same week, and there is nothing worth citing. What exists is vendor comparison pages written by the vendors, plus one unattributed anecdote. So I am not going to imply a shootout that nobody has published.
The argument does not need one. It comes from method.
Two AI visibility tools compute their score from different prompt sets, different run counts per prompt, different sampled models and model versions, different regions and languages. Those are the inputs. Now put the arXiv work next to that. Will Jack, Noah Lehman, Keller Maloney and Sarah Xu ran roughly 6,000 paraphrase runs and roughly 6,000 same-prompt controls across OpenAI and Anthropic models (arXiv 2605.27440), and their abstract states it flatly: "The prompt string, not the underlying buyer intent, is the dominant input to which brands surface."
If the exact wording is the dominant input, and two vendors are not issuing the same wordings, their two scores were never measurements of the same quantity. Not close measurements of one quantity. Different quantities. Both tools can be internally consistent, both can be built honestly, and the numbers still cannot be compared or averaged. The same paper again: counting brand mentions over a fixed set of prompts "produces a metric whose dominant source of variance is which paraphrase the tracker happens to issue, not the model's behavior toward the brand."
The honest limit: a single logged event is close to noise
Now the section that argues against my own recommendation, because leaving it out would make this post the thing it is complaining about.
Throw away the vendor score, start logging raw citation and mention events, and one event on its own is still worth close to nothing.
Dmitrij Zatuchin decomposed where the noise actually comes from: 12,933 LLM responses, 20 Central and Eastern European brands, 8 languages, 3 models (arXiv 2607.13304, 14 July 2026). Query language accounts for 26.5% of the variance in a single response. Brand identity accounts for 1.5%, an ICC of 0.0146. His words: "a single AI answer carries almost no brand-discriminating signal." Brand-ranking reliability from one answer sits near 0.01.
Variance of one AI answer: the language you asked in vs the brand
The reader-level version of that comes from SparkToro. Rand Fishkin, working with Patrick O'Donnell, had 600 volunteers run 12 prompts through 3 tools a combined 2,961 times (ChatGPT, Claude, and Google's AI Overview and AI Mode), published 28 January 2026. The finding: "there's a <1 in 100 chance that ChatGPT or Google's AI, if asked 100X, will give you the same list of brands in any two responses." On ordering it gets worse: "it's more like 1 in 1,000 runs before you'd see two lists in the same order." SparkToro sells an audience research tool, so they sit adjacent to this market too.
Rerunning the same prompt harder does not rescue any of this. The rerun baseline in Jack et al., identical prompt run twice, is a Jaccard overlap of 0.50 to 0.61. That is the ceiling before anybody rewords anything. Change "best CRM" to "top CRM", same buying intent, and it drops to 0.288. Add a constraint, "best CRM for a SaaS startup", and it drops to 0.135. Turning up the model's reasoning effort moved the gap by no more than 0.05 in either direction.
Piling on repeats stops paying almost immediately. In Zatuchin's decomposition, a repeat past the fifth reduces relative-error variance by 0.0003. You are spending API budget to buy a rounding error.
And the scope limits from earlier apply to your own log, not just to somebody else's dashboard. Seer's citation rates exclude Claude and Meta. The citation payloads you can pull describe the calls you made, never what a buyer saw in a consumer app.
Build the log, not the score
The fix is a sampling design, not a bigger single number.
Zatuchin's practical lesson is the one to take: per unit of query budget, adding languages and models cuts relative-error variance far more than adding repeats. Spread the same 500 calls across more phrasings and more surfaces instead of hammering one prompt fifty times. Do that and brand-ranking reliability climbs from near 0.01 on a single answer to about 0.36 at the full crossed design. 0.36 is not precision. It is roughly the distance between a coin flip and a signal, and claiming more than that would repeat the mistake I am objecting to.
So log the event. Every row carries the engine, the exact prompt string you sent, the language, the model version, the date, and whether the brand was mentioned or the domain was cited. Report the distribution across those rows. Do not blend them into one headline figure, because the blending is the step where the auditability dies.
Share of AI voice is the ratio. This is what has to be true about the numerator and the denominator before you trust the ratio. Nothing here introduces a competing metric. It raises the bar on what counts as a citation or a mention inside the share of AI voice number this site already told you to track.
Worth keeping two questions apart while you do it. What predicts AI visibility is a different axis from how you measure the visibility you already have, and the predictor side, branded mentions against backlinks, lives in its own post.
Even the platform of record does not hand you an event log yet: Google Search Console's generative AI performance report shows impressions only, with Pages, Countries, Devices and Dates as its dimensions, no query, no clicks, and Google's own help doc says it is still rolling out "to a subset of website owners."
One thing to sort out before any of the logging is worth doing. A page has to be built to earn a mention or a citation in the first place, and that is a content problem rather than a measurement one. The SEO + AEO Content Rater runs 24 checks in your browser and shows you which ones a page passed or failed and why. It is the structural opposite of a black box: every check is readable. It does not call any AI engine and it does not track citations or mentions, so it is not the measurement tool this post is asking for. It is the step upstream, where you find out your answer is buried on line 40.
Five questions people ask before they drop the score
What is an AI visibility score?
An AI visibility score is a vendor's proprietary 0 to 100 composite meant to summarize how often a brand appears and gets cited across AI answers. It is not a standardized metric. Each vendor defines and weights it differently, on its own prompt set, run count, and sampled models, so a score from one tool cannot be compared to a score from another.
What's the difference between an AI citation and an AI mention?
A mention is your brand named in the text of an AI answer, with or without a link. A citation is a linked or sourced domain the engine used as evidence for what it said. They are different events with different denominators. Semrush reports that on Gemini the overlap between mentioned brands and cited domains can be as low as 30%.
Can I compare AI visibility scores between two different tools?
No, not directly. Each vendor's score is built on a different prompt set, run count, sampled models, and regions. The paper by Jack, Lehman, Maloney and Xu (arXiv 2605.27440) found that the prompt string, not the underlying buyer intent, is the dominant input to which brands surface. Different prompt strings produce different, non-comparable outputs, even when both tools are built in good faith.
How many times should I run the same prompt to track AI visibility?
Repeating one prompt runs out of value fast. In Zatuchin's variance decomposition (arXiv 2607.13304, 12,933 responses), a repeat past the fifth reduces relative-error variance by only 0.0003. Spreading the same query budget across different phrasings, languages, and models reduces variance far more than adding repeats of a single prompt.
Does Google Search Console show AI Overview citations?
Not yet in a usable form. The generative AI performance report in Search Console shows impressions only, with Pages, Countries, Devices and Dates as its dimensions. There is no query dimension and no clicks, and Google states it is still rolling the report out to a subset of website owners.
The citation event was never the clean number. It is the auditable one. It names the engine, the prompt string, and the day, and it tells you exactly where its own noise comes from, which is how you design a sample that survives the noise. A 79 tells you 79. When it slides to 74, you cannot find out why, because the thing that changed might be your brand, or might be the paraphrase the vendor rotated in that week, and nothing in the product tells those apart.
So next time a dashboard hands you one AI visibility number, which do you ask for first: the prompt set, the run count, the models sampled, or the languages?
, Amit