Most AI visibility dashboards report a number that feels like rank tracking and behaves nothing like it. They count how often a model names your brand across a batch of prompts, then present that count as progress. It rarely correlates with what actually pays: whether a model recommends you over a competitor when someone is ready to buy. Here is what separates the metrics worth building a report around from the ones that only look like they matter, drawn from a run of interviews with practitioners who test this for a living.
Prompt tracking borrows a shape that doesn’t fit
The standard AI visibility tool works by feeding a list of prompts into ChatGPT, Perplexity and Google’s AI Overviews, then reporting how often your brand appears. It resembles rank tracking closely enough to sell itself, which is exactly the problem. Technical SEO consultant Jono Alderson has described the approach bluntly: it is “copy-paste the current modality of rank tracking into a new thing. It doesn’t really fit, but it’s better than nothing.” His preferred framing is different — the job is to influence how the machine perceives your brand, not to chase mentions inside a list of guessed prompts.
That guessing problem is real. Most prompt lists are invented by a marketing team hoping to approximate what customers actually type, and an AI prompt is not a keyword — it is closer to a full sentence loaded with context a three-word query never carries. Grounding those prompts in real search data sounds like the fix, except AI traffic is actively corrupting that data. A well-documented bug once caused ChatGPT search queries to leak into Google Search Console as if they were organic searches, inflating impression counts for sites that happened to rank for those terms. Analytics consultant Jason Packer traced and published the mechanics, and outlets including Ars Technica covered the fallout, tying it to the “crocodile mouth” pattern where impressions spike while clicks fall.
The leak was a visible instance of something now constant and mostly invisible: AI systems query Google repeatedly to ground their own answers, fanning one user prompt into several parallel searches no human ever sees. Each counts as an impression on whatever page ranks for it. When your impressions climb and clicks stay flat, some of that climb is machines searching on someone’s behalf and never sending a visitor. Google’s own AI Overviews reporting inside Search Console confirms where this is headed, but it still shows impressions, the number AI is inflating, not the AI-driven clicks that would let you check it. That gap is exactly why treating generative AI data in Search Console at face value is risky.
A citation is not a recommendation
The single most important distinction in AI search measurement is that being cited and being recommended are different events, and most tools quietly conflate them. A citation is a model naming your page as a source underneath its answer. A recommendation is the model telling the user to pick you. Counting the first while implying the second is where prompt-tracking dashboards mislead.
The data backs this up from several directions. SEO analyst Lily Ray pulled AI Overview answers for 100 “best of” business-software queries across three 2026 checkpoints and found that when a brand’s own promotional listicle was cited as a source, that brand was still left out of the actual recommendation 69% of the time — 224 of 323 cited listicles. Google was reading the page, then recommending the competitors named inside it. Separately, Jeff Oxford’s team at Visibility Labs tested 20,000 ChatGPT responses and found recommendations shifted 80.2% once live search was switched on, with only a 0.4 correlation between being cited and being recommended. BrightEdge’s cross-engine analysis found source overlap between engine pairs ranging from 16% to 59%, while the set of actually recommended brands stayed tighter, 36% to 55%. Kevin Indig’s review of 3.7 million citations found 91% of cited URLs show up in only one engine — even your citation footprint rarely travels.
Alisa Scharf, Chief AI Officer at Seer Interactive, ranks citations below page-two Google visibility for exactly this reason: a citation is a leading indicator, not proof your brand was mentioned in the response text at all. She lays out the hierarchy plainly — first the citation, where your page is referenced; then the mention, where your brand name appears in the answer; and only rarely does a model say outright “you should go with X.” Malte Landwehr, who runs product at the AI search platform Peec AI, offers a sharper illustration: a now-defunct tool became one of the most-cited sources behind ChatGPT’s category answers without ever becoming a recommended brand — it gained influence over what got recommended without gaining any visibility of its own.
One measurement is noise, not a ranking
A single check of an AI answer tells you almost nothing, because the answer is different every time you ask. SparkToro founder Rand Fishkin ran the study that quantifies this instability: you are not getting a fixed answer when you prompt a model, you are sampling from thousands of possible answers, and the brand list, order and count can shift between identical prompts. His finding was stark — on average, you would need to ask Claude or ChatGPT roughly 1,500 times before two responses returned the same list of brands in the same order. That does not mean AI visibility is unmeasurable. It means it has to be treated like a poll, not a rank check: ask enough prompts, enough times, with enough variation, and you get a number with a real confidence interval, plus or minus 5%, tighter if you run more samples. The mistake most tools make is running it once and reporting the result as if it were stable.
What to track instead: presence tied to outcome
The metric that should replace raw citation counts is presence — how often you are named across the full answer space — read against whether that presence converts into an actual recommendation and a downstream action. Fishkin calls percentage-of-visibility the honest version of the number, closer to a brand-awareness survey (“have you heard of Nike”) than to a keyword rank. Wil Reynolds, founder of Seer Interactive, adds a detail almost nobody tracks: the composition of the answer over time, not just whether you appear in it. Skip that, and you would miss something like ChatGPT quietly doubling the length of its answers, which lifts your raw visibility number without your position actually improving. And being named is not revenue on its own — somebody still has to act on it. Track presence and conversion together, or the visibility number is decoration. That is the same discipline behind proving AI search is actually working with real tests rather than assuming a mention equals results.
Start with brand accuracy, not recommendation share
Before chasing recommendation share, check whether the model even describes your business correctly. If an AI system holds wrong facts about your company, every number built on top of that — citations, mentions, recommendations — is built on sand, because the model is recommending, or refusing to recommend, a version of you that does not exist.
Duane Forrester, who helped launch Schema.org and built Bing Webmaster Tools, frames the target as becoming the canonical source for your category rather than the top-ranked result: “Your goal should be to be seen as the canonical for whatever your question is. Not rankings, but that you are the source of knowledge.” His reasoning is practical — once a model has spent the computation to trust an answer and users respond well to it, there is little incentive for the system to look elsewhere.
Scharf turns this into something you can actually run: a brand accuracy audit. Build a list of objective, non-negotiable facts — when the company was founded, where it’s based, what it sells, who its real competitors are — and run those questions through each model on a schedule, scoring what it consistently gets right and wrong. That scorecard, not a flattering mention, is the real baseline.
Two blind spots you can’t fully close
Two limitations deserve naming rather than ignoring. The first is the training-data cutoff: a meaningful share of what a model tells a user comes from what it learned before live grounding kicks in, frozen at a date outside your control. You can do everything right today and still be optimizing against a snapshot of the model that is months stale.
The second is platform data access. Frontier model companies have little reason to expose how their systems decided what to recommend — there is no obvious version of OpenAI or Anthropic handing that over. The platforms more likely to share anything are the ones running both the model and a measurement surface with something to protect: Google has folded AI visibility into Search Console, and Microsoft does something similar through Bing Webmaster Tools. It is imperfect data, but more than the pure-play model companies currently offer.
Consistency is becoming the deterministic core of the work
Underneath the metrics debate is a simpler question: does the model know who you are, consistently, across every place it might learn that from? Schema, website copy, social profiles and third-party mentions all need to describe the same company, name and facts, without contradiction.
A legal development is pushing this from “nice to have” toward structural. A German court recently held Google liable for false statements its AI Overview generated about a business, reasoning that the AI-generated answer counts as Google’s own speech. If platforms are now legally exposed for what their AI says about a brand, it is reasonable to expect them to surface only the entities they are confident about, and leave out anything they are not sure of. That is an inference, not a confirmed mechanism, but it points somewhere useful: the most valuable thing to measure may not be how often you appear, but how certain the model is that it knows you.
Frequently asked questions
What’s the difference between an AI citation and an AI recommendation?
A citation is when a model lists your page as a source under its answer. A recommendation is when the model tells the user to choose your brand. Data across several independent studies shows citation and recommendation frequently do not overlap — being cited does not reliably predict being recommended.
Why do Search Console impressions rise while clicks stay flat?
A growing share of impressions comes from AI systems querying Google on a user’s behalf to ground their own answers, without a human ever clicking through. That inflates impressions independently of real demand.
What should teams measure first in AI search?
Brand accuracy. Confirm the model states basic facts about your company correctly — founding details, location, offerings, real competitors — before investing in recommendation-share tracking. An inaccurate entity picture undermines every metric built on top of it.