The standard complaint about AI citation tools is that they can’t do what rank trackers do. A keyword ranking tells you a stable position you can act on; a citation tool tells you who an AI model named this time, and that answer can shift the next time you ask. Read that way, citation tracking looks like a downgrade — a vague brand-awareness signal standing in for a real diagnostic.
Duane Forrester, an AI search analyst who previously worked on Bing’s search team and now runs the AI-citation platform CitationIQ, has been making the case that this framing asks the wrong question entirely. It’s worth walking through, because the underlying data reshapes what “winning” a category even means once AI answers are involved — regardless of which tool anyone uses to measure it.
Wrong question: it’s not “why didn’t I appear,” it’s “has this settled”
The instinct to want rank-tracker-style diagnostics comes honestly from twenty years of SEO, where position was the outcome and the analytical job was figuring out what moved it. Forrester’s argument is that citation tools are answering a different question, and it’s the more commercially useful one: not why a brand did or didn’t appear in a given answer, but whether the query itself is still contested or has already settled on a small set of eligible brands.
If the answer has settled, and it settled on a competitor, no amount of reverse-engineering the “why” changes the outcome. The person asking the question doesn’t care whose answer it was — they wanted the answer, got it, and moved on.
What “settled” looks like in the data
Traditional SEO treated phrasing as expandable — many ways to ask the same thing, each one a countable, separate opportunity, with keyword volume built by aggregating those variations. AI answer systems treat that same phrasing as collapsible instead: they average across many variations and sources and return one synthesized answer the user accepts and acts on.
That convergence isn’t as tidy as it sounds, though. A June 2026 audit of 3,750 AI responses across three models and 250 category queries found all three models agreeing on the single top brand only 41.6% of the time. The more telling number sits underneath that: when the bar drops to majority agreement — at least two of three models naming the same top brand — agreement jumps to 91.6%.
That gap matters. It means the models are converging on which brands are eligible to be named, not on which one comes first. The eligible set is small and comparatively stable; the ordering inside it moves around. Someone rerunning the same prompt and getting a different top answer is usually seeing movement inside a fixed shortlist, not evidence that nothing has settled.
This isn’t quite a repeat of the old featured snippet debate, even though the two get compared often. Snippets collapsed the click but left the underlying phrase contested — one publisher held the box, and a competitor could see who held it and go take it, a dynamic Ahrefs documented at length when snippets were driving down click-through rates. Convergence works differently: the answer is synthesized from several sources at once, so there’s frequently nobody single party holding a position to take.
Does this survive personalization and small sample sizes?
The strongest pushback on convergence is that it might be a measurement artifact — clean test sessions, synthetic prompts, no real user history behind them. If every live user gets a genuinely personalized answer, the argument goes, convergence might only exist inside a controlled test environment.
A separate audit of 2,000 runs across ten buyer personas suggests personalization doesn’t dissolve the effect, but relocates it. Category leaders held roughly 80% consistency regardless of which persona the model believed was asking, while mid-market brands saw up to 75% of the recommendation set change as the persona shifted. The leaders stay put no matter who’s asking; the churn happens in the tier below them — which means personalization concentrates the convergence problem rather than solving it.
The synthetic-prompt objection is harder to dismiss cleanly. No one currently measures AI brand visibility against verified, full-scale real-world query distributions — a limitation that applies across every vendor in this category, CitationIQ included.
The map was always smaller than the list of possible phrasings
Google’s own documentation notes that AI Overviews and AI Mode may issue multiple related searches across subtopics and data sources before assembling a single response — meaning the phrase someone types is frequently not even the phrase the system actually searches on. That’s compression happening a layer earlier than most keyword analysis looks for it.
The harder claim is that the space of genuinely distinct commercial opportunities was always smaller than the space of possible phrasings — AI convergence didn’t shrink that space, it made it visible. Category-level attention concentration is a well-documented pattern from search infrastructure generally: entertainment, autos, and news have historically drawn disproportionate resources because that’s where aggregate demand sits, while niche categories mattered enormously to the people they served and got proportionally less server and editorial attention. That’s ordinary resource allocation in information retrieval, true well before language models entered the picture.
What’s new is that concentration now shapes answers directly, not just budgets. Researchers at Trine University and Texas A&M ran an experiment testing exactly this: they built matched product sets pitting one real brand against nine validated fictional competitors, with identical ratings, prices, review counts, and ingredient descriptions — the only difference was the name. The real brand was recommended in every single one of 670 valid trials, across three models, two languages, and four product categories. Not one fictional brand ever surfaced. The models weren’t evaluating the products on their merits; they were recognizing a name they’d already seen described repeatedly elsewhere.
That reframes what winning looks like: not being the best answer, but being the most-described entity in a category where description has already accumulated. The same June audit found genuine competitive vacuums — category queries with no dominant brand at all — in just 8% of the 250 queries tested, with healthcare technology showing the highest vacuum rate at 20%.
Where the argument is weakest
Forrester is careful to flag the limits of his own case, and they’re worth taking seriously rather than treating this as settled science. Convergence may be temporary — retrieval architectures change and model families diverge, and today’s stable consideration set could fragment again within a couple of years, with no reliable way to predict that shift in advance.
The bigger caveat is query type. The evidence is strongest for informational and category-level questions and considerably weaker for specific, constrained commercial queries — “what is X” convergence says very little about “best X for Y under constraint Z.” Notably, the 41.6% full-agreement figure came from commercial category queries specifically, which cuts directly against the strength of the argument in exactly the area where it matters most for buyers. And nearly all published measurement in this space, including two of the three studies cited above, comes from vendors selling AI visibility measurement — the same disclosed conflict of interest that applies to Forrester’s own platform, and a reason to hold every figure here a little loosely.
What this suggests for strategy
None of this points to a fully worked-out playbook yet, but a few directions look more promising than others.
- Aim to be the source models converge on, not one more source competing for a single phrase. This is the hardest option because it’s earned through years of independent, third-party description rather than produced on a content calendar — the kind of durable trust discussed in how keyword research strategy is already shifting to chase AI visibility rather than volume alone.
- Invest at the entity level, not the page level. If models are selecting on a recognized name rather than page-by-page content quality, the unit worth optimizing is the brand’s overall footprint, not any single URL — a shift that tracks closely with why traditional keyword systems are becoming less central to how visibility actually gets won.
- Look for categories where convergence hasn’t happened yet. Prior research has found that a model’s ability to answer accurately about a topic tracks how many relevant documents it saw during pretraining, and that scaling up models improves recall for popular topics while leaving sparse, under-documented ones roughly where they started. Thin-coverage categories are where the vacuums sit, and that’s a real but temporary opening — closely related to the fragmentation Google has itself acknowledged in how it handles keyword fragmentation and varied user needs in AI search.
- Abandon queries that are already settled rather than keep contesting them. It’s the least satisfying option on this list and arguably the most valuable, since the real cost of fighting for a settled phrase isn’t just the wasted budget — it’s the still-open phrase that budget could have gone toward instead.
The underlying claim is that convergence isn’t a measurement failure. It’s a measurement of something the industry hasn’t previously had a way to see: how much genuine room is actually left in a category. That’s an uncomfortable reframe for a discipline built partly on the promise that a patient operator could always eventually find their niche — but a smaller map that’s actually visible is still more useful than a larger one that was only ever imagined. Anyone weighing whether AI answers are quietly reshaping which brands even have a shot should also look at how Google’s own research is approaching whether AI answers are reliable in the first place, since answer quality and answer concentration are two sides of the same shift.
Frequently asked questions
What is “AI answer convergence”?
It’s the pattern where AI models, when asked variations of the same category question, tend to agree on a small, stable shortlist of eligible brands even though they may disagree on which one ranks first. Research puts full three-model agreement on the top brand around 42%, but agreement on at least a shared top-two rises to over 90%.
Does personalization eliminate this effect?
No — available research suggests it relocates the effect rather than removing it. Category leaders tend to stay consistent across different simulated user personas, while mid-market brands see much more churn depending on who the model believes is asking.
Are all commercial queries affected equally?
No. The evidence for convergence is strongest on informational and category-level questions and weaker on narrow, constrained commercial queries. That’s an important limit, since the underlying commercial-category data is also where full agreement was lowest.
How can a brand find categories that haven’t converged yet?
Thin, under-documented categories are where genuine competitive openings still exist, since model recall tracks how much relevant content existed during training. Category vacuums — queries with no dominant brand — showed up in roughly 8% of tested queries, concentrated more heavily in less-covered sectors like healthcare technology.