Ranking first on Google no longer guarantees a citation in ChatGPT, Copilot, or an AI Overview. That is because AI systems don’t rank pages the way search engines do — they pull out fragments of text from many pages and stitch them into one answer. If your content isn’t built to be lifted in pieces, it can rank perfectly and still never get quoted.
This is the practical follow-up to our look at why websites now need to speak to machines, not just search engines. Here we get into the mechanics: how AI selects what to quote, what the research says actually earns a citation, and the technical and editorial changes that follow from it.
Pages don’t get chosen. Fragments do.
Microsoft’s Krishna Madhavan, principal product manager on the Bing team, explained the process plainly in late 2025: AI assistants break content into smaller, structured pieces — a process called parsing — evaluate each piece for authority and relevance, then assemble the ones that pass into a single answer, often pulling from several sources at once.
That single idea explains most of what looks confusing about AI search. Your page can hold the No. 1 spot on Google and still be skipped by an AI Overview, because the system isn’t asking “which page is best” — it’s asking “which paragraph answers this exactly.” A page written as one long, contextual narrative gives it nothing clean to extract. Related reading: AEO playbook: how to get cited in AI answers.
And the volume involved is no longer trivial. Conductor’s AEO/GEO Benchmarks Report, covering 13,770 domains and 17 million AI responses through January 2026, puts AI traffic at 1.08% of all website sessions and climbing roughly 1% a month. Microsoft separately reported AI referrals to top sites jumped 357% year-over-year by mid-2025, to 1.13 billion visits. One in four Google searches now triggers an AI Overview — nearly one in two in healthcare. Small shares today, growing fast, and the content filling those answers has to be sourced from somewhere.
What the research actually says gets cited
The academic groundwork here is the “GEO: Generative Engine Optimization” paper out of Princeton, IIT Delhi, and Georgia Tech, presented at KDD 2024. It tested nine optimization tactics and found the strongest one by far was citing credible sources — a 115% visibility lift for sites that weren’t already ranking near the top. A less intuitive finding from the same study: writing in an authoritative or persuasive tone did nothing for visibility. These systems don’t respond to rhetoric, only to verifiable claims.
2025 brought several follow-ups tested on live production engines instead of simulations, and they largely converge:
- University of Toronto (September 2025) — the first large-scale study across ChatGPT, Perplexity, Gemini, and Claude. Its headline finding: AI search leans heavily on earned media. In consumer electronics, third-party sources were cited 92.1% of the time versus 54.1% for Google; automotive showed a similar gap, 81.9% versus 45.1%. Whose domain the information sits on matters as much as how it’s written.
- Carnegie Mellon’s AutoGEO (October 2025) — used automated testing to find what generative engines reward, and saw up to a 51% improvement over baseline content. The consistent preferences across engines: full topic coverage, cited facts, and clear logical structure with headings and lists.
- The GEO-16 framework (September 2025) — analyzed 1,702 real citations pulled from Brave, Google AI Overviews, and Perplexity, and isolated 16 on-page factors that predict citation odds. The top three were metadata and freshness, semantic HTML, and structured data — technical factors that matter as much as the writing itself.
- Columbia/MIT ecommerce study (November 2025) — tested 15 common content-rewriting tricks and found 10 did nothing or actively hurt. What worked converged on truthfulness, matching user intent, and genuine competitive differentiation. Not formatting hacks — substance.
Read together, the pattern is consistent: AI systems reward clarity, verifiable facts, and structure. They ignore marketing language and keyword stuffing almost entirely.
Formatting content so it can be lifted whole
Given that AI extracts fragments rather than reading pages top to bottom, a few structural habits make the difference between getting quoted and getting skipped:
- Descriptive headings, not vague ones. “How AI parses content differently than search engines” tells a model exactly what’s in the section. “Overview” or “Learn More” tells it nothing.
- Question-and-answer pairs. Microsoft notes that assistants can often lift a well-formed Q&A pair word for word into a generated response. If a heading mirrors a real question and the paragraph beneath answers it directly, you’ve done half the model’s job for it.
- Front-loaded answers. Lead each section with the actual answer, then add context. A section that opens with two paragraphs of background before stating the number, the step, or the fact will lose the citation to a competitor who states it first.
- Self-contained sections. Each section should stand alone without requiring the reader to have absorbed the one before it — because the model may only ever extract that one section.
- Visible content only. Microsoft’s own guidance is explicit: don’t hide important answers in tabs or expandable menus, because AI systems may never render them. If a fact matters, it needs to live in plain HTML, not behind a click.
Authority now means being cited elsewhere, not just writing well
E-E-A-T — experience, expertise, authoritativeness, trustworthiness — didn’t disappear with the shift to AI search; it just stopped being a Google-only concept. Microsoft’s guidance describes the same baseline: content needs to be fresh, structured, and semantically clear, with claims anchored in measurable facts rather than words like “innovative” or “eco-friendly” that carry no verifiable meaning.
The Toronto research adds an uncomfortable wrinkle for anyone optimizing only their own website: AI systems weight third-party validation over self-published claims. Press coverage, independent reviews, and mentions on established industry sites carry more weight toward a citation than perfecting your own product page’s copy. Getting quoted by others is now part of the optimization work, not a side benefit of it.
Freshness plays a similar role. Madhavan put it bluntly at Pubcon Cyber Week: stale or missing content constrains how much a system can retrieve from a site, and pushes it toward alternative sources instead.
Schema markup and crawler access are the technical floor
None of the editorial work matters if machines can’t reach or parse the page in the first place. Structured data is one part of that: Microsoft describes schema as what “turns plain text into structured data that machines can interpret with confidence,” and the GEO-16 framework independently confirmed structured data as one of the top three predictors of citation likelihood. FAQPage, HowTo, Product, and Article/BlogPosting schema each give a model explicit context instead of forcing it to infer what a page is about. Worth reading in full: our breakdown of where schema actually moves the needle in GEO, and where it doesn’t.
Crawler access is the other half. Most AI platforms now separate search crawling from training crawling, and the two deserve different robots.txt treatment. OpenAI’s split is the cleanest example: you can allow OAI-SearchBot so your content surfaces in ChatGPT search, while blocking GPTBot so it isn’t used for training. Google’s controls are blunter — blocking Google-Extended stops Gemini training but does nothing to AI Overviews, which run on the standard Googlebot. And bot compliance isn’t universal by default; Cloudflare documented Perplexity using undeclared crawlers with rotating IPs to get around no-crawl rules, and separately, OpenAI has said robots.txt rules may not apply to ChatGPT’s fetch bot in every case. Assume access needs verifying, not just declaring.
Measuring whether any of this is working
Ahrefs analyzed 1.9 million citations across 1 million AI Overviews and found 76% came from pages already ranking in Google’s top 10, with a median position of 2 for the most-cited URLs — so classic ranking still feeds AI citation, even though ranking No. 1 alone is, in their words, close to a coin flip for actually getting quoted.
The traffic trade-off is real: Ahrefs measured a 58% lower click-through rate for the No. 1 organic position when an AI Overview is present, and Seer Interactive found a 61% drop in organic CTR on AI Overview queries generally. But being cited inside the Overview itself recovers 35% more clicks compared to not being cited — which is the real argument for chasing citations rather than mourning lost sessions. For a fuller view of that shift, see how to tell if AI Overviews are taking your clicks, and what to do about it.
On tooling, Bing Webmaster Tools is the easiest starting point — it’s free, and its AI Performance Report shows Copilot citation activity directly. For ChatGPT, track utm_source=chatgpt.com in your analytics, since OpenAI appends it automatically to referral traffic. That platform is worth prioritizing specifically: Conductor’s January 2026 data found 87.4% of all AI referral traffic currently comes from ChatGPT alone.
Frequently asked questions
What is answer engine optimization?
Answer engine optimization is the practice of structuring and writing content so AI systems — ChatGPT, Copilot, Perplexity, Google’s AI Overviews — can extract and cite it directly in generated answers, rather than optimizing purely for a ranked position on a results page. More on that in AI answer convergence.
Does traditional SEO still matter for AI citations?
Yes. The majority of AI citations still come from pages already ranking well in Google, so rankings remain a prerequisite. But ranking well doesn’t guarantee a citation — the content also has to be structured in a way the model can extract cleanly.
Does schema markup guarantee more AI citations?
No, but it consistently ranks among the strongest predictive factors in independent research, alongside metadata freshness and semantic HTML. Treat it as part of the technical floor, not a silver bullet on its own.
Should I block AI crawlers from my site?
Most sites benefit from allowing search-focused crawlers (like OAI-SearchBot) while optionally blocking training-focused ones (like GPTBot), since blocking search crawlers removes any chance of being cited at all. See also: the trust signals that get your brand cited by AI.
The takeaway
Traditional SEO asks how to rank. Answer engine optimization asks how to become the fragment a model is confident enough to quote. That comes down to structure a machine can parse, facts it can verify, and a presence beyond your own site that corroborates what you’re claiming — not a single trick, but a shift in how content gets built from the first draft onward.
1 Comment
Accessibility Tree: How AI Agents See Your Website
August 25, 2026[…] site is what makes content easier to cite in AI answers, a topic covered in depth in our guide to getting your content into AI responses. It also connects directly to how machine-first architecture is reshaping site design, and to why […]