How AI answer engines pick and cite sources:
what is documented and what is not
What Google, OpenAI, Perplexity, Anthropic and Microsoft have documented about choosing and citing sources. And where the record stops.
Answer engines
Advice on how to get cited by ChatGPT or included in a Google AI Overview is easy to find. This piece asks a narrower question. What have the companies running these systems actually written down about how they choose and cite sources? It sets out the documented record. Then it marks where the documentation stops.
The shared pattern: search first, then write
Read side by side, the public documentation describes a similar sequence. The system turns a question into one or more search queries. It retrieves pages from a search index. Then it writes an answer that links to some of those pages. This is the idea behind retrieval augmented generation. Material is fetched at the moment of the question and handed to the model with it.
The outline is shared. The details are not. The details are where most of the practical consequences sit.
Google: AI Overviews and AI Mode
Google's guidance for site owners says both features may use a technique it calls query fan-out. That means issuing several related searches across subtopics and data sources while a response is being built. Google says its models identify further supporting pages as the response is generated. That is why the links shown can be wider ranging than those on a classic results page.
Eligibility is stated clearly. To appear as a supporting link, a page must be indexed and eligible to be shown in Google Search with a snippet. Google says there are no extra technical requirements. There is no special schema.org markup to add and no need for new machine-readable or AI text files. It also says that meeting every requirement guarantees nothing. Indexing and serving are never promised.
Two further details matter. Google says AI Overviews and AI Mode may use different models and techniques. So the links they show will vary. It also says AI Overviews only appear when its systems judge them to add something to classic Search. So they often do not appear at all.
Since June 2026 there has also been a clear switch. The Search generative AI control in Search Console lets a site owner include or exclude a site from AI Overviews, AI Mode and the generative AI features in Discover. Inclusion is the default. Google's help page says the control is not used as a ranking or inclusion signal for other parts of Search. It also says the control has been available to all websites worldwide since 31 August 2026.
OpenAI: ChatGPT search
OpenAI's help page on web search says ChatGPT may search automatically when a question would benefit from current information. It explains that ChatGPT search sometimes works with other search providers. When it does, it typically rewrites the user's prompt into one or more targeted queries. After reviewing the first results it may send further, more specific queries. If memory is switched on, saved memories can shape the rewritten query. So can the user's approximate location.
On ranking, the same page says only that results are ranked using multiple factors intended to help people find relevant, reliable information. It adds that placement is not guaranteed. The factors are not listed. The eligibility condition is concrete, though. Allow OAI-SearchBot to crawl the site. Make sure the host or content delivery network accepts traffic from OpenAI's published IP addresses.
OpenAI's crawler documentation adds that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. They can still appear as navigational links. Its publisher FAQ goes a step further. OpenAI may learn the URL of a disallowed page from a third-party search provider or from crawling other pages. If so, it may still show the link and page title in ChatGPT Atlas. A noindex tag prevents that. It only works if the crawler is allowed to fetch the page and read the tag.
For developers, the web search tool documentation shows what a citation is in mechanical terms. It is an annotation carrying the URL, the title and the position in the text it supports. OpenAI requires those citations to be clearly visible and clickable when results are shown to end users. The same page describes three kinds of search. A quick lookup has no planning. An agentic search lets the model decide whether to keep searching. Deep research can draw on hundreds of sources.
Perplexity
Perplexity's help centre describes three steps. It interprets the question, searches the internet in real time and then compiles the most relevant findings into an answer. Each answer carries numbered citations that link to the original sources.
Its crawler documentation says PerplexityBot exists to surface and link websites in Perplexity's search results. It is not used to crawl content for AI foundation models. Any site that wants to appear should allow it in robots.txt. Beyond general phrases such as "authoritative sources", I found no published description of how Perplexity ranks one page above another.
Anthropic and Microsoft
Anthropic documents three bots. Two matter here. Claude-SearchBot indexes content to improve the quality of search results for Claude users. Claude-User retrieves pages when a user's question calls for it. Anthropic says blocking either may reduce a site's visibility in those results.
Of the documents reviewed here, Microsoft's gives the most direct view of the retrieval step. The AI Performance report in Bing Webmaster Tools was announced as a public preview in February 2026. It counts citations across Microsoft Copilot, AI-generated summaries in Bing and selected partner integrations. It also shows a sample of grounding queries. These are the key phrases the AI used when retrieving content that was then cited. Microsoft is clear that citation counts do not indicate ranking, authority or placement within an answer.
What remains unknown
Set against that record, the gaps are easy to list.
- The ranking factors and their weights. None of the documents above lists them. OpenAI's "multiple factors" is as specific as the published material gets.
- The rewritten queries. Google and OpenAI describe the rewriting step but do not show site owners the queries. Bing shows a sample.
- Why one retrieved page is cited and another is not. The documents describe retrieval and citation. None explains the choice between pages that were both retrieved.
- How stable any answer is. Google says links vary between its two features. OpenAI says memory and location can change the queries sent. Neither publishes a measure of how much answers differ between users or between days.
- The balance between training data and live retrieval. None of the pages I reviewed explains how much of a given answer comes from what the model already learned. They do not say how much comes from what it just fetched.
What I take from this
This section is opinion. The documented requirements are basic. Be crawlable by the right bot. Be indexed. Be eligible for a snippet. Keep the important content in text. Do not block yourself at the firewall. Everything beyond that is inference. Some of it is reasonable and some of it is marketing.
My view is that any claim to know the citation formula deserves caution unless it links to an operator's own documentation. A more useful order of work is to confirm access first. Then look at what the operators will actually report back. The glossary defines the terms used here.
Sources
- Google Search Central, AI features and your website: developers.google.com/search/docs/appearance/ai-features
- Google Search Console Help, Search generative AI control: support.google.com/webmasters/answer/16908024
- OpenAI Help Center, Searching the web with ChatGPT: help.openai.com/en/articles/9237897-searching-the-web-with-chatgpt
- OpenAI, Overview of OpenAI crawlers: developers.openai.com/api/docs/bots
- OpenAI Help Center, Publishers and developers FAQ: help.openai.com/en/articles/12627856-publishers-and-developers-faq
- OpenAI, Web search tool guide: developers.openai.com/api/docs/guides/tools-web-search
- Perplexity Help Center, How does Perplexity work?: perplexity.ai/help-center/en/articles/10352895-how-does-perplexity-work
- Perplexity, Perplexity crawlers: docs.perplexity.ai/docs/resources/perplexity-crawlers
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?: support.claude.com/en/articles/8896518
- Microsoft Bing Webmaster Blog, Introducing AI Performance in Bing Webmaster Tools public preview: blogs.bing.com/webmaster/2026/2/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview
- Anthropic, Glossary (retrieval augmented generation): platform.claude.com/docs/en/about-claude/glossary
This essay was researched and drafted with AI assistance, then reviewed and edited by me before publication. The editorial policy explains the process and how corrections are handled. If you think something on this page is wrong, please tell me.