Glossary / Terms used in the essays
A glossary of
AI search.
Twenty terms that come up when discussing how content is found and cited by search engines and AI answer engines.
Each definition follows the usage of the operator or standards body behind the term, restated in plain English.
The definitions are used throughout the essays. For the retrieval terms, see how AI answer engines pick and cite sources.
- AI Mode
- A Google Search feature for questions that need exploration, reasoning or comparison. It returns an AI-generated response with links to supporting websites.
- AI Overview
- An AI-generated summary that Google Search shows for some queries, with links for further reading. Google says it appears only when its systems judge that it adds to classic Search.
- Answer engine
- A search product that responds to a question with a written answer and links to its sources, instead of only a list of results. Perplexity uses the term in its own help centre.
- Citation
- A link to a source, shown with an AI-generated answer. Perplexity numbers its citations. OpenAI's API returns each one as an annotation holding the URL, the title and the position in the text it supports.
- Generative AI impression
- In Google Search Console, a count of how many times links to a site were shown to a user in a generative AI feature on Google Search. AI Overviews and AI Mode are examples.
- Google-Extended
- A robots.txt token from Google with no crawler of its own. It controls whether content Google has crawled may be used to train future Gemini models and for grounding in Gemini Apps and Vertex AI. Google says it does not affect inclusion or ranking in Google Search.
- Grounding
- Supplying a model with retrieved content when a prompt is made, so that the answer is based on that content. Google describes it as a way to improve the factuality and relevance of responses.
- Grounding query
- In Bing Webmaster Tools, a key phrase the AI used when retrieving content that was then cited in an AI-generated answer. Bing shows a sample of these. It does not show the full set.
- Large language model (LLM)
- An AI language model with a very large number of parameters, trained on vast amounts of text. It can generate text, answer questions and summarise information.
- noindex
- A meta tag value asking search systems not to show a page. Google lists it as the way to keep content out of Search completely. OpenAI says it prevents a page's link and title being surfaced in ChatGPT, provided its crawler is allowed to fetch the page and read the tag.
- Preview controls
- Google's name for the nosnippet, data-nosnippet and max-snippet rules, used alongside noindex, that limit how much of a page may be shown in Search. Google says more restrictive settings also limit how content is featured in its AI experiences.
- Query fan-out
- Google's term for issuing several related searches across subtopics and data sources while an AI response is being built. Google says both AI Overviews and AI Mode may use it.
- Retrieval augmented generation (RAG)
- A technique that combines information retrieval with text generation. Relevant documents are fetched when the question is asked and passed to the model with it. The answer can then draw on information beyond the model's training data.
- Robots Exclusion Protocol
- The standard behind robots.txt, published as RFC 9309 in September 2022. Crawlers match their product token against user-agent groups in a file at /robots.txt and follow the most specific rule. The RFC states that the rules are not a form of access authorisation.
- Search crawler
- A crawler that builds the index an AI product searches when it answers with links. OAI-SearchBot, PerplexityBot and Claude-SearchBot are documented by their operators as doing this job.
- Search generative AI control
- A setting in Google Search Console that includes or excludes a site from AI Overviews, AI Mode and generative AI features in Discover. Google says it is not a ranking signal for other parts of Search and does not affect AI training.
- Structured data
- Machine-readable information about a page's content, commonly written with the schema.org vocabulary. Google says it should match what is visible on the page. It also says no special structured data is needed to appear in its AI features.
- Temperature
- A parameter that controls how much randomness a language model uses when generating text. Lower values give more predictable output. Anthropic notes that results are not fully deterministic even at a temperature of 0.
- Training crawler
- A crawler that collects web content which may be used to train AI models. GPTBot and ClaudeBot are documented by OpenAI and Anthropic as doing this job.
- User-triggered fetcher
- An agent that visits a page because a person asked a question or took an action that needed it. It is not part of automatic crawling. ChatGPT-User, Claude-User and Perplexity-User are examples. OpenAI and Perplexity say robots.txt rules may not apply to these fetches.
Sources
- Google Search Central, AI features and your website: developers.google.com/search/docs/appearance/ai-features
- Google Search Central Blog, Top ways to ensure your content performs well in Google's AI experiences on Search: developers.google.com/search/blog/2025/05/succeeding-in-ai-search
- Google, Google's common crawlers (Google-Extended): developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers
- Google Search Console Help, Search generative AI control: support.google.com/webmasters/answer/16908024
- Google Search Console Help, Generative AI performance report (Search): support.google.com/webmasters/answer/16984139
- Microsoft Bing Webmaster Blog, Introducing AI Performance in Bing Webmaster Tools public preview: blogs.bing.com/webmaster/2026/2/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview
- OpenAI, Overview of OpenAI crawlers: developers.openai.com/api/docs/bots
- OpenAI, Web search tool guide: developers.openai.com/api/docs/guides/tools-web-search
- OpenAI Help Center, Publishers and developers FAQ: help.openai.com/en/articles/12627856-publishers-and-developers-faq
- Perplexity Help Center, How does Perplexity work?: perplexity.ai/help-center/en/articles/10352895-how-does-perplexity-work
- Perplexity, Perplexity crawlers: docs.perplexity.ai/docs/resources/perplexity-crawlers
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?: support.claude.com/en/articles/8896518
- Anthropic, Glossary: platform.claude.com/docs/en/about-claude/glossary
- IETF, RFC 9309: Robots Exclusion Protocol: rfc-editor.org/rfc/rfc9309.html
Last updated .