AI search glossary

Guide, updated October 4, 2026

These are short definitions for the terms that come up most in AI search work, written to match the providers' own documentation as of October 2026. Where a term gets fuller treatment elsewhere on Modo AI, the definition links to that guide.

AEO and GEO
Answer engine optimization and generative engine optimization, two labels for work aimed at appearing in AI-generated answers. GEO is also the title of a 2024 research paper that tested page edits against a simulated answer engine. Google's July 2026 guidance says many tactics promoted under these labels aren't effective for Google Search.
AI crawler
A bot run by an AI company that fetches web pages. Most providers now separate training crawlers, search crawlers, and user-triggered fetchers, and each answers to its own robots.txt token, as the crawler and agent access guide lays out provider by provider.
AI Mode
Google's conversational search mode for longer and multi-part questions. It answers with an AI-written response and links to supporting pages, and Google says it may use query fan-out to gather them.
AI Overviews
AI-written summaries at the top of some Google results pages, with links to supporting pages. Google shows them only when it judges them additive to regular results, and a page must be indexed and eligible for a snippet to be linked from one.
Answer engine
Any search product that replies to a question with a written answer and supporting sources. ChatGPT search, Perplexity, Microsoft Copilot, and Google's AI Overviews and AI Mode are common examples.
Browser agent
An AI system that operates a web browser for a person, clicking, typing, and reading pages to complete a task. Google's web.dev guidance says these agents read pages through screenshots, raw HTML, and the accessibility tree.
Citation
A link or source credit attached to an AI answer. Engines usually cite fewer pages than they read, so being retrieved and being cited are separate outcomes, which the guide to how AI search engines choose and cite sources explains.
Google-Extended
A robots.txt product token, with no crawler of its own, that controls whether Google may use crawled content to train Gemini models and to ground answers in Gemini Apps and Vertex AI. Google says it doesn't affect inclusion or ranking in Google Search.
Grounding
Supplying a model with retrieved information at the moment it answers, so the response rests on current sources. Google uses the term as a synonym for retrieval-augmented generation.
llms.txt
A proposed convention, first published in September 2024, for a Markdown file at a site's root that summarizes the site for language models. Google says Google Search doesn't use it, and none of the crawler documentation we checked from OpenAI, Anthropic, Perplexity, or Apple says their bots do.
MCP (Model Context Protocol)
An open protocol, introduced by Anthropic in November 2024, for connecting AI applications to outside tools and data sources. Our essay on the protocol layer between agents and tools looks at why shared protocols matter for how services get found.
Prompt tracking
Sending a fixed set of questions to AI engines on a schedule and recording which brands and sources appear. Answers vary between runs, so results only mean much with repeated runs and a stated margin of error, as the measurement guide explains.
Query fan-out
Splitting one question into several related searches across subtopics and sources before writing an answer. Google documents it for AI Overviews and AI Mode, and OpenAI describes ChatGPT rewriting questions into targeted queries in a similar way.
RAG (retrieval-augmented generation)
A method where a system retrieves relevant documents and passes them to a model so its answer can draw on them. AI search features that look things up on the web work this way.
Reranker
A model that re-scores a set of retrieved results by how well each one fits the question, usually as a final narrowing step. Perplexity has described using cross-encoder rerankers in its search pipeline.
robots.txt
A plain-text file at a site's root, standardized as RFC 9309, that tells compliant crawlers which paths they may fetch. Each crawler follows the group that names its token, and the standard says the rules are not a form of access authorization.
Search generative AI control
A Search Console setting, available to all sites since August 31, 2026, that includes or excludes a site's links and content from AI Overviews, AI Mode, and generative features in Discover. Google says it doesn't affect AI training.
Structured data
Machine-readable markup, usually schema.org vocabulary in JSON-LD, that describes what a page contains. Google says its AI features need no special structured data, though structured data still supports rich results in regular Search.
User-triggered fetcher
A bot that visits a page because a person asked an assistant about it, such as ChatGPT-User, Claude-User, or Perplexity-User. OpenAI and Perplexity say robots.txt may not apply to these fetches, and Anthropic says site owners can use robots.txt to control Claude-User.
WebMCP
A proposed web standard that lets a site declare tools, defined in JavaScript or HTML, that browser agents can call directly. Chrome offers it for testing through an origin trial.