How AI search engines choose and cite sources

Guide, updated October 4, 2026

When an AI assistant answers with links, it's usually working from search results it gathered for that question. The model reads a set of retrieved pages, writes its answer, and credits some of those pages as sources. Which pages make that set, and which of them get the credit, depends on the engine, and the major providers have published far more about the first step than the second.

Three routes from a web page into an answer

A page can shape an answer in three ways. Training is the first, where text a model saw before release influences what it says without any link back. The second is a search index built ahead of time, like the ones OpenAI's OAI-SearchBot and Anthropic's Claude-SearchBot crawl for. The third is a live fetch made while someone waits, which is what ChatGPT-User and Perplexity-User do when a person's question calls for a specific page.

The citations people see come from the second and third routes. OpenAI's documentation for its web search tool lets developers choose between fetching live content and using only cached or indexed results, so an engine can cite a page without visiting it at the moment of the question. That's one reason server logs undercount how often a site gets used.

One question becomes several searches

Engines rarely search for the words a person typed. Google's guidance for site owners says AI Overviews and AI Mode may use a "query fan-out" technique, issuing multiple related searches across subtopics and data sources to develop a response. OpenAI's help center describes a similar pattern for ChatGPT, which typically rewrites a question into one or more targeted queries for its search partners and may send more specific follow-up queries after reviewing the first results.

Personal context shapes those searches as well. OpenAI says ChatGPT may use an approximate location based on the user's IP address, and relevant saved memories when memory is turned on, when it rewrites a query. Two people asking the same question can trigger different searches and see different sources, which is worth remembering before reading much into any single answer.

Where the candidate pages come from

For Google, the pool is its own index. Google describes its generative AI features as rooted in its core Search ranking and quality systems, and it says a page has to be indexed and eligible to show with a snippet in regular Search before it can appear as a supporting link in AI Overviews or AI Mode. Google also says AI Overviews are shown only when its systems judge them additive to classic Search, so on many queries they don't appear at all.

Other engines draw on a mix of sources. OpenAI says ChatGPT search sometimes partners with other search providers, and its crawler documentation says sites that opt out of OAI-SearchBot won't be shown in ChatGPT search answers, though they can still appear as navigational links. Perplexity says its own crawler, PerplexityBot, exists to surface and link websites in Perplexity's search results. Each engine works from a different slice of the web, so strong visibility in one is no guarantee of a place in another, and the crawler and agent access guide lists which bot feeds which product.

Retrieved, read, and cited are different outcomes

An engine typically reads more pages than it credits. OpenAI's web search documentation says the full list of sources a model consulted is often longer than its list of citations. Ahrefs reached a similar conclusion from the outside in an April 2026 study of 1.4 million ChatGPT prompts, which found that ChatGPT cited roughly half of the URLs it retrieved.

How that cut gets made is mostly undisclosed. Perplexity has described its process in some detail. In a September 2025 research post, it explained that it splits documents into self-contained spans that are retrieved and ranked individually, then uses more powerful reranker models to narrow the final set. Google says its systems can find the relevant passage on a long page and that site owners don't need to break content into small pieces for them. OpenAI hasn't published comparable detail about ChatGPT search.

What makes a page easier to use

Since most engines start from search, the basics of being found come first. A page needs to be reachable by the engine's search crawler, indexed, and readable as text in the HTML the server sends. After that, it has to answer something specific. A section that states one clear answer, with the facts, dates, and sources behind it, gives a passage-level ranker something to select, and a page that circles its topic gives it very little to quote. Our essay Ranking for machines makes the case that the opening paragraphs of a page now do more work than its meta description.

Evidence on specific tactics is thinner than the advice online suggests. A 2024 research paper on generative engine optimization (Aggarwal et al., presented at KDD 2024) found that adding quotations, statistics, and cited sources raised a page's share of a generated answer by roughly 30 to 40 percent on its main measure. The test used a 2023 model and a fixed set of five pages per question, though, so it shows what helps a page that's already been selected and says little about getting selected.

Schema markup is a similar case. Google says structured data isn't required for its generative AI features and that there's no special schema.org markup to add. When Ahrefs tracked 1,885 pages that were already being cited and then added JSON-LD, it found no meaningful change in their ChatGPT or AI Mode citations over the following 30 days (study published May 2026). Structured data still makes pages eligible for rich results in regular Google Search, which is reason enough to keep it accurate.

Why the same question returns different sources

AI answers vary from run to run. A 2026 statistical study by Ronald Sielinski, which collected responses through the OpenAI, Perplexity, and Gemini APIs, found that two responses to the same query cited an identical set of domains only 3 to 8 percent of the time on OpenAI's and Perplexity's search models, and almost never on Gemini. The same study estimated that measuring one domain's citation share on OpenAI's search model to within a five-point confidence interval takes 150 or more queries.

For a site owner, that means a single screenshot of an answer proves very little in either direction. Patterns across many questions and repeated runs carry the information, and the guide to measuring AI search visibility covers how to collect them and what the engines' own reports show.