Making a site readable to AI crawlers and agents

Guide, updated October 4, 2026

Before an AI engine can cite a page, something has to fetch it and read it. That might be a search crawler, a training crawler, a fetcher acting for one person in real time, or a browser agent clicking through the site, and each follows different rules. This guide covers the access checks worth running first, based on the providers' own documentation as of October 2026.

Three kinds of AI bot

Most AI companies now separate their bots by job. Training crawlers collect pages that may be used to train future models. Search crawlers build the index an assistant searches when it answers a question. User-triggered fetchers visit a page because a person asked about it, and they don't crawl the web on their own. Each one answers to its own robots.txt token, so a site can allow search crawlers while blocking training crawlers, or the reverse.

ProviderTrainingSearch indexUser-triggered fetch
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
PerplexityNone listed (Perplexity says its bots aren't used for model training)PerplexityBotPerplexity-User
GoogleGoogle-Extended, a robots.txt token with no crawler of its ownGooglebotNot covered here
AppleApplebot-Extended, a robots.txt tokenApplebotNot covered here

Names and roles change, and each company keeps its current list on its own page (OpenAI, Anthropic, Perplexity, Google, and Apple). Check those pages before editing robots.txt, because a misspelled or retired token matches nothing and blocks nothing.

How robots.txt groups apply

robots.txt is standardized as RFC 9309, published in 2022. A crawler looks for the group of rules that names its product token and obeys that group, and it falls back to the User-agent: * group only when no group names it. A site that adds a User-agent: GPTBot group therefore has to repeat any general rules it also wants GPTBot to follow. The RFC also states that these rules "are not a form of access authorization." Compliant crawlers honor them, and nothing technical stops a crawler that doesn't.

A site that wants to stay available to AI search while keeping its pages out of model training could publish groups like the ones below. Applebot-Extended only governs training of Apple's foundation models, so Applebot keeps crawling for Spotlight, Siri, and Safari.

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: *
Allow: /

Changes aren't instant. OpenAI says its search systems can take about 24 hours to adjust after a robots.txt update, and Perplexity gives the same window for its bots.

User-triggered fetchers follow different rules

Because a person asked for the page, some providers treat these fetches more like a browser visit than a crawl. OpenAI's crawler documentation says robots.txt rules "may not apply" to ChatGPT-User and points site owners to OAI-SearchBot for search opt-outs. Perplexity says Perplexity-User "generally ignores robots.txt rules." Anthropic describes Claude-User differently, saying it lets site owners control which sites these user-initiated requests can reach, and that disabling it may reduce a site's visibility in user-directed web search.

Stopping a fetcher that doesn't follow robots.txt takes a server or CDN rule. OpenAI and Perplexity both publish the IP ranges their bots use, and Perplexity's documentation includes example firewall rules that match its user agents and IP ranges together. Before blocking, it's worth deciding whether those visits are unwanted at all, since each one usually means someone asked an assistant about the page, which makes them one of the few direct signals covered in the guide to measuring AI search visibility.

Google's controls work differently

Google doesn't run a separate AI crawler for Search. AI Overviews and AI Mode draw on pages Googlebot crawls, and Google's guidance says robots.txt rules for Googlebot are the control for how a site is crawled for Search as a whole. Google-Extended is only a robots.txt token with no user agent of its own. It governs whether Google may use crawled content to train future Gemini models and to ground answers in Gemini Apps and Vertex AI, and Google says it doesn't affect a site's inclusion or ranking in Google Search.

For AI features inside Search, Google offers two kinds of control. Snippet controls (nosnippet, data-nosnippet, max-snippet, and noindex) limit what can be shown from a page anywhere in Search. Since August 31, 2026, Search Console also has a site-level Search generative AI control that excludes a site's links and content from AI Overviews, AI Mode, and generative features in Discover. Google says an exclusion generally takes a few days to apply, that it doesn't affect AI training, and that Google-Extended is the control for limiting training.

Check the CDN as well as the file

Google's checklist for its AI features asks site owners to make sure crawling is allowed "in robots.txt, and by any CDN or hosting infrastructure." That second part catches more sites than it used to. Cloudflare, for example, lets customers block AI bots by behavior (search, agent, or training) and has changed its defaults for new domains more than once. Under the defaults it set on September 15, 2026, new domains block training and agent bots on pages that show ads while still allowing search bots. A site can end up blocking AI crawlers that its robots.txt welcomes, so the bot settings at the CDN or host deserve the same review as the file.

Do AI crawlers run JavaScript?

Many don't. A December 2024 analysis by Vercel and MERJ of crawler traffic across Vercel's network found that none of the major AI crawlers it measured rendered JavaScript, including OpenAI's three bots, ClaudeBot, and PerplexityBot. Googlebot renders pages, and so does Applebot, which uses a browser-based crawler. Crawlers change, so treat that study as a snapshot, but the safe assumption is that main content, titles, and navigation belong in the HTML the server sends. Fetching a page with curl, or reading its source in a browser, shows roughly what a non-rendering crawler receives.

Browser agents read pages more like people do

A newer kind of visitor drives a real browser on someone's behalf. Google's web.dev guide to agent-friendly websites, updated in April 2026, says these agents read a page through screenshots, raw HTML, and the browser's accessibility tree, and that modern agents combine all three. Its advice is ordinary good front-end practice, such as using real <button> and <a> elements, connecting form labels to their inputs, avoiding transparent overlays that cover controls, and keeping layouts stable. The same guide points to WebMCP, a proposed web standard for declaring actions to agents that Chrome offers for testing through an origin trial. Our earlier essay on the protocol layer between agents and tools looks at the server-side version of that idea.

Where llms.txt fits

llms.txt is a proposal for a Markdown file at a site's root that summarizes the site and links to clean versions of its key pages. Jeremy Howard first published it in September 2024 and revised it in August 2026. Some documentation sites have adopted it, and OpenAI and Perplexity both link one from their own developer docs, but neither company's crawler documentation says its search or training bots use the file. Google is explicit. Its July 2026 guidance says Google Search doesn't use llms.txt or similar files, and that creating one will neither help nor harm a site's visibility there.

Our view is that llms.txt is optional. It can help coding agents and other tools that look for it, and it costs little to maintain on a documentation site, but crawlable HTML is what the engines covered here read.