AEO.wikiThe AEO Reference

AI crawlers and robots.txt

From AEO.wiki, the Answer Engine Optimization reference · Last reviewed · By · Published by Local Blitz · How we research

Most AI companies now run separate crawlers for model training, search indexing and user-triggered fetches, and document robots.txt tokens for each. To appear in AI answers while opting out of training, allow the search bots (OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot, Bingbot) and disallow the training bots (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot). User-triggered fetchers may not follow robots.txt.

Documented user agents

TokenCompanyPurposeRobots.txt
GPTBotOpenAITrainingHonored[1]
OAI-SearchBotOpenAIChatGPT search resultsHonored[1]
ChatGPT-UserOpenAIUser-initiated actions“May not apply”[1]
PerplexityBotPerplexitySearch index; not trainingHonored[2]
Perplexity-UserPerplexityUser fetches“Generally ignores”[2]
ClaudeBot / Claude-SearchBot / Claude-UserAnthropicTraining / search quality / user fetchesHonored, incl. Crawl-delay[3]
GooglebotGoogleSearch, AI Overviews, AI ModeHonored[4]
Google-ExtendedGoogleToken only: Gemini training and grounding; no effect on SearchHonored[4]
BingbotMicrosoftBing, Copilot, groundingHonored[5]
Applebot / Applebot-ExtendedAppleSiri, Spotlight / token only: training opt-outHonored[6]
Meta-ExternalAgent / Meta-WebIndexerMetaTraining and indexing / Meta AI search citationsHonored; Meta-ExternalFetcher may bypass[7]
Amazonbot / Amzn-SearchBot / Amzn-UserAmazonMay train models / Alexa search / live fetchesHonored; no Crawl-delay[8]
CCBotCommon CrawlOpen dataset widely used for trainingHonored[9]

OpenAI and Amazon say robots.txt changes take about 24 hours to apply, and each agent is controlled independently.[1][8] Full user-agent strings are on each engine page, for example ChatGPT and Perplexity.

Tokens that don’t exist

Many templates online list agents such as “Bard”, “Claude-Web” or “Bing-AI”. None appears in current vendor documentation; rules for them do nothing. Use the tokens above. See AEO myths.

Example: visible in AI search, out of training

# Allow AI search and answer bots; opt out of model training
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml

Build your own with the robots.txt builder. Robots.txt is a request, not access control;[10] verify bots by published IP ranges where offered.[2][8]

Check your CDN too

Since July 1, 2025, Cloudflare blocks known AI crawlers by default on new domains.[11] A firewall rule can silently override a permissive robots.txt. Check server logs for 403s to the agents you want. Vendor-documented

Opting out of AI answers entirely

Frequently asked questions

Does blocking GPTBot remove me from ChatGPT search?

No. GPTBot is for training. ChatGPT search uses OAI-SearchBot, which is controlled separately.

Does blocking Google-Extended remove me from AI Overviews?

No. AI Overviews use Googlebot and the Search index. Use Search Console's generative AI control to opt out.

Is there a single token for all AI bots?

No. Each company documents its own tokens, and user-triggered fetchers may not follow robots.txt.

References

Pages accessed September 28, 2026 unless a date is given. See all sources and our editorial policy.

  1. ^ "Overview of OpenAI Crawlers". OpenAI Platform documentation.
  2. ^ "Perplexity Crawlers". Perplexity documentation.
  3. ^ "Does Anthropic crawl data from the web, and how can site owners block the crawler?". Anthropic (Claude Help Center).
  4. ^ "Google’s common crawlers (including Google-Extended)". Google Search Central. Last updated July 14, 2026.
  5. ^ "Bing Webmaster Guidelines". Microsoft Bing Webmaster Tools.
  6. ^ "About Applebot (including Applebot-Extended)". Apple Support.
  7. ^ "Meta Web Crawlers". Meta for Developers.
  8. ^ "About Amazonbot (Amazonbot, Amzn-SearchBot, Amzn-User)". Amazon Developer.
  9. ^ "CCBot". Common Crawl.
  10. ^ "RFC 9309: Robots Exclusion Protocol". IETF (RFC Editor). Published September 2022.
  11. ^ "Content Independence Day: no AI crawl without compensation!". Cloudflare Blog. Published July 1, 2025.
  12. ^ "Search generative AI control". Search Console Help. Rolled out to all sites by August 31, 2026.
  13. ^ "Robots meta tag, data-nosnippet, and X-Robots-Tag specifications". Google Search Central.
  14. ^ "Publishers and Developers – FAQ". OpenAI Help Center.