AI crawlers and robots.txt
Most AI companies now run separate crawlers for model training, search indexing and user-triggered fetches, and document robots.txt tokens for each. To appear in AI answers while opting out of training, allow the search bots (OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot, Bingbot) and disallow the training bots (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot). User-triggered fetchers may not follow robots.txt.
Documented user agents
| Token | Company | Purpose | Robots.txt |
|---|---|---|---|
| GPTBot | OpenAI | Training | Honored[1] |
| OAI-SearchBot | OpenAI | ChatGPT search results | Honored[1] |
| ChatGPT-User | OpenAI | User-initiated actions | “May not apply”[1] |
| PerplexityBot | Perplexity | Search index; not training | Honored[2] |
| Perplexity-User | Perplexity | User fetches | “Generally ignores”[2] |
| ClaudeBot / Claude-SearchBot / Claude-User | Anthropic | Training / search quality / user fetches | Honored, incl. Crawl-delay[3] |
| Googlebot | Search, AI Overviews, AI Mode | Honored[4] | |
| Google-Extended | Token only: Gemini training and grounding; no effect on Search | Honored[4] | |
| Bingbot | Microsoft | Bing, Copilot, grounding | Honored[5] |
| Applebot / Applebot-Extended | Apple | Siri, Spotlight / token only: training opt-out | Honored[6] |
| Meta-ExternalAgent / Meta-WebIndexer | Meta | Training and indexing / Meta AI search citations | Honored; Meta-ExternalFetcher may bypass[7] |
| Amazonbot / Amzn-SearchBot / Amzn-User | Amazon | May train models / Alexa search / live fetches | Honored; no Crawl-delay[8] |
| CCBot | Common Crawl | Open dataset widely used for training | Honored[9] |
OpenAI and Amazon say robots.txt changes take about 24 hours to apply, and each agent is controlled independently.[1][8] Full user-agent strings are on each engine page, for example ChatGPT and Perplexity.
Tokens that don’t exist
Many templates online list agents such as “Bard”, “Claude-Web” or “Bing-AI”. None appears in current vendor documentation; rules for them do nothing. Use the tokens above. See AEO myths.
Example: visible in AI search, out of training
# Allow AI search and answer bots; opt out of model training
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
Build your own with the robots.txt builder. Robots.txt is a request, not access control;[10] verify bots by published IP ranges where offered.[2][8]
Check your CDN too
Since July 1, 2025, Cloudflare blocks known AI crawlers by default on new domains.[11] A firewall rule can silently override a permissive robots.txt. Check server logs for 403s to the agents you want. Vendor-documented
Opting out of AI answers entirely
- Google: the Search Console generative AI control, or
nosnippet/noindex.[12][13] - Bing/Copilot:
NOARCHIVEorNOCACHE.[5] - OpenAI: disallow OAI-SearchBot; use
noindexto keep link-only listings out.[14]
Frequently asked questions
Does blocking GPTBot remove me from ChatGPT search?
No. GPTBot is for training. ChatGPT search uses OAI-SearchBot, which is controlled separately.
Does blocking Google-Extended remove me from AI Overviews?
No. AI Overviews use Googlebot and the Search index. Use Search Console's generative AI control to opt out.
Is there a single token for all AI bots?
No. Each company documents its own tokens, and user-triggered fetchers may not follow robots.txt.
See also
References
Pages accessed September 28, 2026 unless a date is given. See all sources and our editorial policy.
- ^ "Overview of OpenAI Crawlers". OpenAI Platform documentation.
- ^ "Perplexity Crawlers". Perplexity documentation.
- ^ "Does Anthropic crawl data from the web, and how can site owners block the crawler?". Anthropic (Claude Help Center).
- ^ "Google’s common crawlers (including Google-Extended)". Google Search Central. Last updated July 14, 2026.
- ^ "Bing Webmaster Guidelines". Microsoft Bing Webmaster Tools.
- ^ "About Applebot (including Applebot-Extended)". Apple Support.
- ^ "Meta Web Crawlers". Meta for Developers.
- ^ "About Amazonbot (Amazonbot, Amzn-SearchBot, Amzn-User)". Amazon Developer.
- ^ "CCBot". Common Crawl.
- ^ "RFC 9309: Robots Exclusion Protocol". IETF (RFC Editor). Published September 2022.
- ^ "Content Independence Day: no AI crawl without compensation!". Cloudflare Blog. Published July 1, 2025.
- ^ "Search generative AI control". Search Console Help. Rolled out to all sites by August 31, 2026.
- ^ "Robots meta tag, data-nosnippet, and X-Robots-Tag specifications". Google Search Central.
- ^ "Publishers and Developers – FAQ". OpenAI Help Center.