robots.txt for AI crawlers: which bots to allow, block or limit.
"Block AI" and "allow AI" are both the wrong question. Each company runs separate bots for search, training and user requests, and blocking the wrong one removes you from answers.
Half the robots.txt files we audit block the bot that gets them cited and allow the one they meant to block.
Three kinds of AI bot
| Type | Examples | If you block it |
|---|---|---|
| Search | OAI-SearchBot, Claude-SearchBot, PerplexityBot, Applebot, DuckAssistBot | You can disappear from that assistant's search answers |
| Training | GPTBot, ClaudeBot, CCBot, Meta-ExternalAgent, Bytespider | Your content is opted out of future model training; search visibility is unaffected |
| User-triggered | ChatGPT-User, Claude-User, Perplexity-User, MistralAI-User | The assistant can't open your page when a user asks it to |
| Control tokens | Google-Extended, Applebot-Extended | Opts out of Gemini or Apple model use; Google Search and Siri still crawl you |
The details per bot, with each vendor's documentation, are in our robots.txt AI generator. OpenAI's three bots are explained in OAI-SearchBot vs GPTBot vs ChatGPT-User.
The decision
- Want to be cited by AI assistants? Allow all search bots and user-triggered bots. This is the default for almost every business selling something.
- Comfortable with training? Allowing training bots may help models know your brand in future versions. Blocking them is a legitimate choice for publishers whose content is the product.
- Never block Googlebot or Bingbot to stop AI. Googlebot feeds Google Search, AI Overviews and AI Mode; Bingbot feeds Bing and Copilot.
- Use Google-Extended if you want to opt out of Gemini use without leaving Google Search.
A sensible default
# Search and user bots: allowed (be cited)
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /
# Training: your call. Remove these lines to allow.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
Disallow: /
Sitemap: https://www.example.com/sitemap.xml
robots.txt is only the first gate
- Firewalls and CDNs often block AI bots by default, whatever robots.txt says. Our AI crawler checker tests real requests with each bot's user agent.
- JavaScript rendering. Most AI fetchers don't run scripts. If content needs JavaScript, an allowed bot can still see an empty page.
- Compliance varies. Major vendors document that they honour robots.txt; not every scraper does. robots.txt is a request, not a lock.
Questions people ask
Should I block AI crawlers in robots.txt?
Not all of them. Blocking search bots like OAI-SearchBot or PerplexityBot can remove you from AI answers. Blocking training bots like GPTBot or CCBot only opts you out of model training.
Does blocking GPTBot remove me from ChatGPT search?
No. GPTBot is OpenAI's training crawler. ChatGPT search uses OAI-SearchBot, and user-requested fetches use ChatGPT-User.
How do I opt out of Gemini without leaving Google Search?
Disallow the Google-Extended token in robots.txt. Keep Googlebot allowed, because it powers Google Search, AI Overviews and AI Mode.
Want this mapped onto your product?
See how our GEO work runs, or send your store link and a founder will reply personally.