OAI-SearchBot vs GPTBot vs ChatGPT-User: what each one does, and what to block.

Many sites blocked GPTBot to stay out of AI training, then wondered why ChatGPT never mentions them. They usually blocked the wrong thing for the right reason.

By Arsh Goyal, Growth Engineer 3 min read
OAI-SearchBot vs GPTBot vs ChatGPT-User: what each one does, and what to block.

In 2023 a lot of sites added "Disallow: /" for GPTBot to keep their content out of AI training. Plenty of them now wonder why ChatGPT never mentions them. The crawlers have split since then, and robots.txt rules written for one bot don't cover the others.

The OpenAI three

OpenAI documents three crawlers:

CrawlerWhat it doesIf you block it
GPTBotCrawls content that may be used to train OpenAI's modelsYour content shouldn't be used for training. ChatGPT search can still cite you.
OAI-SearchBotIndexes pages for ChatGPT's search featuresYour pages won't appear in ChatGPT search results.
ChatGPT-UserVisits a page when a ChatGPT user's request needs itOpenAI says robots.txt rules may not apply, because a person triggered the visit.

So the common setup, block GPTBot and allow everything else, does what most people actually want: no training, still citable.

Claude and Perplexity work the same way

Anthropic runs ClaudeBot for training, Claude-SearchBot for search indexing and Claude-User for fetches a user asks for. Each can be controlled separately, and Anthropic says all three honour robots.txt.

Perplexity runs PerplexityBot to index pages for its answers, and says it is not used for training. Perplexity-User fetches pages a user asks about and, by Perplexity's own description, generally ignores robots.txt.

Google and Apple are different

Google-Extended isn't a crawler. It is a robots.txt token that controls whether Google can use your content for Gemini training and grounding. It doesn't affect Google Search, AI Overviews or AI Mode, which all use the normal Googlebot index. Block Googlebot to avoid AI Overviews and you leave Google entirely.

Apple has the same split: Applebot powers Siri and Spotlight, and Applebot-Extended controls whether your content trains Apple's models.

How robots.txt decides

robots.txt is now a formal standard, RFC 9309. A crawler follows the group that names its user agent and falls back to User-agent: * only when no group names it. Within a group, the longest matching rule wins. That has a side effect people miss: give GPTBot its own group, and it stops reading your * rules entirely, including the ones protecting /admin/ or /checkout/. Copy those rules into every bot-specific group.

A typical "no training, still citable" policy looks like this:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

Two things robots.txt won't tell you

Your CDN might block them anyway. Several CDNs offer one-click AI bot blocking that ignores robots.txt completely. A site can say "allow" and still return 403 to OAI-SearchBot.

They might reach the page and see nothing. OpenAI, Anthropic, Perplexity and Meta's crawlers don't run JavaScript. A client-rendered site can be fully allowed and still look empty to them.

The free AI crawler checker tests all three problems at once: it reads your live robots.txt for 18 AI crawlers, requests your page as each one to catch CDN blocks, and counts the words a crawler can read without JavaScript. To change your policy, the robots.txt AI generator writes the rules with a comment on every line.

Questions people ask

Does blocking GPTBot remove my site from ChatGPT?

No. GPTBot is for training. ChatGPT search uses OAI-SearchBot. Block GPTBot to opt out of training while staying citable.

Does blocking Google-Extended remove me from AI Overviews?

No. AI Overviews and AI Mode use Googlebot. Google-Extended only affects Gemini training and grounding.

Which AI bots should a business allow?

At least OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User and PerplexityBot, plus Googlebot and Bingbot. Those decide whether AI answers can cite you.

Want this mapped onto your product?

See how our GEO work runs, or send your store link and a founder will reply personally.