Free tools/AI crawler checker

Free GEO toolAI crawler checker: can ChatGPT, Claude and Perplexity read your site?

Enter a domain or page. We read your live robots.txt for 18 AI crawlers, then request the page as each one to see whether your CDN or firewall quietly blocks it. Results split training bots from the bots that put you in AI answers.

Free · no signup · updated 5 October 2026

Check a site

A domain checks the homepage. A full URL checks that page.

Takes about ten seconds. We fetch robots.txt and request the page once per crawler. Results are cached for an hour.

The 18 AI crawlers we check (reviewed 2026-10-05)
CrawlerOperatorPurposeWhat blocking it does
GPTBotOpenAITrainingOpts your content out of OpenAI model training. Does not remove you from ChatGPT search answers. Source
OAI-SearchBotOpenAIAI search indexRemoves your pages from ChatGPT search results and citations. Source
ChatGPT-UserOpenAIUser-requested fetchAsks ChatGPT not to open your pages when a user requests them. OpenAI says robots.txt rules may not apply to user-initiated requests. Source
ClaudeBotAnthropicTrainingOpts your content out of Anthropic model training. Does not affect Claude search or user fetches. Source
Claude-SearchBotAnthropicAI search indexStops Claude indexing your pages for search, so they are less likely to appear in Claude answers. Source
Claude-UserAnthropicUser-requested fetchStops Claude fetching your pages when a user asks about them. Anthropic says it honours robots.txt for this bot. Source
PerplexityBotPerplexityAI search indexRemoves your pages from Perplexity search results and citations. Perplexity says it is not used for model training. Source
Perplexity-UserPerplexityUser-requested fetchFetches pages a user asks about. Perplexity says this agent generally ignores robots.txt. Source
GooglebotGoogleSearch engine (feeds AI answers)Removes you from Google Search, including AI Overviews and AI Mode. Never block it to opt out of AI. Source
Google-ExtendedGooglerobots.txt control tokenOpts out of Gemini model training and grounding. Does not affect Google Search, AI Overviews or AI Mode. Source
BingbotMicrosoftSearch engine (feeds AI answers)Removes you from Bing, and with it from Copilot answers, which are grounded in the Bing index. Source
ApplebotAppleAI search indexRemoves you from Siri and Spotlight suggestions. Source
Applebot-ExtendedApplerobots.txt control tokenOpts out of Apple foundation-model training. Applebot can still crawl you for Siri and Spotlight. Source
CCBotCommon CrawlTrainingKeeps you out of Common Crawl, a public dataset many AI models are trained on. Source
Meta-ExternalAgentMetaTrainingOpts out of Meta crawling for AI training and products. Source
BytespiderByteDanceTrainingOpts out of ByteDance crawling. Studies report it does not always honour robots.txt. Source
DuckAssistBotDuckDuckGoAI search indexRemoves your pages from DuckDuckGo AI-assisted answers. Source
MistralAI-UserMistralUser-requested fetchStops Le Chat fetching your pages when a user asks about them. Source

Does blocking GPTBot remove my site from ChatGPT?

No. GPTBot only collects training data. ChatGPT search results and citations come from OAI-SearchBot, and user-requested page visits from ChatGPT-User. Block GPTBot to opt out of training; block OAI-SearchBot only if you want to disappear from ChatGPT answers. The same split applies to Claude and Perplexity.

How it works

The method, in the open.

1. robots.txt. We fetch /robots.txt from your domain and evaluate it for each crawler the way RFC 9309 describes: the crawler obeys the group that names it, falls back to User-agent: *, and the longest matching rule wins. If robots.txt returns a server error or cannot be reached, crawlers treat that as "crawl nothing", and we flag it.

2. Live request. We request your page once as a normal browser and once as each testable AI crawler. A 403, 429 or challenge page for a bot, when the browser request succeeds, points to a CDN or firewall rule that blocks AI bots regardless of robots.txt.

3. Verdict by purpose. Every bot is labelled training, AI search, user-requested fetch, control token or search engine, so you can see at a glance whether you are opted out of training, out of AI answers, or both.

Worked example

What the numbers look like.

A common, sensible setup. A publisher wants to appear in AI answers but not be used for training. Its robots.txt disallows GPTBot, ClaudeBot, CCBot and Google-Extended, and allows everything else. The checker shows: training blocked for OpenAI, Anthropic, Common Crawl and Gemini; AI search allowed for ChatGPT, Claude and Perplexity; Googlebot allowed, so AI Overviews still work.

A common mistake. A SaaS site turns on its CDN's "block AI bots" switch. robots.txt allows everything, but the live test shows OAI-SearchBot and PerplexityBot get a 403. The site is invisible to ChatGPT search and Perplexity, and nothing in robots.txt explains why.

Sources

What this is based on

  • OpenAI: OAI-SearchBot powers ChatGPT search; GPTBot is for training; robots.txt rules may not apply to user-initiated ChatGPT-User requests. OpenAI
  • Anthropic runs ClaudeBot (training), Claude-SearchBot (search) and Claude-User (user requests), each controllable separately. Anthropic
  • PerplexityBot indexes for Perplexity answers and is not used for training; Perplexity-User generally ignores robots.txt. Perplexity
  • Google-Extended does not affect Google Search; AI Overviews and AI Mode use the normal Search index and Googlebot. Google Search Central
  • robots.txt rules, matching and error handling. RFC 9309
Limitations

What it does not tell you

  • The live test sends each bot's user agent from our server, not from the vendor's IP ranges. CDNs that verify bots by IP may treat the real crawler differently, so a pass is indicative, not proof.
  • Googlebot and Bingbot are checked in robots.txt only. CDNs block spoofed versions of them by design, so a live test would mislead.
  • robots.txt is a request, not enforcement. Studies show some scrapers do not honour it.
  • We test the one URL you enter. Rules for other paths can differ; check important sections separately.
FAQ

Questions people ask.

Does blocking GPTBot remove me from ChatGPT?

No. GPTBot is OpenAI's training crawler. ChatGPT search uses OAI-SearchBot, and page visits a user asks for use ChatGPT-User. Blocking GPTBot opts you out of training while keeping you eligible for ChatGPT answers.

Does Google-Extended affect AI Overviews?

No. Google-Extended only controls use of your content for Gemini model training and grounding. AI Overviews and AI Mode are part of Google Search and use Googlebot. Blocking Googlebot to avoid AI Overviews would remove you from Search entirely.

Why does my robots.txt allow a bot but the live test fails?

Usually a CDN or firewall rule. Cloudflare, Akamai and similar services offer one-click AI bot blocking that works independently of robots.txt. Check your bot management settings and allow the search and user bots you want.

Which AI bots should I allow?

If you want to be cited, allow the AI search and user-fetch bots (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot) plus Googlebot and Bingbot. Training bots are a business decision; blocking them does not remove you from answers.

Should I rate-limit AI bots instead of blocking them?

If crawl load is the problem, yes. Most AI crawlers respect robots.txt but not Crawl-delay, so rate limiting at the CDN is the reliable way to control load without losing visibility.

Does Copilot use its own crawler?

Microsoft Copilot grounds its answers in the Bing index, so Bingbot access and Bing indexing matter. Submitting pages through IndexNow speeds up Bing discovery.

Crawlable is step one

Get a full technical GEO audit.

We check crawl access, indexing in Google and Bing, page structure and where AI engines already cite you, then fix what is missing.