Free GEO toolAI crawler checker: can ChatGPT, Claude and Perplexity read your site?
Enter a domain or page. We read your live robots.txt for 18 AI crawlers, then request the page as each one to see whether your CDN or firewall quietly blocks it. Results split training bots from the bots that put you in AI answers.
Free · no signup · updated 5 October 2026
Check a site
A domain checks the homepage. A full URL checks that page.
Takes about ten seconds. We fetch robots.txt and request the page once per crawler. Results are cached for an hour.
Crawler by crawler
| Crawler | Purpose | robots.txt | Live request |
|---|
Your robots.txt
The 18 AI crawlers we check (reviewed 2026-10-05)
| Crawler | Operator | Purpose | What blocking it does |
|---|---|---|---|
GPTBot | OpenAI | Training | Opts your content out of OpenAI model training. Does not remove you from ChatGPT search answers. Source |
OAI-SearchBot | OpenAI | AI search index | Removes your pages from ChatGPT search results and citations. Source |
ChatGPT-User | OpenAI | User-requested fetch | Asks ChatGPT not to open your pages when a user requests them. OpenAI says robots.txt rules may not apply to user-initiated requests. Source |
ClaudeBot | Anthropic | Training | Opts your content out of Anthropic model training. Does not affect Claude search or user fetches. Source |
Claude-SearchBot | Anthropic | AI search index | Stops Claude indexing your pages for search, so they are less likely to appear in Claude answers. Source |
Claude-User | Anthropic | User-requested fetch | Stops Claude fetching your pages when a user asks about them. Anthropic says it honours robots.txt for this bot. Source |
PerplexityBot | Perplexity | AI search index | Removes your pages from Perplexity search results and citations. Perplexity says it is not used for model training. Source |
Perplexity-User | Perplexity | User-requested fetch | Fetches pages a user asks about. Perplexity says this agent generally ignores robots.txt. Source |
Googlebot | Search engine (feeds AI answers) | Removes you from Google Search, including AI Overviews and AI Mode. Never block it to opt out of AI. Source | |
Google-Extended | robots.txt control token | Opts out of Gemini model training and grounding. Does not affect Google Search, AI Overviews or AI Mode. Source | |
Bingbot | Microsoft | Search engine (feeds AI answers) | Removes you from Bing, and with it from Copilot answers, which are grounded in the Bing index. Source |
Applebot | Apple | AI search index | Removes you from Siri and Spotlight suggestions. Source |
Applebot-Extended | Apple | robots.txt control token | Opts out of Apple foundation-model training. Applebot can still crawl you for Siri and Spotlight. Source |
CCBot | Common Crawl | Training | Keeps you out of Common Crawl, a public dataset many AI models are trained on. Source |
Meta-ExternalAgent | Meta | Training | Opts out of Meta crawling for AI training and products. Source |
Bytespider | ByteDance | Training | Opts out of ByteDance crawling. Studies report it does not always honour robots.txt. Source |
DuckAssistBot | DuckDuckGo | AI search index | Removes your pages from DuckDuckGo AI-assisted answers. Source |
MistralAI-User | Mistral | User-requested fetch | Stops Le Chat fetching your pages when a user asks about them. Source |
Does blocking GPTBot remove my site from ChatGPT?
No. GPTBot only collects training data. ChatGPT search results and citations come from OAI-SearchBot, and user-requested page visits from ChatGPT-User. Block GPTBot to opt out of training; block OAI-SearchBot only if you want to disappear from ChatGPT answers. The same split applies to Claude and Perplexity.
The method, in the open.
1. robots.txt. We fetch /robots.txt from your domain and evaluate it for each crawler the way RFC 9309 describes: the crawler obeys the group that names it, falls back to User-agent: *, and the longest matching rule wins. If robots.txt returns a server error or cannot be reached, crawlers treat that as "crawl nothing", and we flag it.
2. Live request. We request your page once as a normal browser and once as each testable AI crawler. A 403, 429 or challenge page for a bot, when the browser request succeeds, points to a CDN or firewall rule that blocks AI bots regardless of robots.txt.
3. Verdict by purpose. Every bot is labelled training, AI search, user-requested fetch, control token or search engine, so you can see at a glance whether you are opted out of training, out of AI answers, or both.
What the numbers look like.
A common, sensible setup. A publisher wants to appear in AI answers but not be used for training. Its robots.txt disallows GPTBot, ClaudeBot, CCBot and Google-Extended, and allows everything else. The checker shows: training blocked for OpenAI, Anthropic, Common Crawl and Gemini; AI search allowed for ChatGPT, Claude and Perplexity; Googlebot allowed, so AI Overviews still work.
A common mistake. A SaaS site turns on its CDN's "block AI bots" switch. robots.txt allows everything, but the live test shows OAI-SearchBot and PerplexityBot get a 403. The site is invisible to ChatGPT search and Perplexity, and nothing in robots.txt explains why.
What this is based on
- OpenAI: OAI-SearchBot powers ChatGPT search; GPTBot is for training; robots.txt rules may not apply to user-initiated ChatGPT-User requests. OpenAI
- Anthropic runs ClaudeBot (training), Claude-SearchBot (search) and Claude-User (user requests), each controllable separately. Anthropic
- PerplexityBot indexes for Perplexity answers and is not used for training; Perplexity-User generally ignores robots.txt. Perplexity
- Google-Extended does not affect Google Search; AI Overviews and AI Mode use the normal Search index and Googlebot. Google Search Central
- robots.txt rules, matching and error handling. RFC 9309
What it does not tell you
- The live test sends each bot's user agent from our server, not from the vendor's IP ranges. CDNs that verify bots by IP may treat the real crawler differently, so a pass is indicative, not proof.
- Googlebot and Bingbot are checked in robots.txt only. CDNs block spoofed versions of them by design, so a live test would mislead.
- robots.txt is a request, not enforcement. Studies show some scrapers do not honour it.
- We test the one URL you enter. Rules for other paths can differ; check important sections separately.
Questions people ask.
Does blocking GPTBot remove me from ChatGPT?
No. GPTBot is OpenAI's training crawler. ChatGPT search uses OAI-SearchBot, and page visits a user asks for use ChatGPT-User. Blocking GPTBot opts you out of training while keeping you eligible for ChatGPT answers.
Does Google-Extended affect AI Overviews?
No. Google-Extended only controls use of your content for Gemini model training and grounding. AI Overviews and AI Mode are part of Google Search and use Googlebot. Blocking Googlebot to avoid AI Overviews would remove you from Search entirely.
Why does my robots.txt allow a bot but the live test fails?
Usually a CDN or firewall rule. Cloudflare, Akamai and similar services offer one-click AI bot blocking that works independently of robots.txt. Check your bot management settings and allow the search and user bots you want.
Which AI bots should I allow?
If you want to be cited, allow the AI search and user-fetch bots (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot) plus Googlebot and Bingbot. Training bots are a business decision; blocking them does not remove you from answers.
Should I rate-limit AI bots instead of blocking them?
If crawl load is the problem, yes. Most AI crawlers respect robots.txt but not Crawl-delay, so rate limiting at the CDN is the reliable way to control load without losing visibility.
Does Copilot use its own crawler?
Microsoft Copilot grounds its answers in the Bing index, so Bingbot access and Bing indexing matter. Submitting pages through IndexNow speeds up Bing discovery.
Keep checking.
robots.txt AI policy generator
Pick a policy, such as "visible in AI search, no training", and get a commented robots.txt block for every AI bot. Merges with your existing file.
GA4 AI traffic regex generator
Track ChatGPT, Perplexity, Claude, Gemini and Copilot visits in GA4. Generates the regex, channel group steps, a Looker Studio formula and a GTM variable.
Zero-click impact calculator
Paste your Search Console export to estimate clicks lost to AI Overviews and the upside of being cited. Pick the study you trust. Runs in your browser.
Get a full technical GEO audit.
We check crawl access, indexing in Google and Bing, page structure and where AI engines already cite you, then fix what is missing.