Free GEO toolrobots.txt generator for AI crawlers.
Choose what you want, not which bot names to memorise. Pick a preset or toggle each crawler, and get a commented robots.txt block that keeps you in AI answers, out of training, or both. Paste your current file to merge.
Free · no signup · updated 5 October 2026
Your AI crawler policy
Start from a preset, then adjust any crawler.
| Crawler | Purpose | Allow |
|---|---|---|
GPTBotOpenAI · ChatGPT |
Training | |
OAI-SearchBotOpenAI · ChatGPT |
AI search index | |
ChatGPT-UserOpenAI · ChatGPT |
User-requested fetch | |
ClaudeBotAnthropic · Claude |
Training | |
Claude-SearchBotAnthropic · Claude |
AI search index | |
Claude-UserAnthropic · Claude |
User-requested fetch | |
PerplexityBotPerplexity · Perplexity |
AI search index | |
Perplexity-UserPerplexity · Perplexity |
User-requested fetch | |
Google-ExtendedGoogle · Gemini |
robots.txt control token | |
ApplebotApple · Siri, Spotlight |
AI search index | |
Applebot-ExtendedApple · Apple Intelligence |
robots.txt control token | |
CCBotCommon Crawl · Open datasets used by many models |
Training | |
Meta-ExternalAgentMeta · Meta AI |
Training | |
BytespiderByteDance · Doubao and other ByteDance models |
Training | |
DuckAssistBotDuckDuckGo · DuckAssist |
AI search index | |
MistralAI-UserMistral · Le Chat |
User-requested fetch |
Not included on purpose: Googlebot and Bingbot. They power classic search and feed AI Overviews and Copilot; blocking them removes you from search entirely.
Your robots.txt
How do I block AI training but stay in AI search?
Disallow the training crawlers (GPTBot, ClaudeBot, CCBot, Meta-ExternalAgent, Bytespider) and the control tokens Google-Extended and Applebot-Extended, and leave the search and user bots (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot) allowed. Never block Googlebot or Bingbot to opt out of AI.
The method, in the open.
The generator uses the same crawler list as our AI crawler checker, grouped by purpose. Each preset sets every bot at once:
- Visible in AI answers, no training: search and user bots allowed; training bots and control tokens disallowed.
- Fully open: every AI crawler allowed. Right for most businesses that want maximum AI visibility.
- Block all AI: every AI-specific crawler and token disallowed. Googlebot and Bingbot stay allowed, so you remain in classic search.
Each line in the output carries a comment explaining what it does. If you paste your existing robots.txt, the generator removes any groups that only name AI bots it manages, keeps everything else exactly as it was, and appends the new block, so your other rules and sitemap lines survive.
What the numbers look like.
The "visible in AI answers, no training" preset produces groups like:
User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Allow: /
That pair keeps a site eligible for ChatGPT search citations while telling OpenAI not to use it for training. Run the result through the AI crawler checker after you publish it to confirm your CDN agrees.
What this is based on
- Disallowing GPTBot signals content should not be used for training; OAI-SearchBot controls ChatGPT search inclusion. OpenAI
- Google-Extended controls Gemini training and grounding and does not affect Google Search. Google Search Central
- Applebot-Extended opts out of Apple foundation-model training while Applebot still powers Siri and Spotlight. Apple
- Research finds scrapers only selectively respect robots.txt: it is a policy statement, not enforcement. arXiv 2505.21733
What it does not tell you
- robots.txt is voluntary. Well-known AI companies honour it for their main crawlers, but user-initiated fetchers may not, and some scrapers ignore it entirely.
- Changes take time: operators re-read robots.txt on their own schedule, often within a day.
- Blocking a crawler does not delete content already collected.
- The merge only replaces groups that name AI bots exclusively. Review the output before publishing it.
Questions people ask.
Will blocking AI bots hurt my Google rankings?
Not if you leave Googlebot alone. Google-Extended is a separate token that only affects Gemini training and grounding; Google Search, AI Overviews and AI Mode keep using Googlebot.
Should a business block AI crawlers?
Most businesses that sell something want to be recommended by AI assistants, so they should at least allow the search and user bots. Publishers who license content may choose to block training bots. The presets cover both.
Where do I put the robots.txt file?
At the root of each host, e.g. https://www.example.com/robots.txt. Each subdomain needs its own file. It must be plain text and return HTTP 200.
Do I also need an llms.txt file?
It is optional. Large studies have found no measurable citation benefit from llms.txt, and no major AI provider documents it as a ranking signal. robots.txt and your CDN settings decide whether bots can read you at all.
Can I block AI bots on just part of my site?
Yes. Replace "Disallow: /" with the paths you want to protect, e.g. "Disallow: /members/". The longest matching rule wins, so a specific Allow can open a sub-path again.
How do I know it worked?
Publish the file, then run the AI crawler checker on your domain. It evaluates the live file for every bot and tests whether your CDN blocks them.
Keep checking.
AI crawler checker
Is GPTBot, OAI-SearchBot, ClaudeBot or PerplexityBot blocked on your site? Checks robots.txt and tests your CDN, bot by bot.
GA4 AI traffic regex generator
Track ChatGPT, Perplexity, Claude, Gemini and Copilot visits in GA4. Generates the regex, channel group steps, a Looker Studio formula and a GTM variable.
Zero-click impact calculator
Paste your Search Console export to estimate clicks lost to AI Overviews and the upside of being cited. Pick the study you trust. Runs in your browser.
Ask us. We'll set the policy with you.
We will look at your business model, licensing and AI visibility goals and recommend a crawler policy in one short call.