Free tools/robots.txt AI policy generator

Free GEO toolrobots.txt generator for AI crawlers.

Choose what you want, not which bot names to memorise. Pick a preset or toggle each crawler, and get a commented robots.txt block that keeps you in AI answers, out of training, or both. Paste your current file to merge.

Free · no signup · updated 5 October 2026

Your AI crawler policy

Start from a preset, then adjust any crawler.

CrawlerPurposeAllow
GPTBotOpenAI · ChatGPT Training
OAI-SearchBotOpenAI · ChatGPT AI search index
ChatGPT-UserOpenAI · ChatGPT User-requested fetch
ClaudeBotAnthropic · Claude Training
Claude-SearchBotAnthropic · Claude AI search index
Claude-UserAnthropic · Claude User-requested fetch
PerplexityBotPerplexity · Perplexity AI search index
Perplexity-UserPerplexity · Perplexity User-requested fetch
Google-ExtendedGoogle · Gemini robots.txt control token
ApplebotApple · Siri, Spotlight AI search index
Applebot-ExtendedApple · Apple Intelligence robots.txt control token
CCBotCommon Crawl · Open datasets used by many models Training
Meta-ExternalAgentMeta · Meta AI Training
BytespiderByteDance · Doubao and other ByteDance models Training
DuckAssistBotDuckDuckGo · DuckAssist AI search index
MistralAI-UserMistral · Le Chat User-requested fetch

Not included on purpose: Googlebot and Bingbot. They power classic search and feed AI Overviews and Copilot; blocking them removes you from search entirely.

"/" covers the whole site. Use e.g. "/members/" to protect one section.

How do I block AI training but stay in AI search?

Disallow the training crawlers (GPTBot, ClaudeBot, CCBot, Meta-ExternalAgent, Bytespider) and the control tokens Google-Extended and Applebot-Extended, and leave the search and user bots (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot) allowed. Never block Googlebot or Bingbot to opt out of AI.

How it works

The method, in the open.

The generator uses the same crawler list as our AI crawler checker, grouped by purpose. Each preset sets every bot at once:

  • Visible in AI answers, no training: search and user bots allowed; training bots and control tokens disallowed.
  • Fully open: every AI crawler allowed. Right for most businesses that want maximum AI visibility.
  • Block all AI: every AI-specific crawler and token disallowed. Googlebot and Bingbot stay allowed, so you remain in classic search.

Each line in the output carries a comment explaining what it does. If you paste your existing robots.txt, the generator removes any groups that only name AI bots it manages, keeps everything else exactly as it was, and appends the new block, so your other rules and sitemap lines survive.

Worked example

What the numbers look like.

The "visible in AI answers, no training" preset produces groups like:

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

That pair keeps a site eligible for ChatGPT search citations while telling OpenAI not to use it for training. Run the result through the AI crawler checker after you publish it to confirm your CDN agrees.

Sources

What this is based on

  • Disallowing GPTBot signals content should not be used for training; OAI-SearchBot controls ChatGPT search inclusion. OpenAI
  • Google-Extended controls Gemini training and grounding and does not affect Google Search. Google Search Central
  • Applebot-Extended opts out of Apple foundation-model training while Applebot still powers Siri and Spotlight. Apple
  • Research finds scrapers only selectively respect robots.txt: it is a policy statement, not enforcement. arXiv 2505.21733
Limitations

What it does not tell you

  • robots.txt is voluntary. Well-known AI companies honour it for their main crawlers, but user-initiated fetchers may not, and some scrapers ignore it entirely.
  • Changes take time: operators re-read robots.txt on their own schedule, often within a day.
  • Blocking a crawler does not delete content already collected.
  • The merge only replaces groups that name AI bots exclusively. Review the output before publishing it.
FAQ

Questions people ask.

Will blocking AI bots hurt my Google rankings?

Not if you leave Googlebot alone. Google-Extended is a separate token that only affects Gemini training and grounding; Google Search, AI Overviews and AI Mode keep using Googlebot.

Should a business block AI crawlers?

Most businesses that sell something want to be recommended by AI assistants, so they should at least allow the search and user bots. Publishers who license content may choose to block training bots. The presets cover both.

Where do I put the robots.txt file?

At the root of each host, e.g. https://www.example.com/robots.txt. Each subdomain needs its own file. It must be plain text and return HTTP 200.

Do I also need an llms.txt file?

It is optional. Large studies have found no measurable citation benefit from llms.txt, and no major AI provider documents it as a ranking signal. robots.txt and your CDN settings decide whether bots can read you at all.

Can I block AI bots on just part of my site?

Yes. Replace "Disallow: /" with the paths you want to protect, e.g. "Disallow: /members/". The longest matching rule wins, so a specific Allow can open a sub-path again.

How do I know it worked?

Publish the file, then run the AI crawler checker on your domain. It evaluates the live file for every bot and tests whether your CDN blocks them.

Not sure which preset fits?

Ask us. We'll set the policy with you.

We will look at your business model, licensing and AI visibility goals and recommend a crawler policy in one short call.