The Complete robots.txt Playbook for AI Crawlers
Allow the search-relevant AI bots if you want citations, decide separately on training bots, and state every rule explicitly rather than relying on defaults.
Maya Reinholt
Head of Search Research, Toolgram
Quick answer
Allow the search-relevant bots (OAI-SearchBot, PerplexityBot, ClaudeBot, Googlebot) if you want citations; decide separately on training bots (GPTBot, CCBot, Google-Extended, Bytespider) based on your licensing stance. State each rule explicitly rather than relying on defaults.
Key takeaways
- Search bots and training bots are separate decisions; blocking one does not block the other.
- PerplexityBot, OAI-SearchBot, and ClaudeBot drive citations, so allow them if you want AI search visibility.
- GPTBot, CCBot, Google-Extended, and Bytespider are training-oriented and can be blocked with no search downside.
- robots.txt is advisory; enforce real protection with edge rate limits and bot verification.
Answer engines cannot cite what their crawlers cannot read. The robots.txt file is where your AI search visibility is decided, and most sites still ship the default WordPress stub and call it done.
Why AI crawler policy is now an SEO decision
Every answer engine runs its own crawler, and each one feeds a different pipeline: training, search index, or real-time user fetches. A blanket "allow all" gives away training rights you may want to license. A blanket "block all bots" removes you from every synthesized answer on the internet. The correct stance is a explicit, per-bot policy.
The crawlers that matter in 2026:
| Crawler | Operator | Feeds | Typical stance | | --- | --- | --- | --- | | Googlebot | Google | Search + AI Overviews | Allow | | OAI-SearchBot | OpenAI | ChatGPT Search index | Allow if you want citations | | ChatGPT-User | OpenAI | Real-time user fetches | Allow (only fires on user action) | | GPTBot | OpenAI | Model training | Your licensing call | | PerplexityBot | Perplexity | Perplexity answers | Allow if you want citations | | ClaudeBot | Anthropic | Claude answers / real-time | Allow if you want citations | | anthropic-ai | Anthropic | Model training | Your licensing call | | Google-Extended | Google | Gemini training | Your licensing call | | Bytespider | ByteDance | Training + data | Block unless you want the traffic | | CCBot | Common Crawl | Open corpus | Your licensing call |
A production-ready robots.txt
This is the exact policy shape this site ships:
User-agent: Googlebot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: GPTBot
Disallow: /
User-agent: Bytespider
Disallow: /
User-agent: CCBot
Disallow: /
Sitemap: https://example.com/sitemap.xml
Swap the training-bot lines to Allow if you are comfortable contributing to model training. The search bots should stay allowed either way, because that is where citations and referral traffic come from.
robots.txt is advisory. A misbehaving crawler ignores it, which is why edge-level enforcement matters for anything you truly want to block.
Common mistakes
- One blanket rule for all bots. You cannot distinguish search access from training access with
User-agent: *. Name each bot. - Blocking the search bot while keeping the training bot. This is backwards for most publishers: the search bot sends citations and traffic, the training bot sends nothing measurable.
- Forgetting the sitemap line. Every crawler that honors robots.txt can discover your sitemap there. It is the cheapest indexation aid you will ever add.
- Testing in production only. Validate with Google's robots.txt tester and check server logs to confirm each named bot is hitting the pages you expect.
Your action checklist
- List the crawlers actually hitting your server this month (log analysis, reverse-DNS verified).
- Decide your training stance once, as policy, not per-incident.
- Write explicit per-bot rules and deploy.
- Verify citations begin appearing for your top queries within two weeks.
Citations follow structure, not luck.
Frequently asked questions
Does blocking GPTBot remove me from ChatGPT Search?
No. GPTBot collects training data. ChatGPT Search uses OAI-SearchBot, and real-time user fetches use ChatGPT-User. Block or allow each one separately according to your policy.
Will blocking Google-Extended remove me from AI Overviews?
No. AI Overviews follow normal Google crawling rules. Google-Extended only controls use of your content in Gemini training.
How do I stop Bytespider from hammering my server?
Disallow Bytespider in robots.txt and add a rate limit for its IP ranges at your CDN or reverse proxy, since it crawls far more aggressively than other bots.
Do AI crawlers always respect robots.txt?
The major published crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended) do respect it. Verification by reverse-DNS is still recommended, because user-agent strings are trivially spoofed.
Maya Reinholt
Maya Reinholt leads search research at Toolgram. She has spent nine years in technical SEO, the last three mapping how LLM-powered crawlers and answer engines select, parse, and cite web content.