GPTBot, OAI-SearchBot, and ChatGPT-User: What Each One Does
GPTBot trains models, OAI-SearchBot feeds ChatGPT Search results, and ChatGPT-User fetches pages in real time when a user asks.
Maya Reinholt
Head of Search Research, Toolgram
Quick answer
GPTBot collects training data, OAI-SearchBot builds the ChatGPT Search index, and ChatGPT-User fetches pages in real time when a user asks a question. Blocking GPTBot keeps content out of training but does not remove you from search answers.
Key takeaways
- Three distinct OpenAI crawlers feed three distinct pipelines: training, search index, and live user fetches.
- Blocking GPTBot affects training only; it does not remove you from ChatGPT Search answers.
- OAI-SearchBot is the one that sends citations and referral traffic, so keep it allowed.
- ChatGPT-User only fires on explicit user action and respects rate conventions.
OpenAI operates three separate crawlers, and treating them as one bot is the most common robots.txt mistake of the past two years.
The three crawlers, mapped to pipelines
GPTBot crawls the public web to collect training data for future models. It is a bulk crawler with meaningful bandwidth impact and no direct referral traffic payoff. Allowing it is a licensing decision, not a search decision.
OAI-SearchBot builds and refreshes the index behind ChatGPT Search. This is the bot whose access determines whether ChatGPT can cite you. If citations and referral traffic from chatgpt.com are the goal, this bot stays allowed.
ChatGPT-User fetches single pages in real time when a user's request requires it: opening a shared link, reading a page the conversation mentions. It does not crawl; it visits on explicit user action. Blocking it degrades answers about your pages without protecting anything.
| Bot | Pipeline | Fires when | Traffic payoff | | --- | --- | --- | --- | | GPTBot | Model training | Continuous crawl | None directly | | OAI-SearchBot | ChatGPT Search index | Continuous crawl | Citations + referrals | | ChatGPT-User | Real-time fetch | On user request | Indirect (better answers about you) |
The correct robots.txt shape
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: GPTBot
Disallow: /
This configuration keeps you citable in ChatGPT Search while withholding training rights. Invert the GPTBot rule if you would rather contribute to training; there is no measurable search penalty either way.
Policy first, config second. Decide what you owe each pipeline before writing a single line of robots.txt.
Common mistakes
- Blocking
GPTBotand assuming ChatGPT Search lost you. It did not; OAI-SearchBot did the indexing. - Naming only
GPTBotin a blanket block. The other two bots keep crawling, which may be the opposite of the intended policy. - Skipping IP verification. Spoofed GPTBot user-agents are common in scrape attempts. Reverse-DNS against OpenAI's published ranges before drawing conclusions.
Your action checklist
- Grep this month's server logs for the three user-agents and their hit counts.
- Verify the top talkers with reverse-DNS.
- Align your robots.txt with your actual training and search policy.
- Add chatgpt.com as a referral source in analytics so citations become measurable.
Small structural changes compound across every answer engine.
Frequently asked questions
Which OpenAI crawler should I allow for ChatGPT citations?
OAI-SearchBot. It builds the index ChatGPT Search draws from, so allowing it is what makes citations and referrals possible.
What happens if I block GPTBot?
Your content stays out of OpenAI model training. ChatGPT Search visibility is unaffected, because that pipeline uses OAI-SearchBot.
Is ChatGPT-User the same as GPTBot?
No. ChatGPT-User fetches a specific page when a user shares a link or asks about it in real time. It does not crawl or build indexes.
How do I verify a hit really came from OpenAI?
Reverse-DNS the source IP against OpenAI's published verification ranges. User-agent strings alone can be spoofed.
Maya Reinholt
Maya Reinholt leads search research at Toolgram. She has spent nine years in technical SEO, the last three mapping how LLM-powered crawlers and answer engines select, parse, and cite web content.