August 11, 2026
robots.txt for GPTBot, ClaudeBot & Perplexity
Your GPTBot robots.txt rules may be blocking ChatGPT and Claude without you knowing. How to check ClaudeBot and PerplexityBot too — and fix them.
Your robots.txt file controls which AI crawlers are allowed to read your site — and a surprising number of sites block GPTBot, ClaudeBot, and PerplexityBot without their owners ever making that decision. The block usually comes from a CDN default, a security plugin, or a boilerplate file copied years ago. If your GPTBot robots.txt rule says Disallow: /, ChatGPT's crawler cannot fetch your pages, and your content stays out of its answers. Here is how to check, and how to fix it.
The AI crawlers that matter
There are four AI-related user agents worth knowing about, and they do not all do the same thing. Some crawl your site to train models; others crawl it to answer user questions in real time. Blocking one does not block the other.
| User agent | Company | What it reads your site for |
| --- | --- | --- |
| GPTBot | OpenAI | Training data for ChatGPT models |
| OAI-SearchBot | OpenAI | Indexing pages for ChatGPT Search answers |
| ChatGPT-User | OpenAI | Fetching a page when a ChatGPT user asks about it live |
| ClaudeBot | Anthropic | Training and retrieval for Claude |
| PerplexityBot | Perplexity | Indexing and fetching pages for Perplexity answers |
| Google-Extended | Google | Training Gemini (does not affect Search rankings) |
The practical takeaway: retrieval bots (OAI-SearchBot, ChatGPT-User, PerplexityBot) are the ones that decide whether you show up in AI answers today. GPTBot and ClaudeBot matter more for whether your content shapes future models. If you want visibility in AI answers, the retrieval bots are the ones you cannot afford to block.
How to check your own robots.txt in 30 seconds
Your robots.txt is a public text file, and checking it takes under a minute.
- Open
https://yourdomain.com/robots.txtin your browser. - Search the page for
GPTBot,ClaudeBot,PerplexityBot,OAI-SearchBot, andChatGPT-User. - For each one you find, look at the
Disallowline directly under it.
A block looks like this:
User-agent: GPTBot
Disallow: /
Disallow: / means "this bot may not read anything on the site." An empty Disallow: means "read everything." If you find no mention of AI bots at all, they are allowed by default — unless a wildcard rule (see the FAQ) says otherwise.
If you'd rather not eyeball the file, check if ChatGPT can read your site walks through it, or just scan your site free and the audit parses your robots.txt per AI bot for you.
Why sites block AI crawlers without knowing
Most accidental AI blocks come from one of three places — none of them a deliberate decision by the site owner.
- CDN and hosting defaults. Some managed hosts and CDN security layers ship with "block known AI bots" toggles, or generate a robots.txt on your behalf. You may never have seen the file your site actually serves.
- Security and SEO plugins. WordPress plugins and similar tools sometimes add AI-bot disallow rules as part of a "block bad bots" feature, grouping GPTBot in with scrapers.
- Copied boilerplate. In 2023, a widely-shared snippet for blocking AI training made the rounds. Plenty of sites pasted it into their robots.txt and forgot about it. The file outlived the reason.
The common thread: robots.txt is set-and-forget infrastructure. Nobody gets notified when it blocks something valuable. For a deeper look at the access problem, this is one of the three levers covered in the GEO audit — alongside extractability and citability, which an llms.txt file helps with.
Should you block AI crawlers?
It depends on what you're protecting and what you're giving up — and the honest answer is that there is a real trade-off, not an obvious right answer.
Reasons to block: you don't want your content used as training data, your content is your product (paid courses, proprietary docs), or you're concerned about server load from aggressive crawling.
Reasons to allow: blocking retrieval bots makes you invisible in AI answers. When a potential customer asks ChatGPT or Perplexity for a tool like yours, a blocked site simply isn't a candidate for citation. For most solo founders and indie devs, that invisibility costs more than the training-data use does.
One nuance worth understanding: blocking training and blocking retrieval are separate decisions. You can disallow GPTBot (training) while allowing OAI-SearchBot and ChatGPT-User (search and live fetching). OpenAI documents these as separate agents, so you can opt out of training without disappearing from ChatGPT Search. Whether that distinction will hold forever is unknown — but today it is a real, supported option.
How to allow AI crawlers
To explicitly allow the major AI crawlers, add one block per bot with an empty Disallow (or a specific Allow). A clean, explicit setup looks like this:
User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: PerplexityBot
Allow: /
If you'd rather keep training bots out but stay visible in AI answers, use this instead:
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: PerplexityBot
Allow: /
Place these blocks above any wildcard (User-agent: *) rules. Then re-open your live robots.txt and confirm the change actually deployed — if your CDN manages the file, your edit may need to happen in their dashboard, not in your repo.
FAQ
Does Disallow: / under User-agent: * block AI crawlers?
Yes. A wildcard rule with Disallow: / blocks every bot that doesn't have a more specific rule — including GPTBot, ClaudeBot, and PerplexityBot. If your file has a blanket wildcard disallow, AI crawlers are blocked unless you add explicit per-bot allows.
If I have both Disallow and Allow rules, which one wins?
The most specific matching rule wins, and for the major crawlers a longer, more precise path beats a shorter one. Per-bot blocks (User-agent: GPTBot) override the wildcard block entirely. Google, OpenAI, and Anthropic all follow this longest-match convention, though edge cases exist — when in doubt, keep rules simple and explicit.
Does blocking AI crawlers hurt my Google rankings?
No. GPTBot, ClaudeBot, and PerplexityBot are separate from Googlebot, and blocking them has no effect on Search rankings. Google-Extended controls Gemini training only and explicitly does not affect Search either, per Google's own documentation.
What's the difference between OAI-SearchBot and GPTBot?
GPTBot crawls your site to gather training data for OpenAI's models, while OAI-SearchBot indexes pages so ChatGPT Search can cite them in answers. Blocking GPTBot keeps you out of future training runs; blocking OAI-SearchBot keeps you out of ChatGPT Search results today. They are controlled independently.
What if my CDN manages my robots.txt and I can't edit it?
Then the file is generated or proxied by your CDN, and edits must happen in their settings, not in your codebase. Check your CDN or hosting dashboard for a robots.txt or bot-blocking option — some platforms let you customize the file or toggle AI-bot access per bot. If the platform offers no override, contact their support, because the block is their default, not your choice.
Your robots.txt is a five-line file that quietly decides whether AI assistants can see you at all. Check it once, fix it if needed, and scan your site free to confirm GPTBot, ClaudeBot, and PerplexityBot can actually reach your pages — along with the rest of the GEO audit.