August 11, 2026
Can ChatGPT Read Your Site? 3 Free Checks
Three checks tell you if ChatGPT, Claude and Perplexity can read your site: robots.txt, llms.txt and renderable content. Here's how to run them.
How do I check if ChatGPT can read my website? Run three checks: look at your robots.txt for bot rules that block AI crawlers, test whether /llms.txt exists, and confirm your content shows up in the raw HTML without JavaScript. Each check takes under two minutes, and below is exactly how to run them — plus a fourth check for the title and meta description that assistants quote when they do cite you.
Check 1: robots.txt
Open https://yourdomain.com/robots.txt in a browser and read the bot rules. This file tells crawlers what they're allowed to fetch, and AI assistants respect it — a Disallow rule for GPTBot means ChatGPT won't read your site at all.
Blocking looks like this:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
What you want to see is either no mention of AI bots at all (silence means allowed) or explicit Allow rules. These are the user agents that matter:
| User agent | Belongs to | Used for |
| --- | --- | --- |
| GPTBot | OpenAI | Training and crawling for ChatGPT |
| OAI-SearchBot | OpenAI | ChatGPT's search feature |
| ClaudeBot | Anthropic | Crawling for Claude |
| PerplexityBot | Perplexity | Perplexity's answer engine |
| Google-Extended | Google | Control token for Gemini training (not a crawler itself) |
One trap: the block may not be in your file at all. CDNs like Cloudflare offer one-click AI bot blocking at the network edge, which stops the bots before robots.txt is even read. If your file looks clean but the bots still can't reach you, check your CDN's security or AI-crawler settings.
Check 2: /llms.txt
Visit https://yourdomain.com/llms.txt. This is a markdown file that gives AI models a clean, structured summary of your site — who you are, what you do, and which pages matter.
A 404 here does not mean you're blocked. It means you're missing an opportunity: nothing is stopping the assistants, but you're not handing them a map either. If the URL returns plain markdown describing your product, you're in good shape. If it 404s, creating one takes about ten minutes — see our guide to llms.txt.
Check 3: the view-source test
Your content must exist in the raw HTML, because many AI crawlers fetch the page without running JavaScript. If your text only appears after scripts execute, the crawler sees an empty shell.
Two ways to test:
- Browser: right-click your page, choose "View page source" (not "Inspect" — that shows the post-JavaScript DOM), and search for a sentence from your page. If you find it, the content renders without JS.
- Terminal: run
curl -s https://yourdomain.com | grep "a sentence from your page". A match means the text is in the raw HTML.
If neither finds your content, your site is client-rendered — common with React, Vue or Next.js apps that skip server-side rendering. Fixing that usually means turning on SSR, static generation or prerendering, depending on your stack.
Check 4: title and meta description
Check that every important page has a unique <title> tag and meta description in the raw HTML. When assistants quote or summarize a page, these are often the first strings they lift — and they're also what shows when your page appears in an AI search result.
Same test as Check 3: view source and look for <title> and <meta name="description" near the top. They should describe the page in plain language, not a keyword list. If they're missing, duplicated across pages, or stuffed in by JavaScript after load, that's a fix worth making.
The lazy way: one free scan that does all four
If you'd rather not run four manual checks, scan your site free and it runs them for you in one pass. The scan parses your robots.txt per AI bot, checks whether /llms.txt exists, measures how much of your content requires JavaScript, and verifies your title and meta coverage.
What you get back is honest and limited: a set of scores for access, extractability and citability, plus the single biggest issue holding you back. It won't promise rankings or citations — it tells you what's broken and where to start, which is the part most people get stuck on.
What to do when a check fails
Every failure above has a fix measured in minutes, not weeks:
- robots.txt blocks AI bots: remove the
Disallowrules for the bots you want to allow, or add explicitAllowrules. The full per-bot syntax is in our guide to robots.txt for AI crawlers. If the block comes from your CDN, flip its AI-crawler setting instead. - /llms.txt returns 404: write a short markdown file describing your product and key pages, then serve it at the root. Ten minutes of work.
- Content needs JavaScript: enable server-side rendering or static generation in your framework, or add a prerendering step. This is the one fix that can take longer than minutes on a legacy stack, but for most modern frameworks it's a config change.
- Missing title or meta: add them to your page templates. If you use a CMS, this is usually a settings field, not code.
One honest caveat: passing all four checks means assistants can read you — it doesn't mean they will cite you. Access is the entry ticket. What gets you actually mentioned is a separate game of clear positioning, quotable content and real presence across the web. Once your checks pass, the next step is to turn access into citations.
FAQ
How often should I re-check?
Re-check after any change to your hosting, CDN, CMS or robots.txt, and roughly once a quarter otherwise. Blocks often appear silently — a platform update or a new security default can flip your AI access overnight.
Does access guarantee citations?
No. Access means ChatGPT, Claude and Perplexity are technically able to read your site; whether they mention you depends on your content, your positioning and how the web talks about you. Exactly how each model picks sources isn't public, so anyone promising guaranteed citations is guessing.
What about Claude and Perplexity — do the same checks apply?
Yes. robots.txt, /llms.txt, renderable HTML and title/meta work the same way for ClaudeBot and PerplexityBot as for GPTBot — the table in Check 1 lists their user agents. Run the checks once and they cover all three assistants.
My site is on Wix, Shopify or WordPress — can I still do this?
Yes. All four checks work the same way, and all three platforms let you edit robots.txt (Shopify via robots.txt.liquid, WordPress via SEO plugins, Wix via its SEO settings). For /llms.txt, upload a static file or add it through your platform's file or redirect tools.
What if my robots.txt is served by a CDN?
Then you have two places to check: the file itself and the CDN's bot settings. Some CDNs serve a default robots.txt or block AI crawlers at the edge regardless of what your file says — if your origin file looks clean, the CDN dashboard is the next stop.
Access problems are common, cheap to fix, and invisible until you look. Run the four checks once and you'll know exactly where you stand — and if you'd rather skip the manual work, the free scan does it in one pass.