Can ChatGPT read your site?

AI answers are built from the pages their crawlers can read. Check in seconds whether ChatGPT, Claude, Perplexity, Gemini, Google's AI Overviews and Copilot can read yours, and get the exact fix if they can't.

If AI can't read your site, it answers without it.

  • 0 AI crawlers that run JavaScript. The major AI crawlers read the HTML and skip the scripts. Text that needs JavaScript to appear is invisible to them. (Source: Vercel and MERJ, measured on real crawler traffic.)
  • 1 setting can block them all. One AI bot policy on Cloudflare can turn every AI crawler away from a whole site. Cloudflare changed those settings on 15 September 2026. (Source: Cloudflare.)
  • 24h for ChatGPT to see a fix. OpenAI says ChatGPT picks up a robots.txt change in about a day, and many blocks take minutes to fix. (Source: OpenAI crawler documentation.)

What you get

  1. Your score, engine by engine. Which of the six AI engines can read your page, and what stops the others.
  2. The exact fix. The robots.txt lines to paste, the setting to switch, or the page change to make.
  3. The AI search guide, free. Four pages, seven checks to get named in ChatGPT, Perplexity and Google's AI answers. Straight to your inbox.

Blocking AI is a choice. Make sure it's yours.

Big publishers block AI crawlers on purpose. Small businesses usually do it by accident, through a plugin, a CDN setting or a site builder. Checked with this tool on 26 September 2026:

  • The New York Times: 2 of 6. Blocks ChatGPT, Claude, Perplexity and Gemini.
  • BBC: 2 of 6. Blocks ChatGPT, Perplexity and Gemini, keeps its pages out of Copilot.
  • The Guardian: 3 of 6. Blocks Claude and Perplexity, keeps its pages out of Copilot.
  • proofofpixel.agency: 6 of 6. Readable by all six. It's how we build.

How the check works

  1. Reads your robots.txt. Exactly the way each AI crawler reads it, following the official standard (RFC 9309).
  2. Knocks as each crawler. Asks for your page as a browser, then as OAI-SearchBot, GPTBot, Claude-SearchBot, ClaudeBot and PerplexityBot, to catch firewalls and CDN rules.
  3. Turns JavaScript off. Reads the page the way AI crawlers do: the raw HTML, no scripts.
  4. Checks the hidden tags. noindex and nosnippet for Google; noindex, noarchive and nocache for Bing and Copilot.

It checks access, not quality: whether the crawlers can reach and read the page. Whether AI engines quote you depends on your content, and on what other sites say about you.

The answer engines

  • ChatGPT. OAI-SearchBot finds pages for ChatGPT search. ChatGPT-User opens a page when someone asks, and OpenAI says robots.txt may not apply to it.
  • Claude. Claude-SearchBot indexes pages for Claude's search. Claude-User opens a page when someone asks, and follows robots.txt.
  • Perplexity. PerplexityBot indexes pages for Perplexity's answers. Perplexity-User opens a page when someone asks, and generally ignores robots.txt.
  • Google AI Overviews. AI Overviews and AI Mode are part of Google Search: Googlebot reads the page, noindex and nosnippet keep it out. Search Console also has an AI switch only you can see.
  • Gemini. Gemini answers from Google's index. Google-Extended decides whether Gemini may use your pages, for training and for its answers.
  • Microsoft Copilot. Copilot answers from Bing's index, which Bingbot builds. noarchive keeps a page out of Copilot, nocache limits it to the title and snippet.

The crawlers

Search crawlers decide whether you show up in AI answers. Training crawlers only collect text for future models.

  • OAI-SearchBot (OpenAI, search): Finds and indexes pages for ChatGPT search.
  • ChatGPT-User (OpenAI, on request): Opens a page when a ChatGPT user asks about it. OpenAI says robots.txt may not apply.
  • GPTBot (OpenAI, training): Collects text to train OpenAI's models. Independent of ChatGPT search.
  • Claude-SearchBot (Anthropic, search): Indexes pages to improve Claude's search results.
  • Claude-User (Anthropic, on request): Opens a page when a Claude user asks about it. Follows robots.txt.
  • ClaudeBot (Anthropic, training): Collects text that may be used to train Anthropic's models.
  • PerplexityBot (Perplexity, search): Indexes and links pages in Perplexity's answers. Not used for training.
  • Perplexity-User (Perplexity, on request): Opens a page when a Perplexity user asks about it. Generally ignores robots.txt.
  • Googlebot (Google, search): Google Search, including AI Overviews and AI Mode.
  • Google-Extended (Google, training): Decides whether Gemini may train on your pages and use them in its answers. Doesn't affect Google Search.
  • Bingbot (Microsoft, search): Bing's index, which Copilot answers from.
  • Applebot-Extended (Apple, training): Opts your pages out of Apple's AI training. It doesn't crawl, and doesn't affect Siri or Spotlight.
  • CCBot (Common Crawl, training): Builds Common Crawl's open archive of the web, which anyone can download, AI companies included.
  • meta-externalagent (Meta, training): Collects content to train Meta's AI models and to index for its products.

Questions

Does blocking GPTBot remove my site from ChatGPT?

No. GPTBot only collects training data, and OpenAI says each of its crawler settings is independent. ChatGPT search finds pages through OAI-SearchBot, which you allow or block separately. Anthropic splits it the same way (ClaudeBot for training, Claude-SearchBot for search). Google is the exception: Google-Extended covers both Gemini's training and its answers.

Do I need an llms.txt file?

No. It's an optional summary file. Google says Search doesn't use it and that you don't need special AI text files or markup, and OpenAI, Anthropic and Perplexity haven't said their crawlers read one. Your robots.txt and the text in your HTML matter far more.

Why does JavaScript matter?

The AI crawlers read the HTML your server sends and don't run JavaScript: Vercel and MERJ measured this across hundreds of millions of crawler requests. If your text only appears after scripts run, as on many site builders and single-page apps, they see a nearly empty page. Google and Bing do run JavaScript, so a page can rank in Google and still be invisible to ChatGPT.

Why would my site block AI crawlers without me knowing?

Usually a CDN or security setting. Cloudflare's AI bot policies can block these crawlers across a whole site, and since September 2026 its blocking can also reach Googlebot and Bingbot on pages with ads. Hosting firewalls and security plugins sometimes block bots they don't recognise. The check asks for your page as the AI crawlers do, so blocks aimed at them show up.

How long until a fix shows up?

OpenAI and Perplexity say robots.txt changes take effect within about 24 hours; the others recrawl on their own schedule. Being readable makes your pages eligible to be cited. Whether they are depends on your content and on what other sites say about you.

Is it free, and what happens to my email?

It's free. We use your email to send the report and the AI search guide, and add you to The Revenue Layer, our weekly email, which you can leave with one click. We keep your email and the site you checked, and never sell or share them. A result is cached for 30 minutes, so checking again doesn't hit your site twice.

Read next: How to optimize your website for ChatGPT and Perplexity