Can AI Read Your Website? We Checked 9,066 Business Sites

Can AI Read Your Website? We Checked 9,066 Business Sites

We checked 9,066 business websites in four markets. About 1 in 13 cannot be read properly by ChatGPT, Claude or Perplexity, most often because of the homepage.

By Yannis Spanenburg · · 8 min read

We ran our AI crawler checker on 9,066 websites of dental practices, clinics, law firms, accountants and estate agents in Belgium, Dubai, the United Kingdom and New York State. About 1 in 13 failed.

What Does It Mean When AI Cannot Read Your Website?

It means the crawler behind ChatGPT search, Claude or Perplexity cannot get the text of your pages, so the AI has nothing from your site to quote or recommend. The cause is a homepage that shows almost no text without JavaScript, a firewall that turns the crawler away, or a robots.txt rule that keeps it out.

  • In Proof of Pixel's October 2026 study, 7.9% of 9,066 business websites could not be read properly by at least one of ChatGPT, Claude or Perplexity.
  • The most common cause was the homepage: 4.5% of the sites showed AI crawlers fewer than 50 words of text.
  • 2.8% had a firewall that turned an AI search crawler away, and 1.1% kept one out through robots.txt.
  • A further 2.7%, almost all of them UK medical practices, answer every visitor that does not run JavaScript with a challenge page.
  • 16.4% of the sites had an llms.txt file, a kind of file Google says you do not need for its AI features.

How Many Business Websites Can AI Not Read?

About 1 in 13. Of 9,066 business websites that loaded for a normal visitor, 715 (7.9%) could not be read properly by at least one of ChatGPT, Claude or Perplexity. New York State had the highest share at 12.9% and the United Kingdom the lowest at 6.7%.

By market:

  • New York State: 12.9% of 1,022 sites
  • Dubai: 10.1% of 79 sites
  • Belgium: 9.2% of 1,760 sites
  • United Kingdom: 6.7% of 6,205 sites, plus 3.9% behind a JavaScript challenge

By trade:

  • Accounting: 10.4%
  • Real estate: 9.9%
  • Legal: 7.5%
  • Dental: 7.2%
  • Medical: 6.3%, plus 8.7% behind a JavaScript challenge

Why Can't AI Crawlers Read Some Websites?

The main reason is a homepage that holds almost no text in its HTML. On 4.5% of the sites, AI crawlers saw fewer than 50 words, and on 154 of them no words at all. The crawlers behind ChatGPT search, Claude and Perplexity do not run JavaScript, according to an analysis of real crawler traffic by Vercel and MERJ.

Some of these homepages are built entirely in JavaScript. Others are a photo, a slideshow, a frame or a splash page with an Enter button. A visitor sees a full page, while the crawler gets a title and a few menu items. Google and Bing render JavaScript, so such a site can rank in normal search and still be missing from AI answers.

In our client work rebuilding websites for practices and service businesses, this is the gap we find most often: the page looks complete in a browser and the HTML underneath holds a menu. Our own site is a React app, and we render every page to plain HTML at build time for this reason.

Do Firewalls Block AI Crawlers?

Sometimes. On 2.8% of the sites, the server turned away a request sent as an AI search crawler while a normal browser got in. We retested each of these one request at a time before counting it. Cloudflare was in front of 44 of the 255 sites; the rest sat behind other firewalls and hosts. New York stood out again at 5.9%.

What Is a JavaScript Challenge, and Why Does It Matter for AI?

A JavaScript challenge is a firewall check that sends every visitor a small script to run before the page appears. A browser runs it in a second and the visitor sees the site. A crawler that does not run JavaScript gets an empty page.

243 sites in the study answered this way, 240 of them UK medical practices, and they did it for our browser request as well as for the AI crawlers. A real crawler arriving from its own network may be let through, so we report these sites apart from the 7.9%. If your practice is one of them, ask your website provider whether OAI-SearchBot, Claude-SearchBot and PerplexityBot are on the firewall's allowed list.

Which AI Crawlers Do Websites Block?

Perplexity's crawler is blocked most often. Of the 9,066 sites, 94 kept PerplexityBot out through robots.txt or a failing robots.txt, against 67 for ChatGPT's OAI-SearchBot and 62 for Claude-SearchBot. In total only 1.1% of the sites blocked an AI search crawler this way.

New York stood out at 2.9%. Several practice websites there share one robots.txt template that lists the crawlers allowed in, ChatGPT's and Claude's among them, then turns away everyone else. Perplexity's search crawler is missing from that list, so it is shut out without anyone deciding to block it.

Blocking a training crawler is a separate choice. 2.0% of the sites block GPTBot, which collects data to train OpenAI's models, and 142 of those 178 still let OAI-SearchBot in. OpenAI states that each of its robots.txt settings is independent of the others, so those sites can still appear in ChatGPT search. We covered the trade-off in how to block AI crawlers.

Does an llms.txt File Help AI Read Your Site?

Not for Google. Google Search Central states that you do not need AI text files to appear in AI Overviews or AI Mode. Even so, 16.4% of the sites in our study had an llms.txt file, and a third of the sites in Dubai. 110 sites with one still could not be read properly by an AI crawler, because their homepage, their firewall or their robots.txt kept the content out. More in our llms.txt guide.

How to Check If AI Can Read Your Website

The check takes about five minutes and needs only a browser. It covers the causes we found: the text in your HTML, your robots.txt rules and your firewall.

  1. Open your homepage, right-click and choose View page source.
  2. Search the source for the first sentence of your homepage. If it is missing, AI crawlers cannot read that text.
  3. Open yoursite.com/robots.txt and search for OAI-SearchBot, Claude-SearchBot and PerplexityBot.
  4. Look for a final "User-agent: *" group with "Disallow: /", which shuts out every crawler the file does not name.
  5. Run the free AI crawler checker to test your firewall as each crawler.

Proof of Pixel is a marketing and consulting agency for businesses that live on enquiries and orders, building the website, the tracking and the follow-up as one system, with offices in Dubai, New York, London, Antwerp and Malaysia.

Proof of Pixel's position: every page a business wants quoted by AI must show its full text in the HTML, readable by a crawler that does not run JavaScript.

How We Ran the Study

We took every business tagged with a website in OpenStreetMap in five trades (dentists; clinics and doctors; lawyers and notaries; accountants and tax advisors; estate agents) in Belgium, Dubai, the United Kingdom and New York State. We left out public-sector sites and chains, checked 11,594 websites in October 2026 and counted the 9,066 that loaded for a browser.

For each homepage the checker read robots.txt as RFC 9309 prescribes, fetched the page as a browser and as each AI crawler, and counted the words a crawler sees without JavaScript. Every firewall refusal was retested one request at a time. Limits: OpenStreetMap does not list every business, the United Kingdom makes up two thirds of the sites and Dubai only 79, we checked homepages only, and real crawlers come from their own IP addresses, so the firewall figures are the least certain. We name no business.

Where to Start This Week

Open your homepage source and search for your first sentence. If it is there, check robots.txt for the three AI search crawlers. If both pass, run the checker once to rule out the firewall.

Frequently asked questions

Can ChatGPT read my website?

ChatGPT can read your website if your robots.txt allows OAI-SearchBot, your firewall lets it in and your homepage shows its text without JavaScript. In our study of 9,066 business websites, about 1 in 13 failed one of those checks for ChatGPT, Claude or Perplexity.

Why can't AI crawlers read some websites?

AI crawlers most often fail because the page holds almost no text in its HTML. The crawlers behind ChatGPT search, Claude and Perplexity do not run JavaScript, so a page built in JavaScript or made of images shows them a few words. Less often, a firewall or robots.txt turns them away.

Does blocking GPTBot remove my site from ChatGPT?

Blocking GPTBot does not remove your site from ChatGPT search. GPTBot collects training data, and ChatGPT search uses OAI-SearchBot. OpenAI states that each of its robots.txt settings is independent of the others.

Do I need an llms.txt file?

You do not need an llms.txt file for Google, which says no AI text files are needed to appear in its AI features. In our study, 1 in 6 business websites had one, and 110 of those still could not be read by an AI crawler. Make the page itself readable first.

How do I check if AI can read my website?

View your homepage source and search for your first sentence, then check robots.txt for OAI-SearchBot, Claude-SearchBot and PerplexityBot. The free AI crawler checker at proofofpixel.agency/ai-crawler-checker runs both checks and tests the firewall as each crawler.

Keep reading

Where this fits