We have measured two groups of sites the way AI crawlers read them: 104 sites built with AI app builders on 20 September, and 927 x402 Bazaar sites on 27 September. The same five problems came up in both. None of them shows in a browser, so the owner rarely knows. Here is each one, the one-line test, and the fix.
Every test below is a plain curl request: no browser, no JavaScript. That is how GPTBot, ClaudeBot,
OAI-SearchBot and PerplexityBot fetch a page. The full data is in the two measurements:
AI-built sites without JavaScript and
x402 sites and the AI training crawlers.
1. The page is empty without JavaScript
How common: 63 of 104 AI-built sites (61%) had under 300 characters of visible text in their HTML; 69 of 927 x402 sites (7%).
Single-page apps send an almost empty HTML file and build the page in the browser. Google runs JavaScript, so
such a site can rank in Google. The crawlers behind ChatGPT, Claude and Perplexity answers do not run it, so to them
the site is a blank <div id="root"></div>.
Test: fetch the page and look for a sentence you can see on it.
curl -s https://your-site/ | grep -c "a sentence from your page"
0 means the crawlers do not see that sentence either.
Fix: ship the main copy in the HTML itself.
- Lovable or plain Vite: prerender your main routes at build time (vite-plugin-prerender or a
static export), or put the landing copy into
index.html. - Next.js: make the route server-rendered or statically generated, not client-only.
- Nuxt: enable SSR or prerender the route (
nitro.prerender.routes). - Astro: renders on the server by default; check that the page is not one
client:onlyisland.
2. Every address returns 200, even ones that do not exist
How common: 70 of 104 AI-built sites (67%); 43 of 927 x402 sites (5%).
The host answers every path with the app and a 200 OK. Google records these as soft 404s and indexes
copies of your home page, and a broken link never shows up as an error.
Test: ask for a page that cannot exist. The answer should be 404.
curl -s -o /dev/null -w "%{http_code}\n" https://your-site/this-page-does-not-exist
Fix: return a real 404 for unknown paths.
- Netlify: the usual cause is the rewrite
/* /index.html 200. Serve/404.htmlwith status 404 for unknown paths. - Vercel: add a not-found route (
app/not-found.tsxorpages/404.tsx) and make sure no rewrite swallows unknown paths. - Cloudflare Pages: add a
404.htmlat the project root. - Lovable: add a NotFound route that renders a real 404 page.
3. One site on two addresses
How common: among AI-built sites on their own domain, 24 of 60
(40%) served the same page on www and on the bare domain, with no redirect;
7 of 60 (12%) had one of the two not answering at all.
Two copies split the signals between two addresses. A dead one loses everyone who types the other form.
Test: one of the two should answer 301 and point to the other.
curl -sI https://your-site.com/ | head -1
curl -sI https://www.your-site.com/ | head -1
Fix: pick one hostname and 301 the other to it, keeping the path. Point
<link rel="canonical"> at the one you picked. On Cloudflare: a DNS record for both names and a
Redirect Rule.
4. No sitemap
How common: 42 of 104 AI-built sites (40%); 422 of 927 x402 sites (46%).
In a single-page app the routes live inside the JavaScript bundle. Without a sitemap a crawler has no list of them, so it may only ever see your home page.
Test: the answer should be 200, and the body should be XML, not your app.
curl -s -o /dev/null -w "%{http_code}\n" https://your-site/sitemap.xml
Fix: generate sitemap.xml at build time (vite-plugin-sitemap, or app/sitemap.ts
in Next.js) and add a Sitemap: line to robots.txt.
5. The firewall turns AI crawlers away
How common: 86 of 908 x402 sites (9%)
answered GPTBot and ClaudeBot with a 403. 83 of those still let the search and answer crawlers in.
All 86 answered through Cloudflare.
Your robots.txt can say yes while the firewall in front of the site says no. On 15 September 2026
Cloudflare split its AI crawler setting into training, search and agent crawlers, and moved existing sites over on its
own.
Test: request the page as yourself, then with each crawler's complete user-agent string. Use the full string: some firewalls recognise only the full form.
curl -s -o /dev/null -w "%{http_code}\n" https://your-site/
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.1; +https://openai.com/gptbot)" https://your-site/
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot" https://your-site/
200, 403, 200 means training is blocked and search is open. A 403
on the third line is the one to fix: that crawler feeds ChatGPT search.
Fix: on Cloudflare, Security → Settings → AI crawler controls. Set Search (and Agent) to Allow. Training is your choice: allow it if you want models to learn about your service, leave it blocked if you do not.
What this does not cover
These are first-page checks: what one request to your address returns. Whether your pages are actually indexed is in Google Search Console and Bing Webmaster Tools. Whether an AI answer cites you depends on much more than being readable. Being readable is only where it starts.