Home » AI Search (AEO) » AI Crawler Access Audit: Which Bots Can Read Your Website?

AI Crawler Access Audit: Which Bots Can Read Your Website?

Published 2026-10-04 | AI Search (AEO)

For years, "can search engines crawl my site?" meant Googlebot and Bingbot. Now there's a longer list. AI assistants fetch pages to answer questions, to build their own search indexes, and to train models. Each uses a named crawler, and each can be allowed or blocked by you, or by something you forgot about.

An AI crawler access audit tells you what's actually happening on your site today.

The crawlers worth knowing

Names change, so confirm against each vendor's documentation before relying on this list.

Crawler Operator Typical purpose
GPTBot OpenAI Model training
OAI-SearchBot OpenAI Search results in ChatGPT
ChatGPT-User OpenAI Fetching a page on a user's request
ClaudeBot Anthropic Model training
Claude-User Anthropic Fetching pages for a user's question
PerplexityBot Perplexity Building its search index
Google-Extended Google A control for some Google AI uses
Applebot-Extended Apple A control for Apple AI uses
CCBot Common Crawl Open web dataset used by many projects

Step 1: read your robots.txt

Open yourdomain.com/robots.txt and look for lines naming these bots, or a blanket User-agent: * rule. Check three things: is a rule blocking everything, is a section copied from a template that blocks AI by default, and is a more specific rule overriding a general one? The robots.txt guide has the syntax details in robots.txt blocking pages.

Step 2: test beyond robots.txt

A bot can be allowed by robots.txt and still be stopped elsewhere. Look at:

  • Firewall or WAF rules, including bot protection settings on Cloudflare, Sucuri and similar services
  • Rate limiting that returns 429 errors to fast-moving crawlers
  • Hosting security plugins with "block unknown bots" options
  • Geo or IP rules that reject traffic from data-centre ranges

Request one of your pages with the bot's user-agent, for example:

curl -A "GPTBot" -I https://www.example.com/

A 200 response is good. A 403, 429 or a challenge page means something is blocking it, even though robots.txt says it's allowed. Remember that a spoofed user-agent from your own machine tests the rule, but not the bot's real IP, so check your firewall logs too.

Step 3: check what the bot receives

Many of these crawlers don't run JavaScript. If your content depends on scripts, they might see an empty page. Compare the raw HTML to the rendered page, as described in the rendered content gap guide.

Step 4: decide your policy

This is a business decision, not a technical one. Common positions:

  1. Open. Allow search and user-triggered bots, and training bots too, to maximise visibility.
  2. Selective. Allow search and user bots so you can be cited, block training crawlers.
  3. Closed. Block all AI crawlers, accepting you'll likely not appear in their answers.

Whatever you pick, write it down and put it in robots.txt explicitly. Don't leave it to defaults.

Step 5: monitor

Check server logs monthly for AI user-agents. Seeing 200 responses and steady visits confirms access. Seeing none may mean you're blocked, or just not yet discovered.

Common mistakes

  • Blocking all bots on a new site and forgetting to open it up.
  • Allowing the bot in robots.txt while the CDN's bot-fight mode blocks it.
  • Assuming one rule covers every vendor's crawlers.
  • Copying a "block AI" robots.txt snippet from a blog without reading what it does.

You can't be cited by a system that can't read you. Start by checking the door is open.

Frequently Asked Questions

Are training crawlers and search crawlers the same thing?

No. Many AI companies run separate bots for model training, for search indexing, and for fetching a page when a user asks about it. You can often allow one and block another.

Does blocking Google-Extended remove me from Google Search?

No. Google-Extended is a control for whether your content is used for certain Google AI products. It doesn't affect crawling or ranking in regular Search.

How do I know if a bot is real?

Check the user-agent against the vendor's documentation and verify the request IP against their published ranges. Anyone can fake a user-agent string.