Home » AI Search (AEO) » AI-Blocked Content: Pages People Can Read but AI Crawlers Can't

AI-Blocked Content: Pages People Can Read but AI Crawlers Can't

Published 2026-10-04 | AI Search (AEO)

The previous post covered whether AI crawlers can reach your site at all. This one is narrower. It deals with the pages that look perfectly open to a person but are closed to specific bots. The mismatch is easy to miss, because you test your site in a browser and the browser always gets through.

How a page ends up blocked for bots only

Mechanism What happens
robots.txt rule A bot is told not to fetch a path or the whole site
WAF or CDN bot rule The bot gets a 403, or a challenge page, even if robots.txt allows it
Rate limiting A fast bot receives 429 errors after a handful of requests
JavaScript challenge A "checking your browser" screen appears before content
Cookie or consent wall Content loads only after a click on a banner
Soft paywall The first paragraph shows and the rest is hidden or removed
Login requirement Content needs an account
X-Robots-Tag header A server-level directive restricts use
Geo restriction Requests from some regions are refused

Step 1: list the pages that matter

Choose the pages you'd want to be quoted from: key service pages, guides, pricing, FAQs, case studies. Don't try to audit everything at first.

Step 2: request them as a bot would

For each of these pages, make requests with the user-agent strings of the main crawlers and record the result. Use curl -A "OAI-SearchBot" -I https://www.example.com/page/ and similar commands. Log the status code and note any redirects.

A comparison table is helpful:

Page Browser Googlebot OAI-SearchBot PerplexityBot ClaudeBot
/pricing/ 200 200 403 200 403
/guide/ 200 200 200 200 200

Any row where the browser gets in and a bot doesn't is a block.

Step 3: find the source of each block

Work from the outside in:

  1. Check your CDN or firewall dashboard for rules and recent blocked requests.
  2. Read the response body of a failed request. A challenge page often says which service produced it.
  3. Review robots.txt rules by path.
  4. Check response headers for X-Robots-Tag.
  5. Look at the hosting panel for bot-protection toggles.

Step 4: check what the bot sees when it does get in

Fetch the raw HTML and read it. If the content is hidden behind a consent overlay or only partly present, bots will see that version. See the rendered content gap post for how to compare the two.

Step 5: decide, then adjust

For each blocked page, answer honestly: did we mean to do this?

  • Yes, deliberately. Paid content, private data, licensed material. Leave it and keep a note of why.
  • No, by accident. Add an allow rule for verified crawlers, loosen rate limits, or exclude those paths from the challenge.

If you allow bots through a firewall, verify them by published IP ranges where the vendor provides them, rather than by user-agent alone.

Common mistakes

  • Switching on "block AI bots" in a CDN with one click and never checking what's blocked.
  • Forgetting that bot protection also blocks legitimate search crawlers.
  • Treating a partial paywall as invisible to bots when the full text is actually in the HTML.
  • Not re-testing after a hosting or CDN change.

The goal isn't to let everything in. It's to make sure every block is one you chose.

Frequently Asked Questions

What is the noai tag?

Some sites add 'noai' or 'noimageai' directives to meta robots or headers. These are informal conventions, not an official standard, and support varies, so don't rely on them as your only control.

Will a CAPTCHA block AI crawlers?

Usually yes. Most crawlers can't solve challenges, so a page behind a CAPTCHA or a JavaScript check is effectively closed to them.

Should members-only content be open to AI crawlers?

Only if you're happy for it to be quoted. Content behind a login is generally unreachable anyway, and open previews of it should be chosen deliberately.