Payload

← All guides

Are AI crawlers blocked by your robots.txt?

AI search engines and assistants can only cite pages their crawlers are allowed to fetch. One wrong Disallow line can make your whole site invisible to them.

Which crawlers matter

User-agent tokenOperator
GPTBotOpenAI
PerplexityBotPerplexity
ClaudeBotAnthropic
BytespiderByteDance
Applebot-ExtendedApple
Meta-ExternalAgentMeta
Google-ExtendedGoogle

How to read User-agent groups

A robots.txt is a sequence of groups. Each group starts with one or more User-agent lines, followed by the rules (Allow / Disallow) that apply to those agents — until the next User-agent line starts a new group.

The most-specific-match rule: a crawler uses the single group whose User-agent token is the longest match for its own user-agent string. A group for GPTBot beats the * group for GPTBot — even if the * group appears later in the file. Within a group, the longest matching path wins; on a tie, Allow beats Disallow per Google's documented convention, which most AI crawlers follow.

Worked example

User-agent: *
Disallow: /admin/
Disallow: /drafts/

User-agent: GPTBot
Disallow: /

User-agent: PerplexityBot
Allow: /

Resulting matrix for https://example.com/blog/post:

CrawlerStatusWhy
GPTBotBlockedIts own group says "Disallow: /" — the specific group beats *.
PerplexityBotAllowedIts own group says "Allow: /".
ClaudeBotAllowedNo ClaudeBot group, so the * group applies: only /admin/ and /drafts/ are disallowed.
BytespiderAllowedSame as ClaudeBot: falls through to the * group.

The trap in this example: the site owner blocked GPTBot everywhere — perhaps copied from a template — while intending to be AI-friendly. The * group does not override it.

Check your own file

Paste your robots.txt into the free robots.txt checker for AI crawlers — it runs entirely in your browser and shows exactly which AI crawlers are allowed, blocked, or unmentioned, quoting the rule that decided each one.

Common failures

Go further

robots.txt is one of four gates. The AI Search Readiness Audit ($59, one-time, Chrome extension, 100% client-side) also checks raw-HTML coverage, AI Overviews eligibility, and dead schema — the full technical picture of whether AI engines can see, read, and cite your pages.

Built by Payload

Payload builds practical software that makes AI, automation, and business infrastructure safer, cleaner, more reliable, and easier to ship. Support: kylers.partners@gmail.com