AI crawlers blocked in robots.txt
What it is
Your robots.txt file tells one or more AI crawlers not to crawl any part of the site. AI companies use three kinds of crawler, each with its own name in robots.txt: search crawlers that index pages so an assistant can cite them, user-request fetchers that open a page when someone asks about it, and training crawlers that collect content for future models. You can allow one kind and block another.
Why it matters
Blocking an AI search crawler (OAI-SearchBot, Claude-SearchBot, PerplexityBot) keeps your pages out of that assistant’s answers and citations. Blocking a training crawler only opts your content out of model training, which can be deliberate. Google-Extended is the exception: it also controls whether Gemini uses your pages to ground its answers . Blocks are often accidental, from a security plugin, CDN or copied file. Naming AI crawlers does not affect Google or Bing rankings; a User-agent: * block does, and is reported separately.
How Asky checks it
Asky reads your robots.txt and, for each AI crawler below, applies that crawler’s rules: the groups that name it, or the * group when none does. A crawler is reported when those rules block the whole site, as Disallow: / does, even if Allow rules reopen some sections. Path-level blocks are not reported. When Allow and Disallow both match, the longest rule wins, Allow on a tie.
- Search and user-request: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User
- Training: GPTBot, ClaudeBot, Google-Extended, Cohere-AI, CCBot, and the retired Claude-Web and Anthropic-AI
Reported as a warning with medium severity, listing the blocked crawlers.
How to fix it
- Open
https://example.com/robots.txtfor your own domain (each subdomain has its own file) and find the groups that name the crawlers in the issue, or aUser-agent: *group withDisallow: /. - Decide per crawler. To appear in AI answers, allow at least the search crawlers. Blocking training crawlers is your choice. OpenAI and Perplexity say their user-request fetchers may not follow robots.txt, so blocking those is not a reliable opt-out.
- Remove the
Disallow: /rule for the crawlers you want to allow, or give them their own group withAllow: /. A group that names a crawler replaces the*group for it, so repeat any*rules you still want. - If a plugin, a CDN (such as Cloudflare’s managed robots.txt) or a hosting panel writes the file, change it there. Bot blocking in a firewall or CDN never shows in robots.txt, so check those settings too.
- To leave it as it is: if the blocked crawlers are training crawlers you chose to exclude, mark the issue resolved.
Example
This robots.txt on https://example.com blocks every AI crawler it names, including the ones that decide whether the site appears in AI answers:
# Before: search and training crawlers all blocked
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ClaudeBot
User-agent: Claude-SearchBot
Disallow: /
# After: training crawlers blocked, AI search allowed
User-agent: GPTBot
User-agent: ClaudeBot
Disallow: /
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
Allow: /
Sitemap: https://example.com/sitemap.xml