Skip to content
SilktideHelp

AI crawler access

AI assistants such as ChatGPT, Claude and Perplexity read the web using their own automated crawlers. Silktide checks whether your website tells those crawlers to stay out, which usually means AI assistants cannot read, cite or recommend your content.

Websites control crawler access through a file. A single rule in that file can shut out every major AI system:

# Blocks OpenAI's crawler from the entire site
User-agent: GPTBot
Disallow: \/

Blocking can be a deliberate choice, for example to opt out of model training. Because it is sometimes intentional, Silktide reports this as a warning you can approve rather than a confirmed error.

Why this matters

A growing share of people ask an AI assistant instead of a search engine. Those assistants can only describe, cite and link to websites their crawlers are allowed to read. If your site blocks them, competitors' content fills the answers instead, and you lose visitors without ever seeing an error.

Many robots.txt files block AI crawlers by accident: the rules were copied from a template or a security hardening guide, and nobody realised what they excluded. This check surfaces that so the block is a decision, not an oversight.

How to fix it

  1. Open your site's robots.txt file (at https:\/\/yoursite.com\/robots.txt) and find the rules naming the blocked crawlers, or a User-agent: * rule with Disallow: \/ that blocks everything.
  2. If the block is unintentional, remove those rules or narrow them to the paths you genuinely want private, then republish the file:
# Before: blocks the whole site
User-agent: GPTBot
Disallow: \/

# After: allows the site, keeps one area private
User-agent: GPTBot
Disallow: \/internal\/
  1. If the block is deliberate - for example, you do not want your content used for AI training - the finding. Note that some crawlers feed training and others feed live answers, so you can also block selectively: allowing OAI-SearchBot while blocking GPTBot keeps you visible in ChatGPT search without contributing to training.

How Silktide tests this

  1. Fetch the robots.txt file from your website's root.
  2. If the file does not exist or cannot be fetched, the check passes: no robots.txt means no crawlers are blocked.
  3. Parse the rules and test each of a list of around 20 well-known AI crawler user agents (including GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, Applebot-Extended, and CCBot) against your site root.
  4. Warn when one or more of these crawlers is blocked from the site root, listing the blocked crawlers.

Troubleshooting

The check passes but AI assistants still cannot read my site

This check only reads your robots.txt rules. A crawler that is allowed on paper can still be blocked in practice by a firewall or bot protection - see AI crawlers blocked in practice. Content that only appears after runs can also be invisible to AI crawlers - see Content available without JavaScript.

I block AI crawlers on purpose

That is a legitimate choice. Approve the finding and it will not count against your site. Consider whether you want to block all AI crawlers or only the ones used for model training.

Learn more

Last updated

Was this page helpful?