AI crawler blocking
Silktide tests whether requests that identify as the major AI crawlers actually get through to your website. A site can welcome these crawlers on paper while a firewall or security product turns them away at the door.
Your can say are welcome while your firewall refuses them. and security products ship one-click "block AI bots" switches - some enabled by default - that refuse any request identifying as GPTBot, ClaudeBot, PerplexityBot, or similar, regardless of what robots.txt declares. Site owners frequently have no idea it's on: to you the site works perfectly, while to every AI assistant it does not exist.
Why this matters
AI assistants can only cite and recommend content their crawlers can fetch. A firewall-level block is invisible in normal use - your site loads fine in every browser, your analytics look healthy - yet answer engines silently fail to read you and quote competitors instead. Because nothing errors visibly, this is one of the hardest visibility problems to notice without testing for it directly.
Silktide's separate AI crawler access check covers what your robots.txt declares. This check covers what actually happens on the wire. The two commonly disagree, and when they do, the firewall wins.
How to fix it
- Check your CDN or security dashboard for AI bot blocking. In Cloudflare, look under Security for "AI bots" or "Bot Fight Mode"; other providers such as Akamai, Imperva, and DataDome have equivalent switches. These are sometimes enabled by default.
- If the block is unintentional, turn the switch off or add the specific crawlers you want to allow to your allow-list.
- If you have your own firewall or server rules matching bot user agents, remove the AI crawler entries you want to allow.
- If you block AI crawlers deliberately, the finding - it is a legitimate choice.
How Silktide tests this
- Request your homepage with a normal browser identity. If that request is itself refused or challenged, the check is not applicable - nothing can be attributed to crawler discrimination.
- Request the homepage again identifying as each major AI crawler (GPTBot, ClaudeBot, PerplexityBot, and CCBot), using the User-Agent strings their operators publish. All requests come from the same place, so any difference in response is down to the crawler identity.
- Count a crawler as blocked when its request receives an outright refusal
status (for example
403 Forbidden,429 Too Many Requests, or503), a challenge or page a crawler cannot solve, or no response at all while the browser request succeeded. - Warn when any crawler is turned away, listing which ones.
All requests target only the homepage and run once per test.
Troubleshooting
We verify crawlers by their network, not just their name
Some sites admit only requests that come from a crawler operator's verified network addresses. Silktide's probe sends the crawler's published identity from its own network, so such a site may refuse the probe while allowing the real crawler. This setup is much rarer than blanket identity blocking; if you have confirmed you use it, approve the finding.
The check passes but assistants still miss my content
Getting through the door is necessary but not sufficient. Check that your robots.txt allows AI crawlers (AI crawler access) and that your content does not require to appear (Content without JavaScript).