AI crawler access policy
An AI crawler access policy is a deliberate (and sometimes firewall) configuration that says which AI systems may fetch your public pages - and for which purpose. It is situational because the "right" answer depends on whether you care more about being cited in assistant answers, keeping content out of model training, or both. What is not situational is treating a pasted "block all AI" template as if it were a neutral security hardening step. That choice has marketing consequences.
Throughout this page, suppose Fernwood wants to appear in ChatGPT and Claude answers about expense management, but its legal team prefers not to contribute new pages to foundation-model training corpora.
This technique is about access. Once a crawler is allowed in, it still has to receive real HTML - see Content readable without JavaScript - and an optional curated map is covered separately in llms.txt.
Why this is situational
AI products no longer use a single "AI bot." Major vendors split crawlers by job:
| Job | Typical bots | What blocking does |
|---|---|---|
| Model training | OpenAI GPTBot, Anthropic ClaudeBot, Common Crawl CCBot, Google Google-Extended (Gemini training / grounding control - not Googlebot) | Signals that content should not be used to train (or, for Google-Extended, to train Gemini / power certain AI features). Does not, by itself, remove you from classic Google Search. |
| Search / answer indexing | OpenAI OAI-SearchBot, Anthropic Claude-SearchBot, PerplexityBot | Reduces or removes eligibility to be surfaced and cited in that product's search-style answers. |
| User-triggered fetch | OpenAI ChatGPT-User, Anthropic Claude-User, similar "user asked for this URL" agents | Controls live fetches when a person points the assistant at a page. Vendor rules differ on whether robots.txt is honoured for these agents - check the vendor's current docs before assuming a Disallow works. |
OpenAI's own crawler documentation states the independence explicitly: a webmaster can allow OAI-SearchBot to appear in ChatGPT search features while disallowing GPTBot so crawled content is not used to train generative foundation models. Anthropic documents the same three-way split for ClaudeBot (training), Claude-SearchBot (search quality), and Claude-User (user-directed retrieval).
So the policy question is not "AI: on or off?" It is three separate decisions that should match Fernwood's risk tolerance and AEO goals.
Usually allow search/answer crawlers when:
- You want assistants to cite Fernwood's docs, research, and comparison pages.
- Competitors already appear in those answers and you do not.
- Silktide's AI crawler access check is warning because a template blocked everything by accident.
Often block or limit training crawlers when:
- Legal or brand policy forbids contributing site content to third-party model training.
- You publish proprietary research, pricing logic, or gated-adjacent material that must stay out of training corpora even if you still want live citations.
Block broadly when:
- The site is non-public, staging, or has no marketing reason to be in assistant answers (intranet, pure app UI, partner portal).
- You have made a conscious product decision to stay out of AI answers entirely - and you are willing to lose that channel.
How to decide (Fernwood's worksheet)
Work through these in order. Write the answers down; the robots.txt file is just the encoding of the worksheet.
- Do we want to be cited in AI answers at all? If no, disallow the major search/answer bots and approve Silktide's warning. If yes, continue.
- Which assistants matter? ChatGPT, Claude, Perplexity, Google AI features, and others each have their own agents. Prioritise the ones your buyers actually use.
- Is training opt-out required? If legal says yes, disallow training agents (
GPTBot,ClaudeBot,Google-Extended,CCBot, and peers) while keeping search agents allowed where the vendor supports the split. - Are there paths that must stay private either way?
/app/,/internal/, draft preview URLs - disallow those for*(and confirm auth still protects them; robots.txt is not access control). - Is the HTML actually fetchable? A permissive robots.txt with an empty SPA shell still fails AEO - fix rendering next.
Implementing the policy in robots.txt
Robots.txt lives at https://fernwood.example/robots.txt. Rules are per user-agent group. More specific groups override the catch-all for that agent. Prefer named agents over hoping User-agent: * expresses your AI policy - templates that set Disallow: / under * also hobble Googlebot and everyone else.
Pattern A - Visible in answers, out of training (common for AEO-minded sites)
# OpenAI: cite in ChatGPT search, do not use for foundation-model training
User-agent: OAI-SearchBot
Allow: /
User-agent: GPTBot
Disallow: /
# Anthropic: allow search indexing; block training crawl
User-agent: Claude-SearchBot
Allow: /
User-agent: ClaudeBot
Disallow: /
# Google: Googlebot (Search) is separate; Google-Extended controls Gemini training / AI features
User-agent: Google-Extended
Disallow: /
# Keep private app paths closed to everyone
User-agent: *
Disallow: /app/
Disallow: /internal/
OpenAI notes that robots.txt changes for search can take on the order of ~24 hours to take effect. Re-check after publishing.
Pattern B - Fully open (maximum answer-engine reach)
Omit AI-specific Disallow rules. Ensure nothing under User-agent: * accidentally bans the whole site. Still protect private paths.
Pattern C - Fully closed to AI products
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Claude-SearchBot
Disallow: /
User-agent: Claude-User
Disallow: /
User-agent: PerplexityBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: CCBot
Disallow: /
Use this only as a conscious choice. Then Silktide's AI crawler warning so the finding stops counting against the site.
Verify who is actually hitting you
User-Agent strings are trivial to spoof. Prefer vendor-published IP lists when you need certainty:
- OpenAI publishes crawler IP JSON (linked from OpenAI's bot docs).
- Anthropic publishes crawler IPs at claude.com/crawling/bots.json.
Blocking by IP alone, without a robots.txt signal, is brittle - Anthropic notes that IP blocks can interfere with the bot's ability to read your robots.txt and complete an opt-out cleanly.
What robots.txt does not fix
- Firewall / bot-management false positives. A web application firewall (WAF) - the filtering layer services like Cloudflare put in front of a site - that challenges or 403s "non-browser" clients can block GPTBot even when robots.txt allows it. Silktide tests that separately as AI crawlers blocked in practice.
- Empty JavaScript shells. Allowing the crawler into a
<div id="root"></div>helps nobody. Pair this technique with Content readable without JavaScript. - A magical ranking boost. Permitting crawlers makes citation possible. Quotable pages, dates, identity, and evidence still decide whether you get chosen - see Answer-first page structure and Organisation identity markup.
llms.txtas a substitute. An index file cannot override aDisallow. See llms.txt.
How Silktide helps
- AI crawler access parses robots.txt against a list of well-known AI user agents and warns when the site root is disallowed - as a warning you can approve when the block is intentional.
- AI crawlers blocked in practice fetches as those crawlers to catch firewall and edge blocks robots.txt cannot explain.
- Content without JavaScript catches the "allowed but empty" failure mode.
Related
- AI crawler access - the Silktide check this technique expands
- Content readable without JavaScript
- llms.txt - optional curated index after access is solved
- OpenAI crawler documentation
- Anthropic crawler documentation
- Google crawlers overview (including Google-Extended)
- robots.txt specification (RFC 9309)