robots.txt
robots.txt is a plain text file at the root of a website (for example https://example.com/robots.txt) that tells automated which parts of the site they may fetch.
User-agent: GPTBot
Disallow: /
User-agent: *
Allow: /
Each rule names a crawler (or * for all of them) and lists the paths it is allowed or disallowed to visit. Search engines and check this file before fetching pages, so a single Disallow line can make an entire site invisible to an AI assistant - sometimes intentionally, often by accident. robots.txt controls crawling, not indexing: blocking a page here does not remove it from search results the way does, and can even backfire by hiding the noindex tag itself.
Learn more
- Robots.txt introduction and guide - Google's practical guide
- RFC 9309: Robots Exclusion Protocol - the formal IETF standard