Robots.txt
Websites should provide a robots.txt file, which tells search engine crawlers which parts of the site they may visit. Silktide checks that your website serves one.
A file lives at the root of the site - for example.com it must be at example.com\/robots.txt. To allow crawlers everywhere, the file can be as simple as:
# Allow all crawlers to access everything
User-agent: *
Disallow:
Why this matters
Crawlers request robots.txt before anything else, and its absence leaves you without the standard way to steer them - away from search result pages, faceted filters, or other URLs that waste crawl budget. It is also where you declare the location of your with a Sitemap: line.
Even for a site that wants everything crawled, an explicit robots.txt states that intent clearly instead of leaving crawlers to infer it from an error response.
How to fix it
- Create a plain text file named
robots.txtand place it in the root folder of your website so it is served athttps:\/\/example.com\/robots.txt. - Start permissive (the example above) and add
Disallowrules only for areas crawlers should avoid. - Add a
Sitemap:line pointing to your XML sitemap:
User-agent: *
Disallow: \/search\/
Sitemap: https:\/\/example.com\/sitemap.xml
- Confirm the file loads in a browser and returns a normal response rather than an error page. Many content management systems can generate and serve it for you.
How Silktide tests this
- Take the homepage URL of each of your website.
- Request
\/robots.txtat the root of that host. - Pass if the request returns HTTP status 200 (OK); flag the section if the file is missing, returns an error, or cannot be downloaded.
Troubleshooting
I have a robots.txt but the check fails
The file must be served with HTTP status 200. Some servers return a "soft" error page - a page that looks like an error but carries a redirect or error status - or block automated requests entirely. Test the URL directly and check the status code, ideally using your browser's developer tools or from a different network.
Will robots.txt keep pages out of search results?
Not reliably. Disallow stops crawling, but a page can still be indexed from links elsewhere. To keep a page out of results, use instead - see Page set to noindex.