Skip to content
SilktideHelp

Robots.txt

Websites should provide a robots.txt file, which tells search engine crawlers which parts of the site they may visit. Silktide checks that your website serves one.

A file lives at the root of the site - for example.com it must be at example.com/robots.txt. To allow crawlers everywhere, the file can be as simple as:

# Allow all crawlers to access everything
User-agent: *
Disallow:

Why this matters

Crawlers request robots.txt before anything else, and its absence leaves you without the standard way to steer them - away from search result pages, faceted filters, or other URLs that waste crawl budget. It is also where you declare the location of your with a Sitemap: line.

Even for a site that wants everything crawled, an explicit robots.txt states that intent clearly instead of leaving crawlers to infer it from an error response.

How to fix it

  1. Create a plain text file named robots.txt and place it in the root folder of your website, so it is served at https://example.com/robots.txt.
  2. Start permissive (the example above) and add Disallow rules only for areas crawlers should skip.
  3. Add a Sitemap: line pointing at your XML sitemap:
User-agent: *
Disallow: /search/

Sitemap: https://example.com/sitemap.xml
  1. Confirm the file loads in a browser and returns normally rather than an error page. Many content management systems can generate and serve it for you.

How Silktide tests this

  1. Take the home URL of each of your website.
  2. Request /robots.txt at the root of that host.
  3. Pass if the request returns HTTP status 200 (OK); flag the section if the file is missing, returns an error, or cannot be downloaded.

Troubleshooting

I have a robots.txt but the check fails

The file must be served with HTTP status 200. Some servers return a "soft" error page - a page that looks like an error but carries a redirect or error status - or block automated requests entirely. Test the URL directly and check the status code, ideally with your browser's developer tools or from a different network.

Will robots.txt keep pages out of search results?

Not reliably. Disallow stops crawling, but a page can still be indexed from links elsewhere. To keep a page out of results, use instead - see Page set to noindex.

Learn more

Last updated

Was this page helpful?

Robots.txt | Silktide Help