Skip to Content

robots.txt file not found

What it is

Your domain has no robots.txt file: the request for /robots.txt was answered with 404 or another client error, or with a web page instead of a text file. robots.txt is a plain text file at the root of a host, such as https://example.com/robots.txt, that tells crawlers which paths they may fetch and where your sitemap is. Every host serves its own, so a subdomain like blog.example.com needs a separate file.

Why it matters

A missing file does not block anything: Google and the robots.txt standard  treat a 404 as permission to crawl the whole site. What you give up is control and discovery. You cannot tell search or AI crawlers what to skip, and nothing points them to your sitemap, which Bing recommends listing there  for both search and Copilot coverage. A file that disappears is often a side effect of a migration or a new host.

How Asky checks it

Asky requests /robots.txt on your domain during its daily site discovery and at the start of each technical audit, following redirects. The issue is reported when the final answer is 404 or another 4xx other than 401, 403 and 429, or a web page in place of the text file. Each sitemap found is then also reported as not listed there. A 5xx is reported as a robots.txt server error, and a blank file as an empty robots.txt. When the request is refused with 401, 403 or 429, meets a bot-protection challenge or times out, the check is skipped. Only your domain’s own host is read. It clears once the file is found.

Reported as a critical issue with high severity.

How to fix it

  1. Create a plain text file named robots.txt, saved as UTF-8, and serve it at the root of your domain so https://example.com/robots.txt answers with status 200.
  2. Start with a group that allows everything plus a Sitemap: line holding the full URL of your sitemap or sitemap index, as in the example. That file raises no robots.txt issue.
  3. In Webflow, write it in the robots.txt field of the site’s SEO settings. In WordPress, use your SEO plugin’s file editor or upload the file to the web root. On a headless site, such as one built on Sanity, the front-end app serves it as a static file or route.
  4. Open the URL and check it shows the text itself, not your homepage or a 404 page. A web page served at that address is reported as this issue.
  5. Add Disallow rules only for paths crawlers should skip. Disallow: / under User-agent: * blocks the whole site.

Example

https://example.com/robots.txt answered 404. After the fix it answers 200 with this file:

User-agent: * Disallow: Sitemap: https://example.com/sitemap.xml

← Back to Robots.txt

Last updated on