Skip to Content
OpportunitiesTechnicalIssue ReferenceRobots.txtrobots.txt returns a server error

robots.txt returns a server error

What it is

The request for your robots.txt got a server error, an HTTP status from 500 to 599, instead of the file. The file may well exist, but the server or CDN failed to deliver it. A robots.txt built by a CMS plugin or an app route fails this way when that code breaks, and some firewalls answer crawlers they do not trust with 503.

Why it matters

A crawler cannot know what a failing robots.txt would say, so it assumes the worst: the robots.txt standard  tells it to treat the whole site as disallowed. Google  stops crawling the site for the first 12 hours, then relies on its last good copy for up to 30 days. New and updated pages are not fetched meanwhile. AI crawlers that follow the standard are likely to hold back in the same way.

How Asky checks it

Asky requests /robots.txt on your domain during its daily site discovery and at the start of each technical audit. After a 429 or 5xx it waits and retries once with a browser user agent. The issue is reported when the final answer is 500 or higher, with the status in the details. No other robots.txt check runs then: earlier findings and the last good copy are kept, and no sitemap is reported as unlisted. Without the Sitemap lines, Asky uses sitemaps added in Asky or tries the standard paths. A final 429, a timeout or a refused connection skips the check. It clears on the first run that loads the file.

Reported as a critical issue with high severity.

How to fix it

  1. Open https://example.com/robots.txt in a browser and in an HTTP status checker, and note the status each one gets.
  2. If it fails for everyone, check the server or CDN error logs for that request. If a plugin or app route builds the file, fix that code or serve a static file instead.
  3. If it works in a browser but fails for crawlers, a firewall or bot protection is answering them with an error. Exempt /robots.txt from those rules. The AskyBot page shows how to allow Asky’s crawler.
  4. If the site is in maintenance mode, exempt /robots.txt so it keeps answering 200. A 503 during a short planned outage is fine; for days it is not.
  5. Confirm the URL answers 200 with the text of the file.

Example

The status check shows the failure before the fix and the file after it:

curl -s -o /dev/null -w "%{http_code}\n" https://example.com/robots.txt # 500 before the fix, 200 after

← Back to Robots.txt

Last updated on