Skip to Content
OpportunitiesTechnicalIssue ReferenceRobots.txtrobots.txt contains lines crawlers cannot read

robots.txt contains lines crawlers cannot read

What it is

One or more lines in your robots.txt are not valid robots.txt records. Crawlers skip the lines they cannot read and follow the rest, so a mistyped rule quietly does nothing: a misspelled Disalow blocks nothing, and a rule written before any User-agent line applies to no crawler.

Why it matters

The damage depends on the line that is lost. A dropped Disallow leaves sections open that you meant to keep crawlers out of, and a dropped Allow can leave pages blocked that you meant to open. Unsupported lines such as Noindex: give a false sense of control: Google  ignores lines it does not support. AI crawlers that honor robots.txt lose the same rule.

How Asky checks it

Asky reads your robots.txt during its daily site discovery and at the start of each technical audit, ignoring blank lines and anything after #. A line is an error when it has no colon, when an Allow or Disallow comes before the first User-agent, or when its field is not User-agent, Allow, Disallow, Sitemap, Crawl-delay, Host, Request-rate, Visit-time or Clean-param. Field names are not case-sensitive. The details list each error with its line number. Paths and URLs are not validated, and the other robots.txt checks still run on the valid lines.

Reported as a warning with low severity.

How to fix it

  1. Open the issue’s details for the line numbers, then open https://example.com/robots.txt for your own domain.
  2. Write each line as Field: value, for example Disallow: /checkout/. A line without a colon is skipped.
  3. Correct misspelled field names such as Disalow, Useragent or User agent.
  4. Move rules that sit above the first User-agent line into the group of the crawlers they are meant for.
  5. Remove fields robots.txt does not support, such as Noindex:. To keep a page out of search results, put a noindex robots meta tag on the page and leave it crawlable, so crawlers can see the tag.
  6. Check what the corrected rules now do. A repaired Disallow: / under User-agent: * or User-agent: Googlebot blocks the whole site, which is reported as its own critical issue.

Example

In the first version, Disallow /cart/ has no colon and Disalow is misspelled, so neither rule applies:

# Before User-agent: * Disallow /cart/ Disalow: /checkout/ # After User-agent: * Disallow: /cart/ Disallow: /checkout/ Sitemap: https://example.com/sitemap.xml

← Back to Robots.txt

Last updated on