Skip to Content
OpportunitiesTechnicalIssue ReferenceRobots.txtrobots.txt blocks Googlebot from the whole site

robots.txt blocks Googlebot from the whole site

What it is

The rules Googlebot follows in your robots.txt disallow the entire site. Googlebot obeys a group that names it, such as User-agent: Googlebot, and falls back to the User-agent: * group only when no group names it. Either one can cause this, which is how a site ends up blocking Google while other search engines are allowed.

Why it matters

Googlebot is the crawler behind Google Search. While it is blocked, Google cannot read new or changed pages, and URLs it already knows can show without a description . AI Overviews and AI Mode only link to pages that are indexed and eligible for a snippet , so the block reaches them too. Gemini’s use of your content is a separate choice, made with the Google-Extended token.

How Asky checks it

Asky reads your robots.txt during its daily site discovery and at the start of each technical audit. It collects every group whose User-agent is exactly Googlebot, in any letter case, or the * groups when none is, and reports the issue when those rules disallow two sample paths inside the site, /a and /z9/page. The longest matching rule wins, Allow on a tie. Groups for Googlebot-Image, Google-Extended or Googlebot/2.1 are not read, though Google treats the last as Googlebot. A block on some paths only is not reported. When the * group and Bingbot’s rules block the site too, the all-crawlers issue is raised as well.

Reported as a critical issue with high severity.

How to fix it

  1. Open https://example.com/robots.txt for your own domain. Find the group naming Googlebot, or the User-agent: * group if there is none, and its Disallow: / or Disallow: /*.
  2. Remove that rule, or narrow it to the paths Google should skip, such as /checkout/.
  3. A group naming Googlebot replaces the * group for it completely. If you keep one, copy into it the * rules you still want Google to follow.
  4. If a plugin, CDN or hosting panel writes the file, change it there so the block does not come back.
  5. Use the robots.txt report  in Search Console to ask Google to fetch the new file.
  6. To leave it as it is: if this host should stay out of Google on purpose, such as a staging site, mark the issue resolved.

Example

A group copied from another project shuts Google out of https://example.com while every other crawler is allowed:

# Before User-agent: Googlebot Disallow: / User-agent: * Disallow: /checkout/ # After: Googlebot follows the same rules as everyone else User-agent: * Disallow: /checkout/ Sitemap: https://example.com/sitemap.xml

← Back to Robots.txt

Last updated on