robots.txt blocks Googlebot from the whole site
What it is
The rules Googlebot follows in your robots.txt disallow the entire site. Googlebot obeys a group that names it, such as User-agent: Googlebot, and falls back to the User-agent: * group only when no group names it. Either one can cause this, which is how a site ends up blocking Google while other search engines are allowed.
Why it matters
Googlebot is the crawler behind Google Search. While it is blocked, Google cannot read new or changed pages, and URLs it already knows can show without a description . AI Overviews and AI Mode only link to pages that are indexed and eligible for a snippet , so the block reaches them too. Gemini’s use of your content is a separate choice, made with the Google-Extended token.
How Asky checks it
Asky reads your robots.txt during its daily site discovery and at the start of each technical audit. It collects every group whose User-agent is exactly Googlebot, in any letter case, or the * groups when none is, and reports the issue when those rules disallow two sample paths inside the site, /a and /z9/page. The longest matching rule wins, Allow on a tie. Groups for Googlebot-Image, Google-Extended or Googlebot/2.1 are not read, though Google treats the last as Googlebot. A block on some paths only is not reported. When the * group and Bingbot’s rules block the site too, the all-crawlers issue is raised as well.
Reported as a critical issue with high severity.
How to fix it
- Open
https://example.com/robots.txtfor your own domain. Find the group naming Googlebot, or theUser-agent: *group if there is none, and itsDisallow: /orDisallow: /*. - Remove that rule, or narrow it to the paths Google should skip, such as
/checkout/. - A group naming Googlebot replaces the
*group for it completely. If you keep one, copy into it the*rules you still want Google to follow. - If a plugin, CDN or hosting panel writes the file, change it there so the block does not come back.
- Use the robots.txt report in Search Console to ask Google to fetch the new file.
- To leave it as it is: if this host should stay out of Google on purpose, such as a staging site, mark the issue resolved.
Example
A group copied from another project shuts Google out of https://example.com while every other crawler is allowed:
# Before
User-agent: Googlebot
Disallow: /
User-agent: *
Disallow: /checkout/
# After: Googlebot follows the same rules as everyone else
User-agent: *
Disallow: /checkout/
Sitemap: https://example.com/sitemap.xml