Structure

Missing or broken robots.txt: how to find and fix it

Check C1 · Last reviewed

robots.txt is a plain-text file at the root of your domain that tells search engine and AI crawlers which parts of the site they may visit. It should exist, return a 200 response, and parse without errors. A missing or broken one is usually more of a hygiene gap than a ranking cliff, but a misconfigured one can accidentally block your whole site.

Why it matters

Search engines check robots.txt before crawling anything else on the domain. If the file returns an error or times out, some crawlers back off rather than risk crawling something they should not.

The bigger risk is not absence but a wrong rule: a leftover `Disallow: /` from a staging environment, migrated live by accident, blocks every crawler including Google from the entire site.

The same file increasingly controls AI crawlers (GPTBot, PerplexityBot, and others), so it is also where you decide whether AI assistants can read your site at all.

How AuditCrow checks this

AuditCrow fetches `/robots.txt` on the domain, confirms it returns a successful response, and parses the directives to check they are syntactically valid and are not blocking the whole site (a separate check, C2, confirms Googlebot and Bingbot specifically are allowed).

How to fix it

  1. 1Visit `yourdomain.com/robots.txt` in a browser. If it 404s, the file does not exist; if it loads, check for a `Disallow: /` line under a wildcard `User-agent: *` block, which blocks everyone.
  2. 2A minimal, safe robots.txt for most small sites is just a sitemap reference and no disallow rules: `User-agent: *` then `Sitemap: https://yourdomain.com/sitemap.xml`.
  3. 3Only add `Disallow` rules for things you genuinely do not want indexed, such as an internal search results path or an admin area (though admin areas are usually already behind a login).
  4. 4In WordPress, check Settings, then Reading, then "Search engine visibility": if "Discourage search engines from indexing this site" is ticked, WordPress serves a robots.txt that blocks everything. This is the single most common cause of this problem, usually left on from when the site was in development.
  5. 5Use AuditCrow's own robots.txt checker to validate the file's syntax, and to see the current recommended blocks for AI crawlers if you want to allow or disallow them explicitly.

Related terms

robots.txt, Crawl budget

Related checks

Want to see whether your own site passes this check, and everything else?

Join the waitlist

Be first when scans reopen

Scans are paused for a moment. Join the waitlist and we'll tell you when they're back.

We'll email you when free scans are back