SEO

robots.txt versus noindex: what each one does and when to use them

robots.txt controls crawling and noindex controls indexing. What happens when you combine them badly, how to handle internal search and filters, and how to check it.

robots.txt controls which URLs a search engine may crawl; noindex controls which URLs it may index (show in results). They are two different things, and confusing them produces the classic technical SEO mistake: a page blocked in robots.txt can still appear in Google, and a page with noindex only disappears if Google can crawl it to read the tag.

This article explains what each mechanism does, how to combine them and what to use for the usual cases: internal search, filters, admin pages and content under construction.

What robots.txt does

It is a text file at the root of the domain (https://example.com/robots.txt) with rules per user agent stating which paths should not be requested:

User-agent: *
Allow: /
Disallow: /search/
Disallow: /*?q=

Sitemap: https://example.com/sitemap.xml

Key points from Google’s documentation:

  • It is for managing crawling: preventing the bot from spending time on worthless pages or overloading the server.
  • It does not prevent indexing. If other pages link to a blocked URL, Google may index it without visiting it, showing only the URL and a note such as “no information is available for this page”.
  • It is public: anyone can read it. Do not list paths you want to keep secret.
  • Since September 2019 Google no longer supports noindex inside robots.txt; that rule is ignored.
  • The Sitemap: line is the simplest way to declare the sitemap and can be included even if you also submit it through Search Console.

What noindex does

noindex tells the search engine not to include the page in its results. It is declared in two ways:

<!-- In the <head> of an HTML page -->
<meta name="robots" content="noindex, follow" />
# As an HTTP header, for any file type (PDF, images, JSON...)
X-Robots-Tag: noindex

Two important details:

  • For it to work, the bot must be able to crawl the page and read the tag. If you also block it in robots.txt, it will never see the noindex.
  • follow (the default) lets the page’s links keep passing signals. noindex, nofollow cuts that too; use it only when you do not want the links followed.

Over time, Google treats a long-standing noindex as a signal to crawl that URL less often, but that is a consequence, not the purpose.

The wrong combination

Setup Result
Blocked in robots.txt, no noindex May appear in results (URL only), because Google cannot read the page but knows it exists from links.
Blocked in robots.txt and noindex Same as above: the noindex is never read.
Crawlable, with noindex Disappears from results as soon as Google recrawls it. Correct for deindexing.
Crawlable, no noindex Indexed normally.

Practical conclusion: if you want a URL out of Google, use noindex and leave it crawlable. Only once it has disappeared does it make sense, if you want to save crawl budget, to block it in robots.txt.

What to use in each case

Case Recommendation
Internal search results (/search/?q=...) noindex on the page. Optionally Disallow in robots.txt once deindexed, to avoid crawling endless combinations.
Filters and sort orders with parameters (?sort=price) rel="canonical" to the parameter-free URL; if they generate many combinations, noindex.
Admin or login page noindex + real protection with authentication. robots.txt is not a security measure.
Empty or under-construction pages noindex while they have no content; remove it when published.
Files you do not want in results (internal PDFs, JSON) X-Robots-Tag: noindex header.
Resources the bot does not need (scripts of an internal panel) Disallow in robots.txt. Never block the CSS and JS the public page needs to render.
Staging environment HTTP authentication. If not possible, noindex on every page; do not rely on robots.txt alone.

Common mistakes

  1. Blocking CSS or JavaScript in robots.txt. Google renders pages; if it cannot load their resources, it will judge them broken.
  2. Forgetting to remove noindex when going live. The most frequent cause of “my new site does not show up”. Check the template and the server headers.
  3. Using robots.txt to “delete” pages from Google. It does not work; use noindex or, if the page no longer exists, return 404 or 410.
  4. Blocking the sitemap or the home page with a Disallow: / left over from development.
  5. Putting noindex on important pages (categories, author, home) by applying a plugin’s generic rules.

How to check

  • URL inspection in Search Console: tells you whether the URL is blocked by robots.txt, whether it has noindex and whether it is indexed.
  • Pages report (Indexing): groups excluded URLs by reason (“Blocked by robots.txt”, “Excluded by ‘noindex’ tag”, “Indexed, though blocked by robots.txt”). That last category is precisely the symptom of the wrong combination.
  • robots.txt report in Search Console: shows which version of the file Google has and whether it contains syntax errors.

Conclusion

robots.txt says “do not come in here”; noindex says “do not show me”. For something not to appear in Google, it must be crawlable and carry noindex. Keep robots.txt for saving crawl budget on worthless areas and for declaring the sitemap, and check in Search Console that no pages are “indexed, though blocked”: that is the sign the two mechanisms are stepping on each other.

Sources and references

  1. Google Search Central: Introduction to robots.txt developers.google.com
  2. Google Search Central: How to write and submit a robots.txt file developers.google.com
  3. Google Search Central: Block Search indexing with noindex developers.google.com
  4. Google Search Central: Robots meta tag and X-Robots-Tag specifications developers.google.com
  5. Google Search Central Blog: A note on unsupported rules in robots.txt (2019) developers.google.com
Articles in SEO