robots.txt controls which URLs a search engine may crawl; noindex controls which URLs it may index (show in results). They are two different things, and confusing them produces the classic technical SEO mistake: a page blocked in robots.txt can still appear in Google, and a page with noindex only disappears if Google can crawl it to read the tag.
This article explains what each mechanism does, how to combine them and what to use for the usual cases: internal search, filters, admin pages and content under construction.
What robots.txt does
It is a text file at the root of the domain (https://example.com/robots.txt) with rules per user agent stating which paths should not be requested:
User-agent: *
Allow: /
Disallow: /search/
Disallow: /*?q=
Sitemap: https://example.com/sitemap.xml
Key points from Google’s documentation:
- It is for managing crawling: preventing the bot from spending time on worthless pages or overloading the server.
- It does not prevent indexing. If other pages link to a blocked URL, Google may index it without visiting it, showing only the URL and a note such as “no information is available for this page”.
- It is public: anyone can read it. Do not list paths you want to keep secret.
- Since September 2019 Google no longer supports
noindexinside robots.txt; that rule is ignored. - The
Sitemap:line is the simplest way to declare the sitemap and can be included even if you also submit it through Search Console.
What noindex does
noindex tells the search engine not to include the page in its results. It is declared in two ways:
<!-- In the <head> of an HTML page -->
<meta name="robots" content="noindex, follow" />
# As an HTTP header, for any file type (PDF, images, JSON...)
X-Robots-Tag: noindex
Two important details:
- For it to work, the bot must be able to crawl the page and read the tag. If you also block it in robots.txt, it will never see the
noindex. follow(the default) lets the page’s links keep passing signals.noindex, nofollowcuts that too; use it only when you do not want the links followed.
Over time, Google treats a long-standing noindex as a signal to crawl that URL less often, but that is a consequence, not the purpose.
The wrong combination
| Setup | Result |
|---|---|
| Blocked in robots.txt, no noindex | May appear in results (URL only), because Google cannot read the page but knows it exists from links. |
| Blocked in robots.txt and noindex | Same as above: the noindex is never read. |
| Crawlable, with noindex | Disappears from results as soon as Google recrawls it. Correct for deindexing. |
| Crawlable, no noindex | Indexed normally. |
Practical conclusion: if you want a URL out of Google, use noindex and leave it crawlable. Only once it has disappeared does it make sense, if you want to save crawl budget, to block it in robots.txt.
What to use in each case
| Case | Recommendation |
|---|---|
Internal search results (/search/?q=...) |
noindex on the page. Optionally Disallow in robots.txt once deindexed, to avoid crawling endless combinations. |
Filters and sort orders with parameters (?sort=price) |
rel="canonical" to the parameter-free URL; if they generate many combinations, noindex. |
| Admin or login page | noindex + real protection with authentication. robots.txt is not a security measure. |
| Empty or under-construction pages | noindex while they have no content; remove it when published. |
| Files you do not want in results (internal PDFs, JSON) | X-Robots-Tag: noindex header. |
| Resources the bot does not need (scripts of an internal panel) | Disallow in robots.txt. Never block the CSS and JS the public page needs to render. |
| Staging environment | HTTP authentication. If not possible, noindex on every page; do not rely on robots.txt alone. |
Common mistakes
- Blocking CSS or JavaScript in robots.txt. Google renders pages; if it cannot load their resources, it will judge them broken.
- Forgetting to remove
noindexwhen going live. The most frequent cause of “my new site does not show up”. Check the template and the server headers. - Using robots.txt to “delete” pages from Google. It does not work; use
noindexor, if the page no longer exists, return 404 or 410. - Blocking the sitemap or the home page with a
Disallow: /left over from development. - Putting
noindexon important pages (categories, author, home) by applying a plugin’s generic rules.
How to check
- URL inspection in Search Console: tells you whether the URL is blocked by robots.txt, whether it has
noindexand whether it is indexed. - Pages report (Indexing): groups excluded URLs by reason (“Blocked by robots.txt”, “Excluded by ‘noindex’ tag”, “Indexed, though blocked by robots.txt”). That last category is precisely the symptom of the wrong combination.
- robots.txt report in Search Console: shows which version of the file Google has and whether it contains syntax errors.
Conclusion
robots.txt says “do not come in here”; noindex says “do not show me”. For something not to appear in Google, it must be crawlable and carry noindex. Keep robots.txt for saving crawl budget on worthless areas and for declaring the sitemap, and check in Search Console that no pages are “indexed, though blocked”: that is the sign the two mechanisms are stepping on each other.