Should You Use Noindex or robots.txt to Hide a Page?

Robots.txt asks Google not to look. Noindex tells Google what it saw should not be listed. Use the wrong one and the page ends up in results with no title and no description.

Should You Use Noindex or robots.txt to Hide a Page? — Troiana insight cover

In short

To keep a page out of Google's results, use a noindex directive, either a meta robots tag or an X-Robots-Tag header, and leave the page crawlable so Google can see the instruction. Robots.txt only prevents crawling; a disallowed page can still be indexed from links to it, appearing in results with no snippet. Use robots.txt for crawl-budget control on large sites and for resources that should never be fetched, and never combine the two on the same URL, because the disallow stops Google from ever reading the noindex.

Two mechanisms that sound alike

Robots.txt is a file at the root of the site that tells crawlers which paths they may fetch. Disallow: /admin/ asks well-behaved bots not to request anything under /admin/. It controls crawling, which is the act of downloading a page. It says nothing about indexing, which is the act of listing a page in results.

Noindex is an instruction on the page itself, either <meta name="robots" content="noindex"> in the HTML or an X-Robots-Tag: noindex HTTP header. It tells a crawler that has fetched the page not to include it in the index. It controls indexing, and it requires the page to be crawled so the instruction can be read.

That last sentence is the whole subject.

The classic mistake

A page is disallowed in robots.txt and given a noindex tag, on the reasoning that two locks are better than one. Google obeys the robots.txt, never fetches the page, and therefore never sees the noindex. Meanwhile, other pages link to the URL, so Google knows it exists. The result is a page listed in search results with its URL as the title and the text "No information is available for this page" as the description, which is exactly the outcome the two locks were meant to prevent.

If a page must be absent from results, it must be crawlable and carry noindex. If you have already blocked it in robots.txt and it is showing up, remove the disallow, let Google crawl it, and it will drop out once the noindex is read.

Which to use for each case

A staging or development site. Password protection is the correct answer: a crawler that cannot fetch the page cannot index its content, and nobody else can see it either. If that is not possible, noindex on every page, and check the live site does not inherit it at launch.

Thin or utility pages such as internal search results, tag archives, printer-friendly versions, thank-you pages. Noindex. They should stay crawlable so links through them still pass, and so Google reads the instruction.

Duplicate URLs such as filter combinations, tracking parameters, or the same page under two paths. Usually a canonical tag pointing at the preferred version rather than either mechanism, because canonicals consolidate signals while noindex discards them.

Private areas behind a login. The login itself keeps content out; the login page can be indexed or not as you prefer. Robots.txt disallow for the authenticated paths is fine here, because there is nothing to index and it saves the crawler's time.

Large sites with crawl budget problems, such as millions of faceted URLs. Robots.txt is the right tool, because the goal is to stop Google spending its crawl on them, and indexing was never the risk. This is the case crawl budget advice is written for and it does not apply to sites with a few thousand pages.

Files that are not pages: PDFs, images, scripts. Noindex via X-Robots-Tag header if they must not be listed; robots.txt if they merely should not be fetched. Never block CSS and JavaScript, because Google needs them to render pages.

Removing something urgently. Search Console's removal tool hides a URL from results within hours, for about six months. Use it alongside noindex so the removal becomes permanent.

A noindexed page is still crawled and its links are still followed, at least initially; Google has said that long-term noindexed pages are eventually crawled less and treated closer to nofollow. For a page whose links matter, such as a hub page you do not want in results but do want passing authority, this is a consideration, and a canonical or simply leaving it indexed may be better.

What robots.txt does not do

It does not hide content from anyone; the file is public and the disallowed paths are listed in it for all to read, which is why it is not a security measure. It does not stop bad bots, which ignore it. It does not remove pages already indexed. And it does not stop indexing of URLs discovered through links, only crawling of their content.

A checklist

  • Want it out of results: noindex, crawlable, no disallow.
  • Want it out of results now: noindex plus the removal tool.
  • Want the crawler to skip it but do not care about indexing: robots.txt.
  • Want two URLs treated as one: canonical.
  • Want it private: authentication.
  • Never: robots.txt disallow and noindex on the same URL.

After any change, use Search Console's URL Inspection tool to see what Google actually does with the page; it reports both whether crawling is allowed and whether indexing is. A well-written robots.txt is short and rarely changes. If yours is long, it is probably doing jobs that belong to noindex or canonicals, and book a call if you would like it untangled.

Common questions

Does robots.txt stop a page from being indexed?

No. It stops Google from crawling the page's content, but if other pages link to the URL, Google can still index it and show it in results with the URL as the title and no description. To keep a page out of results, use a noindex directive and leave the page crawlable so Google can read it.

Can I use noindex and robots.txt together?

Not on the same URL. If robots.txt disallows the page, Google never fetches it and never sees the noindex tag, so the page can remain indexed indefinitely. Choose one: noindex for pages that must not appear in results, robots.txt for paths the crawler should not spend time on.

How long does it take for a noindex page to disappear from Google?

Usually days to a few weeks, depending on how often Google recrawls the page. Requesting indexing through Search Console's URL Inspection tool after adding the tag speeds it up. For urgent removals, the Search Console removal tool hides the URL within hours for about six months while the noindex makes it permanent.

What is the difference between noindex and nofollow?

Noindex tells search engines not to list the page in results. Nofollow tells them not to pass ranking signals through the page's links, or through a specific link when used on an anchor. They are independent: a page can be noindex and still have its links followed, which is the default and usually what you want.

Should I block my staging site with robots.txt?

Password-protect it instead. Robots.txt is public and does not stop indexing of URLs that get linked, and it is easy to copy the file to the live site by mistake. If protection is impossible, put noindex on every staging page and confirm at launch that the live site does not carry it.

Have something worth building right?