Inhoudsopgave

Why remove a page from the index?

There are many reasons to remove a page from Google’s index (or that of other search engines). The most common reasons are:

  • A page is no longer relevant, so users no longer need to find it.
  • A page was accidentally indexed when it shouldn’t have been.
  • It appears that your website’s staging environment has been indexed.
  • This page is a duplicate of another page on your website.
  • The page contains low-quality content.
  • The URL in the index returns a 404 error.
  • There is a legal reason why the page should not be included in the index (anymore).

Ways to Remove a Page from the Index

There are several ways to remove a page from the index. Each method has its pros and cons, and the most effective method is to add a noindex tag.

When the content is no longer relevant to your visitors, deleting the page is the most effective solution. Most CMSs will then display a 404 or 410 error. When Google encounters a 404 error, the page is usually removed from the index within a few weeks.

Itโ€™s often thought that a 404 error is harmful, but thatโ€™s not the case. If there simply isnโ€™t a relevant URL to redirect to, you should display a 404 error. If you do set up a redirect, Google will treat that as a soft 404 error.

Adding a noindex tag is the most effective way to remove URLs from the index and keep them out of the index. Search engines respect noindex tags and usually process the change within a few weeks. The noindex tag must be placed in the meta robots tag within your pageโ€™s and tells search engines that the page must literally be โ€œunindexable.โ€

The robots.txt file contains guidelines for crawlers, including instructions on which URL paths on your website should not be crawled. Do you have entire groups of pages that you want to remove from the index? If so, the robots.txt file is an effective way to do this. The downside of using โ€œdisallowโ€ directives in your robots.txt file is that search engines ignore them when there are conflicting signals. Therefore, this is not the most effective or reliable solution for removing URLs from the index.

The robots.txt file is, however, a good way to prevent search engines from crawling URL paths and, as a result, potentially indexing them. So it is primarily a way to prevent indexing.

Do you want to quickly remove a URL or a group of URLs from the index? If so, you can use the URL Removal tool in Google Search Console. With this tool, you can request the temporary removal of a single URL or a group of URLs:

URL Removal Tool in Google Search Console

Google will then remove the URLs from the index for about 6 months. Do you want to keep the URL out of the index after that? If so, itโ€™s important to add a noindex tag in the meantime. The URL Removal Tool is an effective way to remove a URL from the index in a very short amount of time.

An X-Robots-Tag HTTP header is the same as a meta robots noindex tag. However, this X-Robots-Tag HTTP header is added at the server level and can be used for both HTML pages and files. That is why this is an effective way to, for example, remove a PDF file from Googleโ€™s index. In practice, the X-Robots-Tag HTTP header is also used when it is not feasible to modify meta robots tags.

Important to keep in mind

In practice, conflicting actions often occur when removing pages from the index. Keep the following points in mind:

  • Do not use the noindex tag and the robots.txt “disallow” directive interchangeably. Because of a “disallow” directive in the robots.txt file, Google can no longer view the page, which means you’ll see a message in the search results stating that a description of the page cannot be displayed due to the robots.txt file.
  • Do not use the URL removal tool as a permanent solution. Google’s URL removal tool removes a URL from the index for about 6 months. If the page is still indexable after that, it will be re-included in the search results.
  • Do not redirect a URL to an irrelevant page. Are you redirecting a URL to a page that isn’t relevant? If so, Google will treat it as a soft 404 error. This confuses search engines, results in a poor user experience, and wastes your crawl budget.
  • Exclude parts of a website using the robots.txt file. Do you want to exclude entire URL paths from crawling and indexing? If so, you can add a “disallow” directive to do this. This way, you prevent parts of your website from being crawled.
  • Did you accidentally get your staging website indexed by Google? In that case, it’s important to add a noindex tag or HTTP authentication, and to have it removed more quickly, you can submit the URL for removal using Google’s URL Removal Tool.
Jarik Oosting

This article is written by Jarik Oosting

With a passion for SEO and an unmatched drive for results, Jarik Oosting is the driving force behind SmartRanking. With over 15 years of experience in the field, he has built an extensive body of knowledge spanning technical SEO to complex site migrations. As the founder of SmartRanking, he has assembled a team of like-minded SEO specialists who help businesses achieve sustainable online growth.

His academic background in information science at the University of Groningen, with a specialization in natural language processing, gives him a unique perspective on the world of SEO. For Jarik, it’s not just about visibility in search engines, it’s about sharing knowledge and guiding businesses toward sustainable online success. That mission also led him to write a Dutch book about GEO.

More about SmartRanking