Inhoudsopgave

Don’t be alarmed by this conclusion. About 25 to 30% of all content on the internet is duplicate, and this does not result in severe Google penalties. However, it can potentially have a serious impact on your ranking in search results.

Google provides the following definition of duplicate content: “Duplicate content typically refers to substantial blocks of content within or across domains that are either exactly the same or substantially similar” (source).

Exactly how much that is isn’t entirely clear. Google examines the entire page, including the header, footer, sidebar, images, videos, etc. It’s clear that Google doesn’t have a problem with you copying a few sentences from another source into a long piece of text (though this may be different from a copyright perspective!). If you’re copying entire paragraphs, however, Google will likely notice, and it will affect your search visibility.

What are the different types of duplicate content?

Meta Titles and Meta Descriptions

Duplicate meta titles—also known as title tags—and meta descriptions are a sign of duplicate content. Larger websites have a huge number of pages and almost always use a script to generate automated meta titles and meta descriptions that haven’t been manually filled in. Often, the choice is made to set all meta tags to the same value on pages where no manual tags have been entered. This makes it appear to search engines that these pages are identical. It’s also possible that all pages or a group of pages are accidentally assigned the same meta title or meta description.

Example of duplicate content in meta descriptions at the AD

For example, AD.nl has a meta description that’s the same on almost all of its pages. This isn’t necessarily a major problem, but it does make it harder for Google to distinguish between your different pieces of content. This is often relatively easy to fix.

Multiple URLs with the same destination (URL duplicates)

Internal URL duplicates are always a challenge. They’re often overlooked by web developers and online marketers, but it’s very important to resolve them. And fortunately, it’s usually relatively simple, too.

Basically, it comes down to this: if you have two or more versions of a URL, they’re considered different pages and therefore duplicate content. This is because you have two or more pages that contain exactly the same information or serve the same purpose. For example, if you sell bicycles, it’s likely that the URLs below display the same page or the same content (such as products):

  • https://fietsenshop.nl/stadsfietsen/gazelle/primeur
  • https://fietsenshop.nl/merken/gazelle/primeur/

To let Google know which page should rank, you use a canonical tag. You’ll learn more about this later in the blog. It’s also possible that all the pages on your website are duplicates. This is often the case if you can approach your website in the following way:

  • https://fietsenshop.nl
  • https://fietsenshop.nl
  • https://www.fietsenshop.nl
  • https://www.fietsenshop.nl

If you can access your website in any way without a 301 redirect and the URL remains exactly the same, then all the pages on your website are duplicates four times over! This calls for immediate action: Have all domains redirect to your preferred URL. For search engine optimization, it doesn’t really matter which option you choose—as long as you choose one!

Filter pages, search results pages, AMP pages, print-friendly pages, tag pages, mobile-friendly pages, tracking parameter pages, and session ID pages are all examples of pages that are very likely to be duplicates but are not the main page. As you can imagine, these URLs are not SEO-friendly. Make sure you have proper canonical tags, that your robots.txt file is set up correctly, or that you’re handling the parameters in Google Search Console. Want to be sure a page is blocked via your robots.txt file? Then use our robots.txt checker tool.

Location Pages

Suppose you have a website in the Netherlands and in Flemish-speaking Belgium with the same products and content. There are minor differences tailored to each country, such as the payment page and the terms and conditions. This confuses search engines, causing the Flemish site to rank in the Netherlands and vice versa. You don’t want that to happen. To prevent this, you should use an hreflang tag. Add this to your website, and search engines will understand that it’s focused on a specific country and isn’t duplicate content.

Useful? Yes! Because with your website, you can also easily conquer the Flemish search engines. And that doesn’t require any extra SEO effort! Keep in mind, though, that even with an hreflang tag, Google sometimes gets confused and treats pages as duplicates or ranks them in the wrong regions.

External duplicate content

The first three causes of duplicate content mentioned above were examples of internal duplicate content. External duplicate content occurs, for example, when different sources republish a press release, when standardized product information is copied, or when someone in bad faith blindly copies large chunks of text. Ranking in search engines is difficult in all these cases, because Google will often choose the “oldest” source.

Duplicate content in the form of product descriptions

Is duplicate content bad for your SEO?

There’s no clear-cut answer, but in many cases, duplicate content hurts your SEO results. At the very least, you’re not making it any easier for search engines, because you’re leaving it up to them to decide which page should rank. You have less control over this, and it’s entirely possible that Google will pick the “wrong” page. You can work around this with a redirect, but you’ll lose some authority in the process. So it’s not an ideal solution.

Your duplicate pages will also be indexed by Google, but they’ll appear lower in the search results than the duplicate that Google has chosen as its top pick. Note: This wastes your valuable crawl budget, especially if your website has three or four duplicate versions of every page! Every website has a specific crawl budget from Google. Once this is used up, Google won’t crawl all your pages. And with a lot of duplicate content, this happens much faster.

Furthermore, duplicate content leads to internal cannibalization. Multiple pages rank for the same search terms, which hurts performance for those terms. So you’ll never achieve a top ranking. On top of that, duplicate content puts you at risk of an algorithmic duplicate content penalty. As a result, one or more of your pages will no longer appear in the search results.

Finally, duplicate content isn’t ideal for your internal link structure. You decide which pages on your website receive link juice. Are you passing link juice to duplicate pages? That’s a waste! Because that way, you’re wasting valuable link juice on a page that isn’t meant to rank.

How do you resolve duplicate content issues?

Fortunately, there are several ways to fix your duplicate content issues. Below, we explain the five most important solutions for duplicate content issues. For more specific information on this topic, feel free to contact us—we’d be happy to help!

The rel=”canonical” attribute

As briefly mentioned above: the canonical tag. This is a tag that tells search engines which page is the original source of your duplicate content.

If URL #1 is the main page, you can add a canonical tag to URL #2 that points to the original page. By using a canonical URL, you’re indicating that https://fietsenshop.nl/stadsfietsen/gazelle/primeur is the original page. Add a canonical tag to the page https://fietsenshop.nl/merken/gazelle/primeur/ pointing to the first URL.

Then the search engines will know that the first URL is the main page. Incidentally, for SEO purposes, it’s always a good idea to include a canonical tag on all pages—a self-referencing canonical tag. This is simply a page that links to itself, thereby indicating to the search engine that this is the original page. This is important so that Google doesn’t treat URLs with parameters (such as UTM codes) as duplicates.

A 301 redirect

You can also easily resolve duplicates with a 301 redirect. This way, you keep a single page and pass all the link juice from the duplicate pages to that page. This is the most powerful and easiest way to reduce duplicate pages. However, this is often not possible due to unique parameters in a URL that are necessary for a page to function optimally, such as on a filter page.

Adding a noindex tag

Of course, you can also remove pages from your website from Google’s index using a noindex tag. The noindex tag is a meta robots tag that indicates that a page should not be included in the SERP, and it looks like this:

This way, the search engine will crawl your page but won’t index it. You could also set it to content=”noindex,nofollow,” but that’s not Google’s preferred method. A content=”nofollow” tag means that Google isn’t allowed to follow the links on that page. Google wants to see everything that’s happening on a website and wants to follow the links on a page.

Exclusions via the robots.txt file

Another good way to ensure that Google doesn’t crawl pages with duplicate content is the robots.txt file. This file allows you to direct crawlers across your website and deny access to certain parts of it. This saves your crawl budget because the entire page doesn’t need to be fetched and makes it easy to resolve large portions of your duplicate content. For online stores, for example, it’s highly recommended to exclude all filter parameters via the robots.txt file.

Writing Unique Content

The last—and most obvious—solution is to make your content unique. In any case, unique content that meets the visitor’s needs is always important. Do you have the same content on multiple pages? Rewrite the content and make sure it aligns with your visitor’s search intent and search query. This will not only improve your search rankings but also lead to a higher conversion rate.

How do I check if there is duplicate content on my website?

It’s fairly easy to check if there’s duplicate content on your website. There are a number of free tools available, but paid tools like Ahrefs, Semrush, or Deepcrawl are a necessary addition. The advantage of working with an SEO agency is that it has all the tools in-house and makes them available to you. Here’s what you can do without expensive tools:

  • Check the "Coverage" section in Google Search Console to see if Google has detected duplicate content. You can do this in the "Excluded" report. Next, look at the pages listed under “Duplicate page with no user-selected canonical version” and “Duplicate page, submitted URL not selected as canonical.”
  • Use the Siteliner website to quickly check whether your website contains any internal duplicate content.
  • Take a critical look at your website to see if it has multiple pages with duplicate content. Check whether these pages already have a canonical tag, a 301 redirect, or a "noindex" tag. If so, you’re all set. If not? Take action!
  • Enter four different URLs to access your website and see if they all redirect to one of these four. If not, take action!
  • You can often start a free 15- or 30-day trial with the tools mentioned above! Take advantage of this to significantly optimize your website.
  • Use tools like Screaming Frog and ContentKing to quickly identify duplicate content issues involving meta titles, meta descriptions, and HTML headings.
Jarik Oosting

This article is written by Jarik Oosting

With a passion for SEO and an unmatched drive for results, Jarik Oosting is the driving force behind SmartRanking. With over 15 years of experience in the field, he has built an extensive body of knowledge spanning technical SEO to complex site migrations. As the founder of SmartRanking, he has assembled a team of like-minded SEO specialists who help businesses achieve sustainable online growth.

His academic background in information science at the University of Groningen, with a specialization in natural language processing, gives him a unique perspective on the world of SEO. For Jarik, it’s not just about visibility in search engines, it’s about sharing knowledge and guiding businesses toward sustainable online success. That mission also led him to write a Dutch book about GEO.

More about SmartRanking