Duplicate content

Updated: 3 min read SEOFuxx editorial team

Duplicate content means that the same or very similar content can be reached at several addresses, on the same website or on different ones. Most of the time it arises unintentionally from the technology. Google does not punish it across the board, but it picks one version for the search results, and that is not necessarily the one you want.

Is it a penalty?

No. In Google's documentation, duplicate content is not a violation in itself. The spam policies target other things, such as scraping: copying content from others without adding value of your own, changing it only slightly, or taking over feeds from other sites without benefit to visitors. Mass-produced, barely distinguishable pages (doorway pages) are also a violation. Anyone who does none of this has no penalty problem with technical duplicates but a selection problem: Google decides which address it shows.

Typical causes

  • several spellings of the same address: http and https, with and without www, with and without a trailing slash
  • URL parameters for tracking, sorting, filters or sessions, for example ?gclid=
  • one page in several categories or paths
  • print versions and separate mobile addresses
  • products with almost identical descriptions or manufacturer texts
  • copied content on other websites

Why you should specify an address

If you do not give an address, Google picks what it considers the best one. If you specify it, you get benefits:

  • You decide which address appears in the search results, for example the clean one rather than the one with a tracking parameter.
  • The signals, such as backlinks to different variants, come together on one address.
  • Metrics are easier to read when a piece of content has only one address.
  • The crawler spends less time on duplicates (see crawl budget).

Methods compared

MethodStrengthUse
301 redirectstrong signalwhen the duplicate page should go away
rel="canonical" in the HTMLstrong signalwhen both pages stay reachable, HTML pages only
rel="canonical" as an HTTP headerstrong signalfor files without HTML such as PDFs
Inclusion in the sitemapweak signalas a supplement to support the preferred address

The methods can be combined. Of the two canonical variants you should use only one, because both together are more error-prone. Google also recommends a self-referencing canonical, that is, a page pointing to itself.

<!-- in the head of the duplicate AND the main page -->
<link rel="canonical" href="https://www.yourdomain.com/your-page">
# Apache: redirect http to https with 301
RewriteEngine On
RewriteCond %{HTTPS} off
RewriteRule ^(.*)$ https://www.yourdomain.com/$1 [R=301,L]

How Google picks the address

Google evaluates the signals mentioned and adds characteristics of the website. It generally prefers https over http, and it prefers addresses that are part of a hreflang cluster. A canonical is therefore a strong hint, but not a guarantee.

Multilingual pages and similar cases

Pages in different languages are not duplicates, even if they correspond in content. Link them with hreflang. For very similar pages in the same language for different countries, hreflang helps as well. Pure variants that only a filter produces, on the other hand, you should consolidate or, if they have no search value, mark with noindex. With a canonical pointing to a different page, the variant disappears from the results.

Common mistakes

  • Canonical to a completely different page: The content should match the page you point to.
  • Canonical and noindex at the same time: That sends contradictory signals.
  • Duplicate addresses in the sitemap: The sitemap should contain only the main address.
  • Canonical in the body instead of the head: The element belongs in the head of the page.
  • Reacting to penalties that do not exist: More important is that the right version appears in the results.

Sources

Is this content helpful?

·