What is duplicate content?
Duplicate content means pages, articles or products on a site that carry content identical to another page, article or product.
Plenty of people assume content duplication only means copying text from other sites. In reality most Google penalties for duplicate content come from duplicate content within your own site.
Think about it: every page on your site has the same menu, the same footer, the same sidebar. So every page on your site carries a fair amount of duplicate content.
It’s true that Google is able to weigh the different parts of a page and give them more or less importance, but a percentage of duplicate content exists on every page of your site. On duplication specifically, Google gives far less weight (proportionally) to what’s written in the footer than to what’s in the body text.
Want to become an SEO guru?
Thinking about starting a career in SEO?
Got a site you want to optimize for search yourself?
Find out all about our SEO courses!
I want more information
An example of duplicate content: ecommerce
Cases of duplicate content turn up very often in ecommerce, for instance (see the ecommerce SEO lesson).
Picture a hypothetical store selling Converse All Star shoes. And think how many All Star colours there are… a great many.
Inside that store you’ll find the “Chuck” model in red, blue, yellow, black, white, cream and so on.
So you end up with a series of products (Converse All Star, Chuck model) loaded into the store with the exact same description (because the model is identical), the same shoe sizes, the same price, the same menu, the same footer, the same sidebars, and so on. The only difference is the colour.
As you’ll appreciate, this is 100% a case of content duplication inside your own site.
SEO and Google penalties for duplicate content
Duplicate content is a very important subject in SEO.
When search engine spiders crawl your site’s URLs and find duplicate content, they move to apply a penalty.
What can happen:
- if the crawler finds one page with duplicate content, it penalizes that page’s ranking;
- if the crawler finds many pages with duplicate content, it penalizes the ranking of the whole site.
Which is why avoiding duplicate content matters so much.

How to avoid duplicate content: write original material and don’t copy
The first solution is also the most obvious one.
Since the beginning of time Google has told us “Content is King”. So the content of the page is the single most important thing for ranking.
Content has to be informative, it has to answer a need the user had when they searched on Google, and above all it has to be original.
So copying content from another page — a competitor’s, say — is not the way to go. Do it and you’ll be penalized.
It’s true that practically everything has been written about everything, so if you cover a topic that’s been covered before, something similar will exist.
Google doesn’t penalize you if part of your page, article or product resembles part of another site. Google judges content duplication as a percentage. If you have 20% duplicate content, you have 80% original content, and you won’t be penalized for it. As you’ll know, Google gives no details about its algorithm, so we don’t know the threshold beyond which it treats content as duplicate.
Some advice on not copying content:
- search Google for the keyword you want to write an article about;
- open the first 10 results;
- go through the pages Google served you one by one, use them to work out which topics to cover, and build yourself a paragraph structure;
- take the parts that interest you, but rewrite them completely in your own words. Google can’t grasp the meaning of a passage — it weighs the words that are written. So “we grabbed a quick beer” and “we had a beer” say the same thing, but to Google they’re two completely different sentences.
Here’s a question I get asked constantly: if you translate an English text word for word into Italian, is that duplicate content?
No. As we wrote above, Google doesn’t grasp the meaning of a text — it reads the words. (True, it has learned to handle long tails better and group them semantically, but we’re still a long way from it understanding meaning.) So “Red shoes” and “Scarpe rosse” are, to Google, 100% different things.
Want an SEO analysis of your site?
Want to know how many keywords you rank for on Google?
Want to know your website’s organic traffic?
Want to know who your competitors are and how they’re performing?
Want to know the number of backlinks pointing to your site?
No problem, we’ll tell you. FREE!
Content duplication and Rel Canonical
A more “technical” way of avoiding duplicate content is the Rel Canonical.
What is the Rel Canonical tag?
First things first: the rel canonical is a tag.
A rel canonical is a way of telling search engines that a specific URL (a page, article or product of yours) is a copy of a “main page”.
Using the canonical tag solves problems caused by identical or “duplicate” content.
In practice, the rel canonical tag (rel=canonical) tells search engines which version of a URL matters to you (call it the mother) and which URLs are derived from it (the children).
So going back to the store selling All Star Chucks: you can tell Google that the product that matters to you is the red one (at, say, www.website.com/chuck/red) and that the other URLs (www.website.com/chuck/yellow, /green, /white and so on) are derived pages. That way you avoid duplicate content.
Using the SEO canonical tag helps you keep your duplicate content under control.
Image from moz.com
Rel Canonical: things worth considering
Duplicate content problems can get extremely complicated, but as we’ve seen they can be resolved with canonicalization.
Here are some important things to keep in mind when using the canonical tag:
- Canonical tags can be self-referential
It’s fine for a rel canonical to point at the current URL. In other words, if URLs X, Y and Z are duplicates and X is the canonical version, it’s correct to place a rel canonical pointing to X on URL X itself.
So on your important page you can put a rel canonical that points to that very page. It sounds obvious, but it’s a common source of confusion.
- Canonicalize your home page proactively
Home page duplicates are very common, and people can link to your homepage in many ways you can’t control, so it’s usually a good idea to put a canonical tag on your home page template to head off unexpected problems.
- Check your dynamic canonical tags
Faulty code sometimes makes a site write a different canonical tag automatically for every version of a URL — which defeats the point of the tag entirely. Check your URLs carefully, particularly on ecommerce sites and CMSs.
- Avoid mixed signals
Search engines may ignore a canonical tag or read it wrongly if you send mixed signals.
In other words, don’t canonicalize page A to page B and then page B to page A.
By the same token, don’t canonicalize page A to page B and then 301 redirect page B to page A.
It’s generally a bad idea to chain canonical tags (A to B, B to C, C to D) if you can avoid it. Send clear signals, or you force search engines into bad choices.
- Canonical tags across domains
If you own and control two sites, you can use the canonical tag between the two domains.
Say you’re a publisher that often runs the same article across half a dozen sites. Using the canonical tag concentrates your ranking power on one site.
Bear in mind that canonicalization will stop the non-canonical sites ranking, so make sure that fits your business case.
Canonical tags and 301 redirects
Are rel canonical and the 301 redirect the same thing? Absolutely not.
What’s the difference between rel canonical and a 301 redirect?
As we’ve seen, the rel canonical tells Google which URL matters to you without removing the less important one.
A 301 redirect, on the other hand, forwards the URL that matters less to the one that matters more. Click the less important URL and you’re taken to the important one.
So, simplifying heavily to make the point (and not entirely accurately), you could say the 301 redirect deletes the less important URL, while the rel canonical keeps both and tells Google which one counts.
A common SEO question is whether canonical tags pass link juice (PageRank, authority and so on) the way 301 redirects do.
The tests that have been run suggest they do.

Useful tools for canonicalization
Several tools can help if you need to do canonicalization work.
- If your site was built on the WordPress CMS, you can use a plugin called Yoast SEO (official site).
Among the advanced options there’s a field for entering the URL of the main page to canonicalize to.
- A useful tool for finding duplicate content is Google Search Console (official site).
You can look into individual pages to see whether they’ve lost a lot of ground in a short time. If you find that pattern, there’s a good chance you’ve been penalized for duplicate content.
Conclusion
In this article you’ve learned what duplicate content is, why it turns up on your own site far more often than you’d expect, what it costs you in rankings, and how the rel canonical tag resolves it — along with how it differs from a 301 redirect and which tools help you put it in place.
In the other lessons of our SEO Academy you’ll find plenty more that will help with your SEO project.
Want to optimize your site, get found by your customers on Google and beat the competition? Get in touch! We’ll work out the best SEO strategy to finally do what you haven’t managed so far.
Want to sharpen your SEO knowledge? Get in touch! We’ll find the right course for you.
Want to ask us something or just talk to us? Get in touch! That’s what we’re here for.
