Duplicate Pages and How a Canonical URL Is Chosen
A canonical URL is the address Google chooses as the representative copy when the same content is available at more than one URL. Google calls the process canonicalization, or deduplication. Some duplicate content is normal, and Google says it is not, by itself, a violation of spam policies [1]. The problem is the mess: people cannot tell which link is the real one, and you cannot tell which URL earned the visit.
What counts as a duplicate?
Google lists the usual ones. Region variants that are the same language at two URLs. A mobile version and a desktop version. HTTP and HTTPS. Sorting and filtering a category. A demo or staging copy left where crawlers can reach it [1]. Different languages are duplicates only when the body is still the same language and only the chrome is translated [1]. A real translation is not a duplicate of the original.
How does Google pick?
It groups pages whose main content looks the same, then chooses the one that looks most complete and useful. That copy is crawled more often. The duplicates are crawled less, to save your server [1]. Signals include whether the page is HTTPS, whether redirects point at it, whether it is in the sitemap, and whether you published a rel=canonical link. You can state a preference. Google may choose a different URL anyway. The preference is a hint, not a rule [1].
Which redirect supports that hint?
A permanent one. A 301 or 308 is a permanent canonical signal. A 302, 303, or 307 is temporary and is not [2]. If http still answers as its own page, or a 302 sends people to https "for now" a year later, you are feeding the duplicate instead of closing it. Server-side redirects are the ones Google prefers [2].
Does the sitemap settle it?
It is one of the signals, not the decision [1]. A sitemap also does not guarantee crawl or index [3]. Listing both the http URL and the https URL, or every sorted copy of a collection, asks Google to consider all of them. List the URL you want as the representative. Take the copies off the list.
What if the extra copy is a staging site?
Google names an accidentally public demo as one way duplicates happen [1]. A robots.txt disallow does not reliably keep that demo out of results [4]. noindex works only when Google can crawl the page and see it. If robots.txt blocks the demo, the noindex is invisible [5]. Close the demo, or let the crawler see a noindex. Do not leave it up and hope the canonical tag on the real site wins.
What if one of the copies is an error?
A 4xx other than 429 is ignored, and a URL that stays in that state is dropped over time [6]. A soft 404, where the server says success and the page is empty or an error, is also a failed page [6]. Do not canonical-tag a real product at a URL that 404s. Fix the address, or redirect it if there is a true replacement [2].
What about a regional copy in the same language?
Google says to use both a canonical signal and hreflang when the same language is published for two regions at two URLs [1]. hreflang says who the copy is for. The canonical says which URL is the representative when the pages are still duplicates. One does not replace the other.
Will search always link to the canonical?
Usually, and not always. Google says a result usually points at the canonical page, unless another copy is a better match for that searcher. A person on a phone may see the mobile URL even when the desktop URL is the canonical [1]. That is a display choice. It is not a reason to publish a third copy.
- HTTPS — one copy, with the other version redirecting.
- Filters — do not give every sort order its own indexable URL.
- Canonical tag — a hint. Google can pick a different URL.
- Sitemap — list the representative URL, not every twin.
- Demo — noindex it in a way the crawler can see, or take it down.
Where does this show up on a small site?
After a redesign that left the old paths alive, which is one of the signs a redesign needs a URL list. A store with filters creates duplicates faster than a brochure site, so the catalog needs a rule for sort URLs. The build should pick one hostname and stick to it. The launch checklist should include "one address for the homepage."
What should you compare this week?
Open the homepage with http and with https, and with and without www. If more than one of those answers as a real page, you have the duplicate Google describes [1]. Pick one, redirect the others with a 301 or 308 [2], and put only the winner in the sitemap [3]. Then check that a staging link is not the second homepage [1].
Sources & references
- Google Search Central: URL canonicalization.
- Google Search Central: Redirects and Google Search.
- Google Search Central: Learn about sitemaps.
- Google Search Central: Introduction to robots.txt.
- Google Search Central: Block search indexing with noindex.
- Google Search Central: HTTP status codes and network errors.
Canonicalization rules are on the linked Google page. Read the current page before you add tags across a site.
If the same page is in search under two addresses, we can see which one Google is treating as the main copy.
Talk to us arrow_right_alt