Technical SEO · Analysis

How Google picks a canonical URL, and why your preference is only a hint

When the same content is reachable at several URLs, Google clusters them and chooses one to show. You can state a preference, but the choice stays with Google, so the useful work is making every signal point the same way.

By Dean Cruddace · Published · Updated · Read 6 min

Key findings

  • Google treats a canonical preference as a hint, not a rule: it weighs redirects, rel=canonical, sitemap inclusion and whether a URL is served over HTTPS, and may pick a different URL from the one you chose.
  • Its documentation ranks redirects and rel=canonical as strong signals and sitemap inclusion as a weak one, and says combining methods raises the chance that your preferred URL is chosen.
  • The Google-selected canonical in URL Inspection is the only direct evidence of the outcome, and it comes from indexed data: the live test cannot predict it.

Most sites can show the same page at more than one address. A product might be reachable with and without a tracking parameter, on HTTP and HTTPS, or through a sorted and an unsorted category view. Google has to notice, because it wants to show one version of the content, not five. This article sets out how Google describes that choice, which signals it lists, and what we think follows for anyone managing a site.

Why one page ends up at several URLs

Google’s documentation on URL canonicalisation defines canonicalisation as the process of selecting the representative URL of a piece of content. It lists the usual causes of duplication: region variants (separate US and UK URLs with essentially the same content in the same language), device variants, protocol variants such as HTTP and HTTPS, site functions such as the results of sorting and filtering on a category page, and accidental variants such as a demo version of a site left open to crawlers.

The page is careful on one point. Some duplicate content on a site is normal and is not a violation of Google’s spam policies. The cost is practical: people may wonder which page is the right one, and it can become harder to track how a piece of content performs in search results. So the aim is not to eliminate duplicates at all costs. It is to make sure the version Google shows is one you would be happy to have shown.

How Google clusters duplicates and chooses one

According to the same page, Google determines the primary content of each page it indexes. If it finds pages that seem to be the same, or whose primary content is very similar, it clusters them together. It then picks the page that, based on the signals gathered during indexing, is the most complete and useful for search users, and marks it as the canonical. The canonical page is crawled most regularly, and duplicates are crawled less often to reduce the load on the site.

The factors that play a role, as listed there, are whether the page is served over HTTP or HTTPS, redirects, the presence of the URL in a sitemap, and rel=canonical link annotations. Google’s wording is plain: indicating a canonical preference is a hint, not a rule. A search result usually points to the canonical page, but not always. The documentation gives the example that a result will probably point to the mobile page for someone on a mobile device, even if the desktop page is the canonical.

Language matters too: different language versions count as duplicates only if the primary content is in the same language. For same-language regional variants, Google says to use both canonicalisation and hreflang.

Strong signals, weak signals and what they stack to

Google’s page on specifying a canonical lists the methods in order of how strongly they can influence the outcome:

  • Redirects are a strong signal that the target should become canonical. Google says to use permanent redirects only when deprecating a duplicate page.
  • rel=canonical annotations are a strong signal that the specified URL should become canonical. They can be a link element in the HTML head or an HTTP header, which suits files such as PDFs.
  • Sitemap inclusion is a weak signal, though it is simple to maintain on large sites.

The methods can stack, and using two or more increases the chance that your preferred URL appears in search results. Google also says none of them is required: if you specify nothing, it will identify what it considers the best version to show. There are, though, listed reasons to state a preference, including choosing which URL people see, consolidating signals such as links from other sites into one URL, simplifying tracking, and avoiding crawling time spent on duplicates.

The documentation recommends choosing either the HTML element or the HTTP header for rel=canonical, because using both is more error prone. It prefers absolute URLs to relative ones, and recommends a self-referential canonical on the canonical page itself. It also says that rel=canonical annotations carrying hreflang, lang, media or type attributes are not used for canonicalisation.

Signals that disagree, and other things to avoid

This is the part of the documentation we would pin above a developer’s desk. Its list of things not to do includes:

  • Do not specify different URLs as canonical for the same page using different techniques, for example one URL in a sitemap and another in rel=canonical.
  • Do not use robots.txt for canonicalisation. Google may still index URLs disallowed in robots.txt, without their content.
  • Do not use the URL removal tool for canonicalisation, because it hides all versions of a URL from search.
  • Do not specify a URL fragment as canonical, as Google generally does not support fragments.

Google also does not recommend noindex to steer canonical selection within a site, since it blocks the page from search completely; rel=canonical is the preferred solution. Internally, link to the canonical URL rather than a duplicate, because consistent linking helps Google understand your preference.

If the site relies on client-side rendering, the documentation says the best approach is to put the canonical in the HTML source and make sure JavaScript does not change it. If that is not possible, it advises leaving the canonical out of the source and setting it only with JavaScript, so that the information is as clear as possible.

Our reading, not Google’s wording: since the choice is a weighing of signals, a mismatch between your preference and Google’s choice is usually a prompt to look for a signal that disagrees, such as a sitemap entry, an internal link pattern or a redirect that points somewhere else. Consistency improves the odds; the documentation promises nothing more.

Checking what Google actually chose

Preferences are easy to declare and hard to verify without Google’s own data. The URL Inspection tool in Search Console shows, under page indexing, the Google-selected canonical for an inspected URL. The help page is specific about a limit: you can determine the canonical only in the indexed data, and the live test cannot predict whether the tested version will be considered canonical.

Two further cautions come from the same page. The indexed result reflects the most recently indexed version of the page, not the live one, so a recent fix will not show until Google has processed it. And if a URL redirects, the results describe the tested URL in the index, not the redirect target. There is a separate inspect link for the canonical of a redirected page.

Our reading, not Google’s wording: a sensible routine is to inspect a sample of URLs, compare the declared canonical with the Google-selected one, and group differences by pattern rather than treating each as a one-off.

Where this fits at Cultured Digital: choosing the version that counts

Technical SEO is our flagship service, and canonicals and duplication sit in its crawling and indexing group. On the technical SEO page we frame the index step as one question: which version counts. Tools report symptoms; the work is finding which step is failing, and why. For canonicals, that means reading the signals a site sends rather than only checking that a tag exists.

Our crawling and indexing guide covers the surrounding mechanics, including robots.txt, sitemaps and log files, and we write findings as tickets with the fault, the evidence, the fix and the acceptance test. If your developers have no capacity, we can make the changes. For a formal investigation, see our SEO audit page.

Further reading on canonicalisation

Written by Dean Cruddace

Founder of Cultured Digital. Working in SEO since 2001, across independent consultancy, in-house and agency roles, with a focus on technical SEO, strategy and development.

About Dean →