Glossary

Canonical URL

Canonical URL is the one address you want search engines to treat as the main version of a page that is reachable at several URLs. You declare it with a link element whose rel attribute is canonical, placed in the head of the page. The link relation is described in RFC 6596, an informational RFC published in April 2012. Google treats the tag as a strong hint, not a command.

How it works

The same content often has many addresses: tracking parameters, sort and filter options, http and https, a trailing slash, www and non-www. Without guidance, a search engine picks one itself and may split ranking signals such as inbound links across the copies. A canonical URL tells it which one to keep.

  • Link element: <link rel="canonical" href="https://example.com/shoes"> in the head of every duplicate page, pointing to the preferred URL.
  • Redirect: a permanent redirect from the duplicate to the preferred URL is a stronger signal, because visitors move too.
  • Sitemap: listing only preferred URLs in the sitemap is a weaker signal that supports the others.
  • Absolute URLs: use the full address with scheme and host. Google supports relative paths but does not recommend them, because they can cause problems later.

The script below shows why duplicates appear. It removes a tracking parameter with the WHATWG URL parser, so two orderings of the same query collapse to one address, while the http version and the trailing-slash version stay different.

const urls = [
  "https://example.com/shoes?utm_source=news&color=red",
  "https://example.com/shoes?color=red&utm_source=news",
  "http://example.com/shoes?color=red",
  "https://example.com/shoes/?color=red",
];
for (const u of urls) {
  const x = new URL(u);
  x.searchParams.delete("utm_source");
  console.log(u, "->", x.href);
}
https://example.com/shoes?utm_source=news&color=red -> https://example.com/shoes?color=red
https://example.com/shoes?color=red&utm_source=news -> https://example.com/shoes?color=red
http://example.com/shoes?color=red -> http://example.com/shoes?color=red
https://example.com/shoes/?color=red -> https://example.com/shoes/?color=red

Four addresses, three distinct results. Stripping parameters is not enough: you still choose https, a trailing-slash rule and a host, then point every variant at that choice.

Does Google always follow the canonical tag?

No. Google calls a canonical preference a hint, not a rule, and may choose a different page as canonical. Redirects, internal links, HTTPS and the sitemap all feed that choice. Keep those signals consistent so they all name the same URL.

Common pitfalls

  • Relative URLs in the href: a value like /shoes works, but it can point at the wrong host if a test copy of the site gets crawled. Write the full https address.
  • Canonical pointing to an error page: the target must exist, not return an error or a soft 404, and must not carry a noindex robots tag.
  • Conflicting signals: a canonical to page A, a sitemap listing page B and internal links to page C make the engine choose for you. Make all three agree.
  • Multiple canonical tags: Google says that when a page has more than one rel=canonical, all of them are ignored. Output exactly one, and check that plugins do not add a second.
  • Canonicalizing different content: Google says not to point the pages of a paginated list at page 1. Give each page its own canonical, and canonicalize only true duplicates.
  • Blocking the duplicate in robots.txt: a crawler that cannot fetch the page cannot read its canonical tag, and Google may still index the blocked URL. Let it crawl, or use a redirect.

Related terms

  • Hreflang — marks language and region versions. Google says each should name a canonical in the same language.
  • robots.txt — controls crawling, and blocking a page there hides its canonical tag.
  • Structured data — separate markup that should describe the canonical page.
  • HTTP — the protocol whose redirect status codes consolidate duplicate URLs.
  • Open Graph — the og:url tag should match the canonical URL.

See also