Glossary

Hreflang

Hreflang is an annotation, usually an attribute on a link element, that tells search engines which language and optional region each alternate version of a page targets. Google documents three ways to declare it: HTML link tags, an HTTP Link header, or a sitemap. The language is an ISO 639-1 code and the region an ISO 3166-1 Alpha 2 code, for example en-US. The special value x-default marks the fallback page.

How it works

Each page lists every version of itself, including itself, with rel="alternate" and an hreflang value. The format is a language code, then an optional dash and a region code.

  • Language only: hreflang="de" targets German speakers anywhere.
  • Language and region: hreflang="en-GB" targets English speakers in the United Kingdom. A region alone is not valid.
  • x-default: matches any language not listed. Google recommends it especially for language or country selectors and auto-redirecting home pages.
  • Fully qualified URLs: each alternate URL needs its scheme and host, so a path like /foo or a scheme-less //example.com/foo is not accepted.
  • Return links: if page X points to page Y, page Y must point back to X. Google ignores the pair if the two pages do not both point to each other.

The three methods are equivalent to Google, and you can use any one of them. Using more than one gives no extra benefit in Search. Google says the HTML tags must sit in a well-formed head section.

from html.parser import HTMLParser

class Alt(HTMLParser):
    def __init__(self):
        super().__init__()
        self.alts = {}
    def handle_starttag(self, tag, attrs):
        a = dict(attrs)
        if tag == "link" and a.get("rel") == "alternate" and "hreflang" in a:
            self.alts[a["hreflang"]] = a["href"]

pages = {
    "https://example.com/en/": '''<link rel="alternate" hreflang="en" href="https://example.com/en/">
<link rel="alternate" hreflang="de" href="https://example.com/de/">
<link rel="alternate" hreflang="x-default" href="https://example.com/">''',
    "https://example.com/de/": '''<link rel="alternate" hreflang="de" href="https://example.com/de/">''',
}
parsed = {}
for url, html in pages.items():
    p = Alt(); p.feed(html); parsed[url] = p.alts
for url, alts in parsed.items():
    print(url, "->", alts)
    for lang, target in alts.items():
        if target != url and target in parsed and url not in parsed[target].values():
            print("  no return link from", target, "to", url)

Output from Python 3:

https://example.com/en/ -> {'en': 'https://example.com/en/', 'de': 'https://example.com/de/', 'x-default': 'https://example.com/'}
  no return link from https://example.com/de/ to https://example.com/en/
https://example.com/de/ -> {'de': 'https://example.com/de/'}

The English page lists the German page, but the German page does not list the English one. Missing return links are one of the most common hreflang errors, and Google ignores an annotation that is not confirmed from the other page.

What does hreflang x-default do?

The x-default value names the page to show when no listed language or region fits the user. In Google's example it points at a generic page such as a country selector. It is one more link element added to the same set, for instance hreflang="x-default" with the root URL. Since every version lists the full set, include it on each one.

Common pitfalls

  • Missing return links: the checker above flags the German page. Add the full set of alternates, including a self-reference, to every version.
  • Region without language: hreflang="US" is invalid. Use en-US.
  • Unsupported codes: Google supports only ISO 639-1 languages and ISO 3166-1 Alpha 2 regions, so es-419 is not supported. A language code by itself, such as es, is valid.
  • Relative URLs: hreflang="en" with href="/en/" is not accepted. Write the full https URL.
  • Case worries: Google treats the value as case-insensitive, so en-gb and en-GB both work. Uppercase regions follow the ISO convention.

Related terms

  • Canonical URL — the other per-page URL signal declared in the head of a page
  • Structured data — separate markup that describes a page to search engines
  • robots.txt — controls crawling, a separate mechanism from language targeting

See also

  • Term: Canonical URL — the closest related term, covering which URL a page declares as preferred