Glossary

robots.txt

robots.txt is a plain-text file served at /robots.txt on a host that tells crawlers which URL paths they may request. The format is the Robots Exclusion Protocol, standardized as RFC 9309 in September 2022. The file is a request that well-behaved crawlers honor, not access control. It controls crawling, not whether a URL appears in search results.

How it works

The file is a list of groups. A group starts with one or more user-agent lines and is followed by allow and disallow rules.

  • Location: the file must be named /robots.txt in lowercase at the top level of the host, UTF-8 encoded, served as text/plain.
  • Matching a group: the crawler picks the group whose user-agent value matches its product token, case-insensitively. If none match, it uses the group. If there is no group, no rules apply.
  • Paths: matching should be case-sensitive and must start at the first octet of the path. The most specific match wins, meaning the one with the most octets. If an allow and a disallow rule are equivalent, allow should be used.
  • Special characters: # starts a comment, * matches zero or more characters, and $ marks the end of the match pattern.
  • Other lines: a Sitemap line is not part of the core rules, but RFC 9309 says crawlers may interpret it and it must not end a group.

This Python code applies the RFC's longest-match rule to three rules and prints the verdicts:

import re
rules = [("disallow", "/private/"), ("allow", "/private/press/"), ("disallow", "/*.pdf$")]

def matches(pat, path):
    rx = re.escape(pat).replace(r"\*", ".*")
    if rx.endswith(r"\$"):
        rx = rx[:-2] + "$"
    return re.match(rx, path) is not None

def allowed(path):
    best_len, best_allow = -1, True
    for kind, pat in rules:
        if matches(pat, path):
            n = len(pat)
            if n > best_len or (n == best_len and kind == "allow"):
                best_len, best_allow = n, kind == "allow"
    return best_allow

for p in ["/private/a.html", "/private/press/kit.html", "/blog/", "/files/report.pdf", "/files/report.pdf?v=2"]:
    print(f"{p:26} {'allowed' if allowed(p) else 'blocked'}")
/private/a.html            blocked
/private/press/kit.html    allowed
/blog/                     allowed
/files/report.pdf          blocked
/files/report.pdf?v=2      allowed

What happens when robots.txt returns an error?

A 4xx response means crawlers may access anything, while a 5xx or network failure means they must assume everything is disallowed. RFC 9309 treats the 400 to 499 range as "unavailable" and 500 to 599 as "unreachable". Crawlers should follow at least five consecutive redirects. They should not use a cached copy for more than 24 hours unless the file is unreachable, and must parse at least 500 KiB.

Does Disallow remove a page from Google?

No. Google says robots.txt is mainly for managing crawl load and is not a way to keep a page out of Google. A blocked URL can still be indexed if other pages link to it, though the result has no description. Use noindex or password protection, and do not block a page you want crawlers to read the noindex on.

Common pitfalls

  • Using it for secrecy: RFC 9309 says listing paths in the file exposes them publicly. Protect sensitive paths with authentication.
  • Blocking the page that carries noindex: the crawler never fetches the page, so it never sees the tag.
  • Writing noindex in robots.txt: Google retired handling of unsupported rules such as noindex on September 1, 2019.
  • Trusting parser libraries to match the RFC: Python's urllib.robotparser applies rules in file order. With Disallow: /private/ first and Allow: /private/press/ second, it blocked /private/press/kit.html. Reversing the lines allowed it, while the RFC 9309 longest-match rule allows it in either order.
  • Forgetting the dollar sign: Disallow: /.pdf blocks any URL containing .pdf, while /.pdf$ stops at the end. In the run above, /files/report.pdf?v=2 was allowed by the anchored rule.
  • Serving a 5xx: a server error during a deploy makes compliant crawlers stop crawling the whole host.

Related terms

  • Canonical URL — a page-level hint about the preferred URL, separate from crawl rules.
  • HTTP — the status codes decide how crawlers treat a missing or failing file.
  • HTTPS — RFC 9309 defines the file's URI as scheme, authority and /robots.txt, so it is fetched per origin.
  • Structured data — markup crawlers read only if they are allowed to fetch the page.

See also