robots.txt is a plain-text file served at /robots.txt on a host that tells crawlers which URL paths they may request. The format is the Robots Exclusion Protocol, standardized as RFC 9309 in September 2022. The file is a request that well-behaved crawlers honor, not access control. It controls crawling, not whether a URL appears in search results.
The file is a list of groups. A group starts with one or more user-agent lines and is followed by allow and disallow rules.
group. If there is no group, no rules apply.# starts a comment, * matches zero or more characters, and $ marks the end of the match pattern.This Python code applies the RFC's longest-match rule to three rules and prints the verdicts:
import re
rules = [("disallow", "/private/"), ("allow", "/private/press/"), ("disallow", "/*.pdf$")]
def matches(pat, path):
rx = re.escape(pat).replace(r"\*", ".*")
if rx.endswith(r"\$"):
rx = rx[:-2] + "$"
return re.match(rx, path) is not None
def allowed(path):
best_len, best_allow = -1, True
for kind, pat in rules:
if matches(pat, path):
n = len(pat)
if n > best_len or (n == best_len and kind == "allow"):
best_len, best_allow = n, kind == "allow"
return best_allow
for p in ["/private/a.html", "/private/press/kit.html", "/blog/", "/files/report.pdf", "/files/report.pdf?v=2"]:
print(f"{p:26} {'allowed' if allowed(p) else 'blocked'}")
/private/a.html blocked
/private/press/kit.html allowed
/blog/ allowed
/files/report.pdf blocked
/files/report.pdf?v=2 allowed
A 4xx response means crawlers may access anything, while a 5xx or network failure means they must assume everything is disallowed. RFC 9309 treats the 400 to 499 range as "unavailable" and 500 to 599 as "unreachable". Crawlers should follow at least five consecutive redirects. They should not use a cached copy for more than 24 hours unless the file is unreachable, and must parse at least 500 KiB.
No. Google says robots.txt is mainly for managing crawl load and is not a way to keep a page out of Google. A blocked URL can still be indexed if other pages link to it, though the result has no description. Use noindex or password protection, and do not block a page you want crawlers to read the noindex on.
Disallow: /private/ first and Allow: /private/press/ second, it blocked /private/press/kit.html. Reversing the lines allowed it, while the RFC 9309 longest-match rule allows it in either order.Disallow: /.pdf blocks any URL containing .pdf, while /.pdf$ stops at the end. In the run above, /files/report.pdf?v=2 was allowed by the anchored rule.