๐Ÿ“š General & Other

robots.txt Matching Rules Are Not What You Expect

The format looks simple and its matching semantics are genuinely surprising in several places.

The format looks simple and its matching semantics are genuinely surprising in several places.

Only one group applies

A crawler picks the single most specific matching User-agent group and ignores every other group entirely, including the wildcard one. Rules in User-agent: * do not apply to Googlebot if a User-agent: Googlebot group exists anywhere in the file โ€” a very common misunderstanding that leaves intended rules silently inactive.

Longest match wins, not first

Where Allow and Disallow both match a URL, the rule with the longer path pattern wins regardless of order. Allow: /blog/public/ beats Disallow: /blog/ for a URL under the former.

Wildcards are limited

* matches any sequence and $ anchors the end. There is no character class, no alternation and no capture. Anything more expressive has to be handled elsewhere.

It is per-origin and per-port

https://example.com and http://example.com have separate robots.txt files, as does each subdomain and each non-standard port. A rule on one does not govern the others.

Try it: Robots.txt Tester on SeoWolf's Notepad.