๐Ÿ“š General & Other

Test Whether robots.txt Blocks a URL

A robots.txt tester takes a set of rules and a URL and reports whether a crawler would be allowed to fetch it.

A robots.txt tester takes a set of rules and a URL and reports whether a crawler would be allowed to fetch it.

Why testing beats reading

robots.txt matching rules are less intuitive than they look. Which group applies to which crawler, how wildcards behave, and which of several conflicting rules wins are all easy to get wrong by inspection.

What blocking actually does

Disallow prevents crawling, not indexing. A blocked URL can still appear in results if other pages link to it โ€” without a description, because the crawler was never allowed to read it.

The rule worth remembering

To keep a page out of the index, allow crawling and use a noindex directive. Blocking a page in robots.txt prevents the crawler from ever seeing that directive.

Try it: Robots.txt Tester on SeoWolf's Notepad.