Key takeaways
- A crawler obeys exactly one group: the one naming its user-agent most specifically, otherwise the * group. Rules from other groups don't add up.
- Inside that group the longest matching path pattern wins. "Allow: /admin/help" beats "Disallow: /admin" for /admin/help/faq.
- When an Allow and a Disallow are equally long, Google picks Allow (the least restrictive rule).
- * matches any characters and $ anchors the end of the URL. Matching is case-sensitive and starts at the beginning of the path.
- robots.txt controls crawling, not indexing. A blocked URL can still appear in Google if other sites link to it.
How Google decides whether a URL is blocked
Google's rules are now standardized in RFC 9309 (the Robots Exclusion Protocol, published in 2022) plus a few documented Google behaviors. The tester implements them in three steps:
- 1.Group selection. The file is split into groups — one or more User-agent lines followed by rules. The crawler picks the group whose user-agent token matches it most specifically; groups naming the same token are merged. Google's sub-crawlers fall back to their parent, so Googlebot-Image uses a Googlebot group if there's no Googlebot-Image group. If nothing matches, the * group applies; if there's no * group either, everything is allowed.
- 2.Rule matching. Each Allow and Disallow pattern is compared with the URL's path and query string from the start. * matches any sequence of characters, $ at the end means "the URL ends here", and characters are compared after percent-encoding normalization.
- 3.Precedence. The matching rule with the longest pattern wins. On a tie between Allow and Disallow, Allow wins. No matching rule means allowed. /robots.txt itself is always allowed.
Worked examples
| Rules in the group | URL | Result | Why |
|---|---|---|---|
| Disallow: /admin, Allow: /admin/help | /admin/help/faq | Allowed | "/admin/help" (11 chars) is longer than "/admin" (6). |
| Disallow: /admin | /administrator | Blocked | Patterns are prefixes: /admin matches /administrator. Use /admin/ to block only the folder. |
| Disallow: /*.pdf$ | /files/report.pdf?v=2 | Allowed | $ requires the URL to end in .pdf; the query string breaks the match. |
| Disallow: /*?sort= | /shop?sort=price | Blocked | * spans "shop", then ?sort= matches literally. |
| Disallow: /page, Allow: /page | /page | Allowed | Equal length — Google applies the least restrictive rule. |
| Disallow: /Private | /private | Allowed | Paths are case-sensitive. |
| Disallow: (empty) | /anything | Allowed | An empty Disallow matches nothing. |
The most surprising rule: groups don't inherit from *
If your file has a User-agent: * group that disallows /admin and a separate User-agent: Googlebot group that disallows /staging, Googlebot may crawl /admin. It only reads its own group. This catches many sites that add a Googlebot or GPTBot section and accidentally open up everything they had blocked for everyone else. Copy the shared rules into each named group.
How to test your robots.txt
- 1.Open https://yourdomain.com/robots.txt in a browser and paste the whole file into the tester. Each hostname (www, a subdomain, http vs https) has its own robots.txt.
- 2.Enter a URL or path. Include the query string if your rules use ? or *, because patterns match against path + query.
- 3.Enter a crawler. Use a token like Googlebot, Bingbot, GPTBot or ClaudeBot, or paste a full user-agent string from your logs — the tester finds the group whose token appears in it.
- 4.Read the verdict, the group it obeyed and the deciding rule (shaded in the parsed file). If the result isn't what you intended, add a longer, more specific rule rather than reordering lines — order doesn't matter to Google.
- 5.Fix the syntax warnings: rules before any User-agent line, paths not starting with /, and directives Google ignores such as crawl-delay and noindex.
Common robots.txt mistakes
- Blocking a page to remove it from Google. Disallow stops crawling, so Google can't see a noindex tag on the page and may keep the URL indexed from links. Allow crawling and use noindex instead.
- Blocking CSS and JavaScript. Google renders pages; blocking /assets/ or /_next/ can make it see a broken page.
- Relying on order. Google ignores the order of rules within a group — only length and type matter. (Some other crawlers use first-match, which is why specific rules are safest.)
- Using noindex or crawl-delay in robots.txt for Google. Google dropped noindex support in 2019 and has never honored crawl-delay; Bing and some others do honor crawl-delay.
- Forgetting trailing slashes. Disallow: /blog also blocks /blog-post-ideas and /blogroll.
- Assuming every bot complies. Well-behaved crawlers do; scrapers don't. robots.txt is a public request, not access control.
Testing rules for AI crawlers
The same semantics apply to GPTBot, ClaudeBot, PerplexityBot and other AI crawlers that honor robots.txt (RFC 9309 is the standard they cite). Build AI-specific groups with the AI crawler robots.txt generator, then paste the result here to confirm each bot gets the access you intended. To see which crawlers really visit and what they fetch, use server-side AI crawler tracking — robots.txt only states your intent.
Frequently asked questions
How does Google handle conflicting Allow and Disallow rules?
Google applies the most specific rule, meaning the one with the longest matching path pattern. If an Allow and a Disallow rule are exactly as long, Google uses the least restrictive one, which is Allow. The order of the rules in the file doesn't matter.
Does a specific User-agent group inherit the rules from User-agent: *?
No. A crawler follows only the single most specific group that matches it. If you add a Googlebot group, Googlebot ignores everything in the * group, so repeat any shared rules inside the named group.
Are robots.txt paths case-sensitive?
Yes. The path in Allow and Disallow rules is case-sensitive, so Disallow: /Private doesn't block /private. User-agent names, on the other hand, are matched case-insensitively.
What do * and $ mean in robots.txt?
* matches any sequence of characters, including none, and $ at the end of a pattern means the URL must end there. So Disallow: /*.pdf$ blocks URLs ending in .pdf but not /file.pdf?download=1.
Does blocking a URL in robots.txt remove it from Google?
No. robots.txt prevents crawling, not indexing. Google can still index a blocked URL without its content if other pages link to it. To remove a page from results, allow crawling and add a noindex meta tag or X-Robots-Tag header.
Why does my test say allowed when there's a Disallow rule for that folder?
Most often because a longer Allow rule also matches, because the crawler obeys a different group than you expected (a named group instead of *), or because of case or a trailing-slash difference. The tester shows the group it used and the deciding rule so you can see which.
Does the tester fetch my robots.txt?
No. You paste the file, and everything runs in your browser. Nothing is sent to our servers, which also means you can test a draft before publishing it.
Related tools and guides
- AI crawler robots.txt generatorBlock training bots while staying visible in AI search.
- User-agent bot checkerIdentify the crawler behind a UA string.
- llms.txt generatorA Markdown map of your site for AI assistants.
- SERP snippet previewCheck title and description truncation.
- AI crawler trackingSee which bots fetch which pages.
- Analytics glossaryPlain-English definitions of SEO and analytics terms.