Enter a domain and a path. You get the answer for every crawler at once, the exact line that decided it, why that line won, and whether the crawler you asked about honours robots.txt in the first place. The rules come from RFC 9309; the matcher was tested against Google's own parser.
The file is fetched from /robots.txt at the origin you give, because one file covers one scheme, host and port. Paste a full URL and its path becomes the first path tested.
Quick Answer: how do I test whether robots.txt blocks a URL?
Enter the domain and the path in the tester above. It fetches /robots.txt from that origin and answers for every crawler at once, naming the line that decided it. When an Allow and a Disallow both match, the longer rule wins — RFC 9309 measures specificity in octets, not in line order — and only an exact tie goes to the allow.
Jessica Wright
Cybersecurity Threat Researcher
Jessica works on bot and crawler verification at TrustMyIP, including the GPTBot, ClaudeBot and Googlebot verifiers, the access log analyzer and the AI crawler blocker.
This tester was built by implementing RFC 9309 rule by rule, then checking the result against Google's own robots.txt parser, which was cloned and compiled for the purpose. The two were compared over 700,226 generated files and 3,852,034 questions of the form "may this crawler fetch this path": there is no case left where they disagree without a documented reason. Google's own test file was then extracted and run as well, and this engine agrees with 140 of its 142 assertions; the two it does not are the case where that suite hands its parser an unescaped URL, which this page escapes first. The comparison found eleven real faults in this matcher, including specificity being measured on the rule as typed rather than as compared, and the multi-token rule above being missing entirely. A further 333 offline assertions cover both example files printed in RFC 9309 itself, and 96 cover the network layer against a local server. It has not yet been run against a public site from this codebase, and the page says so under "What this tester cannot tell you".
Last reviewed 26 September 2026 · Published 26 September 2026
View all articles by Jessica WrightA robots.txt file is a plain text file at the root of a host telling automated clients which paths they may fetch. Unlike most things in this corner of the web it is a real standard: RFC 9309, published in 2022 on the IETF Standards Track after twenty-five years as a convention. This robots.txt tester implements it rule by rule.
One word matters more than any other here: robots.txt asks. It does not prevent, stop or block anything. A compliant crawler reads your file and chooses to obey it; one that ignores it fetches the page anyway. That is the design, not a flaw — a notice on a door, not a lock. If you need a lock, you have to refuse the request at the server or the edge, which is what the guide to blocking an IP address at every layer covers.
The file is made of groups. A group starts with one or more User-agent: lines, followed by the Allow: and Disallow: rules for them. A crawler obeys the group naming its own product token and ignores every other group — including the * group. That catches people out constantly: give Googlebot a group of its own and it stops reading your * group entirely.
So a tester has two jobs: does the file say what you think, and will the crawler you care about ever see it? The second is about HTTP, and it does the real damage.
Key fact: one robots.txt covers exactly one scheme, host and port. https://example.com/robots.txt says nothing about blog.example.com, and nothing about the http:// version of the same host if that answers separately instead of redirecting.
When you test robots.txt here you get four answers rather than one, because "robots.txt is not working" is four problems with four different fixes.
| Verdict | The question it answers | Who decides |
|---|---|---|
| Delivery | Did your server hand over the file at all — right status, right type, no detour? | RFC 9309 section 2.3.1 |
| Syntax | Is the robots.txt syntax valid, and does every rule you wrote actually do something? | RFC 9309 sections 2.2 and 2.5 |
| Your path | For this crawler and this path, which line wins, and why that one? | RFC 9309 section 2.2.2 |
| Reality | Will that crawler honour the answer at all? | Each vendor's own documentation |
Every finding is labelled with its authority: RFC 9309 for the standard, vendor docs for behaviour a named crawler documents and the standard does not, HTTP layer for how your server answered, and TrustMyIP view for our own opinion with its threshold beside it. Only findings against RFC 9309 make a file invalid; ours never do. That is the main reason one robots.txt checker calls a file broken and another calls it fine — most present their preferences as errors, so a robots.txt validator that never separates the two leaves you guessing which findings matter.
The per-path answer prints three things other testers leave out: the line that decided it, that rule's length in octets, and the rule it beat. If you have ever wondered why a rule you wrote is ignored, that is the answer.
Search Console will not give you this. Its robots.txt report shows the file Google last fetched and offers "Request a recrawl", but does not test a URL against the rules; its help page sends you to URL Inspection for that, which answers for Google's crawler on a property you have verified. Nothing there answers for GPTBot.
Worth knowing: if the delivery verdict is red, stop and fix that first. The syntax and path answers are withheld on purpose when what came back was not a robots.txt file — checking an error page against RFC 9309 sends you hunting for a fault in a file your server never served.
This is the question the whole tool exists to answer, and the rule is short. RFC 9309 section 2.2.2:
The most specific match found MUST be used. The most specific match is the match that has the most octets.
Octets, not lines. Not first-wins, not last-wins, not allow-beats-disallow: count the bytes in each matching value and the longest decides. Only on an exact tie does the standard say the allow "SHOULD be used" — a SHOULD, not a MUST.
Take the pattern nearly every WordPress site serves:
For /wp-admin/admin-ajax.php both rules match. The allow is 24 octets, the disallow 10, so the allow wins — in either order. A great deal of advice online says rule order matters. It does not.
Group selection is case-insensitive on the crawler name. With no group naming the crawler, RFC 9309 says it "MUST obey the group with a user-agent line with the * value, if present"; with no * group either, nothing applies.
Two refinements are not in the RFC's text. Google states that "if there are multiple groups in a robots.txt file that are relevant to a specific user agent, Google's crawlers internally merge the groups", so rules split across two User-agent: Googlebot blocks all apply. And a value is read up to the first character that is not a letter, hyphen or underscore, so User-agent: Googlebot/2.1 names Googlebot and User-agent: Foo Bar is the group for Foo. A strict grammar parser rejects both. This tester follows Google, whose parser the standard was written from, and says when your file depends on the difference.
Only two characters are special inside a path. The * wildcard matches "zero or more instances of any character", slashes included, and $ at the very end anchors the match to the end of the URL. A $ anywhere else is an ordinary dollar sign. There is no regular-expression support at all — matching already starts at the first octet of the path.
Sitemap: belongs to no group: it is an extension, takes an absolute URL, and applies wherever it appears.
# starts a comment anywhere, including mid-path: Disallow: /a#b is the rule /a, which blocks far more than you meant.
A bare Disallow: blocks nothing: the grammar allows an empty value, and an empty pattern matches nothing. It is the conventional way to say "no restrictions". A value not starting with / or * also matches nothing, because comparison begins at the first octet — Disallow: admin is silently dead.
A rule and a URL are escaped before comparison, and Google escapes them differently: in a rule it encodes only the non-ASCII octets, while a real request has everything encoded. So Disallow: /my page matches nothing, while Disallow: /café/ matches the real request. Nothing is decoded, so /a%2Fb is not /a/b. The tester prints the compared form when it differs from what you typed, and flags a rule that can never match.
Section 2.5 requires a crawler to parse "at least 500 kibibytes", and Google enforces exactly that. On a site that generates a rule per URL, the rules at the bottom may never be read, so the important ones belong at the top.
Most questions about robots.txt now are really about AI, and the honest answer has three parts. A green tick without them is not the truth.
Some agents document that they ignore robots.txt. Perplexity states that Perplexity-User generally ignores it, because those fetches are triggered by a person asking a question rather than by a crawl. A perfect Disallow for that token changes nothing, and the reality column says so.
Some tokens are not crawlers. Google-Extended and Applebot-Extended control whether already-collected content may train models. Nothing fetches under either name, so a Disallow there is a use rule, not a crawl rule, and has no effect on search. Our AI crawler blocker builds the rules and explains which can be enforced by IP.
Some crawlers follow another crawler's group. Apple documents that if your instructions do not mention Applebot but do mention Googlebot, "the Apple robot will follow Googlebot instructions". So a Googlebot group written for search reasons quietly governs Apple too. The tester marks that answer as coming from the Googlebot group rather than from *, because you would never guess it from the file.
The vendors that document compliance are consistent:
| Token | What it is for | Follows robots.txt? | Crawl-delay? |
|---|---|---|---|
Googlebot | Google Search, Images, News and Discover | Yes | No — Google does not use it |
ClaudeBot | Anthropic model training | Yes, per Anthropic | Yes, and rules must be repeated per subdomain |
PerplexityBot | Surfacing and linking sites in Perplexity | Yes, per Perplexity | Not documented |
Perplexity-User | Fetching a page a person asked about | No — documented as generally ignoring it | Not applicable |
Google-Extended | Whether collected content may train Gemini | Not a crawler — nothing fetches under this name | Not applicable |
Applebot | Apple search features | Yes, and falls back to the Googlebot group | No — Apple states it does not follow it |
AdsBot-Google | Google Ads landing-page quality | Its own rules only — the * group is ignored | Not applicable |
Google-CloudVertexBot | Crawling for Vertex AI Agents | Yes, and a Googlebot rule reaches it | Not applicable |
AhrefsBot | Backlink and SEO crawling | Yes, per Ahrefs | Yes |
SemrushBot | Backlink and SEO crawling | Yes, per Semrush | Yes, up to ten seconds |
So the same line can be honoured, ignored or meaningless depending on who reads it, which is why this page answers per crawler. Where a vendor publishes nothing, it says "unverified" rather than guessing. For the practical version, the guide to blocking AI crawlers such as GPTBot and ClaudeBot walks through which methods hold.
robots.txt has a counterpart. Where it refuses, llms.txt invites: a curated list of the pages worth reading. Unrelated files, opposite purposes, and our llms.txt checker tests that one against its own format.
Rarely explained, and the most damaging thing in the file. RFC 9309 section 2.3.1.4:
If the robots.txt file is unreachable ... the crawler MUST assume complete disallow.
Unreachable means a 5xx, a 429, a timeout, a refused connection or a failed TLS handshake, and "complete disallow" means every URL on the host. A crawler that has seen this for more than thirty days may fall back to cached rules or treat the site as allowed — but a month of not being crawled is not a recovery plan.
A 4xx is safer. Section 2.3.1.3 treats an unavailable file as no rules, and a crawler "MAY access any resources on the server". A missing robots.txt is valid; a broken one is not.
Three more delivery failures the tester looks for:
Content type matters too: the standard calls robots.txt a text file "served as text/plain", and text/html is usually the fingerprint of the soft 404 above. Our headers analyzer shows what your server sends. On caching, section 2.4 asks crawlers not to reuse a cached copy for more than twenty-four hours, and permits longer when the file cannot be fetched — so a change takes about a day, and only affects future crawls.
Two things follow from the answers. A rule that needs enforcing rather than requesting belongs in the server configuration — the walkthrough for refusing an address in .htaccess is the shortest route on Apache. And robots.txt cannot tell you whether a visit claiming to be a named AI crawler was genuine; a reverse-DNS and IP check can, which is what the ClaudeBot verifier does for one of the busiest.
User-agent: * plus Disallow: / is occasionally intended. Far more often it is a staging configuration that got deployed. Reported as a warning rather than an error, because it is legal — but on a live site it is almost certainly wrong.
It is also not as total as it looks. Of AdsBot-Google, AdsBot-Google-Mobile, Mediapartners-Google and APIs-Google, Google says "the global user agent (*) is ignored", and of Google-Safety that it "ignores robots.txt rules". A site-wide block stops Search and leaves those four crawling; only a rule naming them reaches them.
A crawler that finds its own group stops reading *:
Googlebot is not kept out of /wp-admin/ here. It found its own group, and that group says nothing about /wp-admin/. An empty group is sharper still: name a crawler, give it no rules, and it is completely unrestricted.
The reverse trap is as common. Google lists a second token for several crawlers, and says "you need to match only one crawler token for a rule to apply". So a Googlebot rule also governs Googlebot-Image, Googlebot-Video, Googlebot-News, Google-InspectionTool and Google-CloudVertexBot — the last crawls for Vertex AI, which no Search rule hints at. Writing a separate Googlebot-Image group does not release it from the Googlebot one: the two are read together.
Noindex: to workNoindex: is not in RFC 9309 and Google does not support it here. The line does nothing at all. Deindexing belongs in a page-level tag or header instead, and the FAQ below covers why blocking a page makes that impossible.
A Disallow on /assets/ or /static/ stops a crawler rendering the page as a browser would, and Google renders before deciding what a page says.
Google's parser corrects a short list of frequent misspellings on purpose — Dissalow and Disalow among them — matches field names by prefix, and accepts a space for the colon when the line has exactly two words. A stricter parser drops those lines. So a typo can leave a rule working for Google and missing everywhere else, which is worse than one that fails everywhere. The tester reports each leniency it applied.
curl -sI https://example.com/robots.txt shows the status and content type by hand. It will not tell you which rule wins, nor that Disallow: /Private leaves /private wide open — paths are compared byte for byte.
Four honest limits:
It cannot make a crawler obey you. Everything here is what a compliant crawler concludes; a scraper that ignores robots.txt fetches the page regardless.
It cannot tell you what actually visited. The rules say what should happen; only your access logs say what did. A user-agent string is trivially forged, so a log line claiming to be Googlebot proves nothing on its own — that takes our bot IP verifier.
It cannot know undocumented behaviour. Where a vendor publishes how its crawler treats robots.txt, this page uses that and names the source; otherwise the answer says "unverified". One token is marked as in common use rather than vendor-verified, because Bing's crawler page is rendered by JavaScript and could not be read on the review date.
It cannot remove anything from search. Deindexing is a different mechanism, which blocking actively prevents.
And not to overclaim: this matcher has been compared against Google's own parser over more than two million questions with no unexplained disagreement, and against Scrapy's parser over the same set. It has not yet been run against a live website from the server it is published on.
The longer one. RFC 9309 requires the most specific match to win, and defines specificity as the number of octets in the rule. So rules are compared by length in bytes — not by position in the file, and not by allow beating disallow. Allow: /admin/public at 13 octets beats Disallow: /admin/ at 7 octets, whichever comes first. Only when two matching rules are exactly the same length does allow win, and the RFC puts that as a SHOULD rather than a MUST. This tester prints the winning line, its octet count, and the rule it beat.
No. A robots.txt file applies to one scheme, host and port. blog.example.com needs its own file at https://blog.example.com/robots.txt, and so does the http:// version of a host if it answers separately rather than redirecting. Anthropic says this outright in its crawler documentation: rules must be repeated for every subdomain. A rule on the main domain does nothing for a subdomain, which is how staging subdomains end up indexed.
The paths are, the field names and crawler names are not. Disallow: /Private does not block /private. RFC 9309 requires case-insensitive matching only when a crawler is looking for the group that names it, so the spelling of a product token does not matter. The path after the colon is compared byte for byte instead. On a server that treats URLs case-insensitively this is a real gap, because the page is reachable at both spellings and only one is closed.
Because robots.txt controls crawling, not indexing. A blocked URL can still be listed if other pages link to it - Google keeps the address and the anchor text without fetching the page. Worse, blocking it removes the only way for a crawler to see a noindex tag on that page, so blocking a page you want out of search actively prevents its removal. Let it be crawled and serve noindex in a meta tag or an X-Robots-Tag header instead.
Plan on a day. The 24-hour cache ceiling in RFC 9309 is the number to work from, and a crawler is allowed to hold a copy longer than that while your file cannot be fetched at all. Google publishes the same figure and lets you ask for a re-fetch in Search Console. Remember too that a change only affects future crawls - it does not undo what has already been fetched, and it does not remove anything already indexed.
Your whole site stops being crawled. RFC 9309 section 2.3.1.4 is explicit: when robots.txt is unreachable, which includes any 5xx and a 429, "the crawler MUST assume complete disallow". Not the paths you named - everything. A 404 is far safer, because an unavailable file means there are no rules and section 2.3.1.3 then lets a crawler "access any resources on the server". When the file comes out of your application rather than off disk, any error in that application closes the whole site to search until it is fixed. Serving it as a static file, or from the edge, removes that failure mode entirely.
robots.txt asks. These tools help you check whether anyone listened, and refuse the ones that did not.
AI Crawler Blocker
Build the rules for AI bots
llms.txt Checker
Test the file that invites them
GPTBot Verifier
Confirm a GPTBot visit is real
ClaudeBot Verifier
Confirm a ClaudeBot visit is real
User Agent Checker
Identify a bot from its string
Access Log Analyzer
See which bots really visited
Bot IP Verifier
Verify any crawler by its IP
All Tools
The complete free toolkit
This page tells you what your rules say. Only your server log says what came and what it asked for — and because a user-agent string can be typed by anyone, the visits that matter still have to be verified by IP.
Last updated 26 September 2026 · Matching rules from RFC 9309 and Google’s robots.txt documentation. Crawler tokens and behaviour come from each vendor’s own pages, named beside every answer.