Advertisement
Checked against v2 of the proposal

llms.txt Checker
and what the other validators get wrong

Enter a domain or paste your file. You get four separate answers: how your server delivered it, whether it matches the v2 format, whether it is actually useful to an agent, and what Google Lighthouse will say. Every finding names the rule book it came from, so a recommendation is never dressed up as an error.

Validate an llms.txt file

Enter a domain and the file is fetched from /llms.txt. Enter a path such as example.com/docs/ and it is looked for there instead, because a file covers the URLs under its own path.

Or paste a file to check it before you publish

Nothing is fetched and nothing is stored. The format, the Lighthouse prediction and the usefulness signals are all checked on this server and shown below. The delivery and discovery checks need a URL.

Quick Answer: how do I check whether my llms.txt file is valid?

Enter your domain in the llms.txt checker above and it fetches /llms.txt and reports four things separately: how your server served it, whether it matches the v2 format, how useful it is to an agent, and what Google Lighthouse will say. Only one thing is actually required — an H1 with your site's name. Everything else is a recommendation.

Advertisement
Jessica Wright, Cybersecurity Threat Researcher, on llms.txt checking and AI crawler verification at TrustMyIP.com
Written & Verified By

Jessica Wright

Cybersecurity Threat Researcher

Jessica works on bot and crawler verification at TrustMyIP, including the GPTBot, ClaudeBot and Googlebot verifiers and the access log analyzer.

This checker was built by reading the v2 proposal at llmstxt.org and the source of Google Lighthouse's own llms.txt audit, then writing the parser to the ordered section list the proposal prints. It was verified with 318 offline assertions and 103 assertions against a local server covering redirect chains, wrong content types, HTML served in the file's place, a real timeout and a self-signed certificate. Three cases where Lighthouse disagrees with the proposal — an underlined H1, a byte-order mark, and a file with no links — are reported rather than hidden. It has not yet been run against a public site from this codebase, and the page says so under "What this checker cannot tell you".

Last reviewed 25 September 2026 · Published 25 September 2026

View all articles by Jessica Wright
Advertisement

What Is an llms.txt File, and What Does This Checker Validate?

An llms.txt file is a short markdown file that hands an AI agent a curated map of your site: a title, a one-paragraph summary, and lists of links to the pages worth reading. This llms.txt checker validates it against version 2 of the proposal at llmstxt.org, by Jeremy Howard of Answer.AI, and separates what the format requires from what other tools merely prefer.

Two things surprise most people. It is a community proposal, not a web standard: no standards body has adopted it. And it does not have to sit at the root — the proposal covers files "at the root path /llms.txt of a website or at any subpath (e.g. /docs/llms.txt)", where "a file covers the URLs under its path, and where more than one file applies, agents should use the most specific one". A site can legitimately have nothing at the root and a complete file under /docs/.

What it is for is narrower than the hype suggests. The proposal expects agents to "view or search llms.txt to find the information they need, then follow the relevant links", with the detail behind those links so the file stays small enough to sit in a context window. It is a reading list, not a permission file.

Advertisement

llms.txt is constantly confused with robots.txt, and they pull in opposite directions: robots.txt is where you refuse a crawler, llms.txt is where you invite one. If your goal is to keep AI crawlers out, the file you want is robots.txt. Our AI crawler blocker builds the rules, and the guide to blocking an IP address at every layer covers the case where a polite refusal is not enough. It is not sitemap.xml either: a sitemap lists indexable pages for search engines, leaves out the markdown versions, and is far too large in aggregate to sit in a context window.

Key fact: the proposal lists exactly one required section, and it is the H1 carrying your project or site name. A file containing nothing but # Your Site is valid. Everything else in the format is optional.

How to Read Your llms.txt Validation Result

The checker gives four answers instead of one pass or fail, because "my llms.txt is broken" is really four different problems with different fixes.

Advertisement
VerdictThe question it answersWho decides
DeliveryDid your server hand over the file at all — right status, no detour, sensible content type?HTTP, not the proposal
v2 formatDoes the file match the ordered section list the proposal prints?The proposal
UsefulnessWould an agent get anything out of it?TrustMyIP, thresholds shown
LighthouseWill Google's own audit pass it?The Lighthouse audit source

Every finding carries the same labelling. A finding tagged llms.txt v2 spec quotes a rule from the proposal. Spec guidance is advice the proposal gives without making it part of the format. Google Lighthouse means the audit will object even though the proposal does not. TrustMyIP view is our opinion, and the threshold behind it is printed next to it so you can disagree.

Only errors against the proposal make a file invalid; our own observations never do. Fix those first, then decide which recommendations are worth your time.

Worth knowing: if the delivery verdict is red, ignore the format findings entirely until you have fixed it. When a server returns its home page in place of your file, the format checks are describing that home page.

What the llms.txt v2 Spec Actually Requires

The proposal defines the format as a sequence of sections "in the specific order" below. That is the whole of it: no formal grammar, no schema, no separate conformance document — one reason public validators diverge so widely.

  1. An optional byte-order mark.
  2. An H1 with the name of the project or site. This is the only required section.
  3. A blockquote with a short summary, "containing key information necessary for understanding the rest of the file".
  4. Zero or more markdown sections "of any type except headings" holding more detail.
  5. Zero or more sections delimited by H2 headers, each containing a "file list": a markdown list whose items each carry a required markdown hyperlink [name](url), then optionally a : and notes.

Read items 4 and 5 together: this is where most validators break. A bullet list that appears before the first H2 is a detail section, and its items are not required to contain links. The proposal's own reference file for FastHTML does exactly that: five bullets of guidance, none of them a link, above the first ## Docs heading. A checker that treats every list as a file list reports five errors on the proposal author's own file.

Here is the mock example the proposal itself prints:

# Title > Optional description goes here Optional details go here ## Section name - [Link title](https://link_url): Optional link details ## Optional - [Link title](https://link_url)

The Optional heading is not special syntax. The proposal says it "is used, by convention, for secondary information: links an agent can skip when a shorter context is needed". In v1 it had a mechanical meaning for a bundled tool called llms_txt2ctx; v2 removed both, so it is now just a section with a conventional name.

Markdown versions of your pages, and how agents find them

The links in the file "should therefore point to LLM-friendly content", which in practice means the markdown copy of each page. The proposal allows two naming forms: append .md to the whole URL (page.html.md) or replace the extension (page.md); for a URL with no file name, index.html.md or index.md.

v2 then added the piece v1 was missing: how an agent finds any of this without guessing. Two standard link relations: rel="alternate" type="text/markdown" points at a page's markdown version, and rel="describedby" points at the llms.txt file that covers it. Either may be delivered as an HTML <link> element or as an HTTP Link: response header. The proposal's own example:

Link: </docs/page.html.md>; rel="alternate"; type="text/markdown", </docs/llms.txt>; rel="describedby"

The checker looks for both on the page your file covers, in the head and in the response headers — the part of v2 that almost nothing else tests.

Limit: the proposal sets no maximum size anywhere. It says only that the file should stay "small enough to fit in context". Any checker quoting you a hard kilobyte limit is using its own number, not the specification's.

Google Lighthouse Checks llms.txt Too, and Its Rules Are Different

Lighthouse now includes an llms.txt audit among its agentic browsing checks, titled "llms.txt does not follow recommendations" when it fails. Its own description tells you the framing: the file "should be a Markdown file containing at least one H1 header". The audit's source is public, and reading it explains a lot of confused support threads.

It does three things: finds the H1 with the regular expression /^\s*#\s+.+/m, looks for a link with /\[.+\]\(.+\)/, and fails outright on a fetch error, a missing status, or a status of 500 or above. It also has a "File is suspiciously short." message, whose length threshold is not stated among the audit's constants.

Those two regular expressions are where Lighthouse and the proposal part company, in three ways worth knowing before you go looking for a fault in your file:

Your fileThe v2 proposalLighthouse
Title written underlined — a setext heading, Title then =====Valid — a real markdown H1Fails. Its regex matches only the # form
Starts with a byte-order markValid — the ordered list begins "An optional byte-order mark"Fails. Its regex allows only whitespace before the #, and a byte-order mark is not whitespace
An H1 and a summary, no links yetValid — the H1 is the only required sectionFails with "File does not appear to contain any links."

The byte-order mark case is the cruel one: editors on Windows add it silently, the proposal explicitly permits it, and the audit reports a missing H1. That is why the checker reports the Lighthouse column separately and names which of its three checks failed.

There is an odd tension inside Google itself: its Search documentation says you do not need extra machine-readable files to appear in AI features, while Lighthouse audits your site for one. Both are true, from different products.

Practical takeaway: if you want to satisfy both, write the title as # Your Site, save as UTF-8 without a byte-order mark, and include at least one link. That file is valid under v2 and passes the audit.

Does Anything Actually Read Your llms.txt?

Mostly not yet, and you should know that before you spend a morning on the file. Ahrefs studied 137,000 domains, found 38,360 with a valid llms.txt, and reported that 97% of them received zero requests for it in May 2026. The study was updated on 15 June 2026; its own conclusion was that llms.txt is "a solution in search of a problem".

Google's position on Search is on the record in its AI features documentation: "You don't need to create new machine readable files, AI text files, or markup to appear in these features." That page does not name llms.txt, so treat it as a statement about extra files in general.

The other side is real too, and the proposal states it plainly: thousands of sites publish a file, platforms including Mintlify, GitBook, Wix, Yoast SEO and AIOSEO generate one automatically, and OpenAI, Anthropic and Gemini all publish one for their own developer documentation. Chrome's Lighthouse audits for one. The file is not a dead letter; it is a supply that arrived well ahead of demand.

The honest position: publishing one is cheap and harmless, and there is no evidence yet that it wins you anything in search. Where it earns its keep is coding and documentation agents pointed at your docs — a curated reading list beats making them scrape your navigation.

Finding out for your own site

The answer is in your own access log, and it is a two-step job, because a request that claims to be GPTBot often is not. First see which bots hit which paths: our access log bot analyzer reads an Apache or Nginx log and groups the traffic by crawler, so a request for /llms.txt shows up if one ever happened. Then verify the ones that matter, because the user agent string is just text anyone can send — the GPTBot verifier confirms a visit really came from OpenAI's published address ranges rather than from something wearing its name.

If nothing reads your file while unverified crawlers take your content anyway, the question changes from guiding them to controlling them. Our guides on blocking AI crawlers such as GPTBot and ClaudeBot and on verifying Googlebot and AI crawlers before you block them cover both halves — verify first, block second, or you will block a customer.

Why Your llms.txt Is Not Working: the HTTP Layer Nobody Checks

Most llms.txt problems are not markdown problems. The file is fine and the server never hands it over, and because the usual validators only parse what arrives, they describe the symptom instead of the cause. Four failures account for nearly all of it.

1. Your server returns a web page instead of the file

This is the big one, and it belongs to every single-page application and every catch-all rewrite: an unknown path answers 200 OK with the site's index page. A structural checker parses that HTML, finds no # heading and reports "missing H1" or "invalid markdown" — so you go hunting for a fault in a file that was never served. The checker tests the body separately from the status. The fix: exclude /llms.txt from the rewrite, or upload it as a static file.

2. Redirects

A 301 from http to https, or from a no-slash path to a slashed one, costs a round trip. Some clients stop following after a few hops, and a chain that ends on your home page delivers something that is not your file at all. The checker follows up to five redirects by hand and prints every hop with its status.

3. Content type

The proposal does not require a content type, so this is never a specification error — but a file announced as text/html is announced as a web page, and some clients treat it as one. text/markdown or text/plain are the sensible choices. To see every header your server sends, including the Link: header carrying the v2 discovery relations, our HTTP headers analyzer shows the raw response.

4. Bot rules and certificates

A 403 from a firewall, or a certificate error, stops an agent exactly as it stops this checker. Fix a TLS problem first: no client that validates certificates will read the file.

Checking it by hand

Two commands settle most of it: the first shows the status, content type and any Link: header without downloading the body; the second prints the first lines of what actually arrives, so you can see whether you were handed markdown or HTML.

# status, content type and Link header only

curl -sSI https://example.com/llms.txt

# follow redirects and look at the first lines of the body

curl -sSL https://example.com/llms.txt | head -20

# does the covering page advertise the file?

curl -sSL https://example.com/ | grep -i 'describedby\|text/markdown'

Watch for: a 200 with a body that starts <!DOCTYPE html>. That is not a valid llms.txt no matter how good your markdown is, and it is the single most common reason a file that looks correct fails every checker you try.

What This llms.txt Checker Cannot Tell You

Four things, stated plainly. A checker that pretends to know more than it does is worse than none.

  • Whether anything has read your file. That lives in your access log. And a log line claiming to be GPTBot proves nothing on its own: the bot user agent checker tells you what a string claims to be, and the verifier tools tell you whether the claim holds up.
  • Whether an agent found it useful. The six usefulness signals are our own proxies with their thresholds printed — not a measurement of agent behaviour, which nobody publishes.
  • Whether your llms-full.txt is correct. There is no format to check it against: that file is a platform convention and appears nowhere in the proposal. Point the tool at one and it applies the llms.txt rules so you can see the difference. That is the most an honest checker can offer.
  • Lighthouse's exact "too short" threshold. The audit has the check; its length is not stated in the audit source we read. The short caution above uses our own 200-character line and is labelled as ours, not predicted.

Two notes on the parser. It treats lines indented four spaces or more as a code block, per CommonMark, so a heading behind that indentation is deliberately not read as one. And it does not resolve reference-style links ([name][ref]): the file list calls for an inline [name](url), so such an item is reported as having no link.

Finally, a fetched result is held on our server for a few minutes, so if you have just changed your file and the result looks stale, use the "Check again" link shown with it.

Frequently Asked Questions About llms.txt

What is an llms.txt file?

A plain markdown file that tells an AI agent which pages on your site are worth reading, and where to find clean versions of them. It was first written in 2024 by Jeremy Howard of Answer.AI and is now in its second version. It can sit at /llms.txt or under any path, such as /docs/llms.txt, in which case it covers only what is beneath that path. No standards body has adopted it, so treat it as a convention rather than a requirement.

Is an llms.txt file required for my website?

No. Nothing requires one, and Google states that you do not need to create new machine-readable files or AI text files to appear in its AI features. Publishing one is cheap and low-risk if your content is already structured, but Ahrefs studied 137,000 domains and found that 97% of valid llms.txt files received zero requests in May 2026. Treat it as an optional extra, not a ranking task.

How is llms.txt different from robots.txt?

They answer different questions. robots.txt is permission: it tells automated tools which paths they may fetch, and it is checked before crawling. llms.txt is an invitation: it tells an agent which pages are worth reading, and it is used on demand while the agent is helping someone. The proposal puts it plainly - llms.txt information is "used on demand, when an agent needs information about a topic".

Does Google use llms.txt?

Google has not said that it does. Its documentation on AI features tells site owners they do not need to create extra machine-readable or AI text files to appear in those features, and it never names llms.txt. Meanwhile Google Lighthouse has added an llms.txt audit to its agentic browsing checks, so one Google product tests for the file while another says it is not needed for Search. Both are accurate; they come from different teams.

Do I need an llms-full.txt file as well?

No. llms-full.txt does not appear anywhere in the proposal - it is a convention some documentation platforms adopted for a single large file containing the content itself rather than links. If you publish one, nothing validates it, because there is no format to validate it against. This checker will read it if you point it at one, and applies the llms.txt format rules so you can see how far it diverges.

Why does another llms.txt validator say my file is invalid?

Usually because it is enforcing its own preferences as rules. The most common one: a missing summary blockquote reported as an error, when the proposal says the H1 "is the only required section". Two others are a byte-order mark, which the proposal explicitly allows, and a bullet list above the first H2, which is a detail section and needs no links. Every finding on this page names the authority behind it so you can tell the difference.

Related AI crawler and bot tools

Guiding agents with llms.txt is one half of the job. Verifying which ones actually turned up, and refusing the rest, is the other.

Now find out whether anything actually fetched it

A valid file that no crawler ever requests is worth knowing about. Upload your access log to see which bots came and what they asked for, then confirm the ones claiming to be AI crawlers really are.

Last updated 25 September 2026 · Format rules from the v2 proposal at llmstxt.org, audit behaviour from the Lighthouse llms.txt audit, adoption figures from Ahrefs' 137,000-domain study, all read 25 September 2026