[ Definition ]
llms.txt: what it does, what it does not, and why I publish one anyway
[ In short ]
llms.txt is a markdown file indexing a site's canonical pages. No major vendor uses it as a visibility signal, and a study across three hundred thousand domains finds no citation lift attributable to the file. It stays useful as an entity-consistency artifact, at almost no cost. That is why I publish one.
- Published
- Reading
- 8 min
What is llms.txt?
A markdown text file, published at the site root, listing canonical pages and describing each in one line. The proposal dates to 2024 and the starting idea is reasonable: give a system that reads text a curated index instead of leaving it to infer structure from menus, breadcrumbs, and an XML sitemap.
The name created half the problem. It resembles `robots.txt`, a de facto standard respected for twenty years, and that resemblance implies an authority that does not exist. `robots.txt` works because crawlers decided to read it. `llms.txt` is a proposal, and a proposal is worth as much as its adoption on the consuming side, not the publishing side.
Who uses it and who does not?
It is worth separating documented positions from hearsay here, because on this topic the two get mixed constantly. The right-hand column states the level of certainty, and it matters more than the middle one.
| Vendor | Position on consuming the file | Level of certainty |
|---|---|---|
| Stated as unsupported and not planned, no effect on Search or AI Overviews | Public statements from spokespeople | |
| OpenAI | No stated support, crawlers are governed via robots.txt | Official documentation, which never mentions it |
| Anthropic | No public commitment to consuming it as a signal | Absence of documentation, not a denial |
| Perplexity | Reported retrieval of the file to prioritise pages | Secondary sources, not verified first-hand |
Where does the confusion start?
From a true fact read the wrong way. Anthropic and OpenAI publish an `llms.txt` for their own technical documentation. This gets cited constantly as proof that both support it, and the conclusion does not hold: publishing a file is the opposite of consuming one.
These are two different uses sharing a name. The first is a documentation delivery format: a vendor puts its docs in clean markdown because it knows they will be pasted into a model, and it serves developers. It works, and it is measurable, because the benefit is immediate for the reader. The second is a retrieval signal: the idea that an engine reads your `llms.txt` and weighs it when deciding whom to cite. Nobody has stated they do that.
The distinction is not pedantry, it changes the budget decision. As a delivery format, `llms.txt` makes sense if you have technical documentation and users who paste it into a model. As a visibility lever, nothing supports it, and treating it as one means moving attention away from work with demonstrable effect.
What does the data say?
The widest study available analyses roughly three hundred thousand indexed domains. Adoption went from under one per cent in early 2025 to around ten per cent by mid 2026: fast growth, which explains why the question comes up so often.
The interesting result is not the adoption, it is the absence of effect. Controlling for site authority, structured data density, and content recency, no citation lift attributable to the file's presence emerges. In the model used for the study, removing the `llms.txt` variable improved prediction accuracy, which is the statistical way of saying that variable was adding noise rather than information.
There is a second figure that matters more than the first, and it concerns who adopted. Adoption is high in the middle band of the web and close to zero among the highest-traffic domains. That asymmetry is decisive: an emerging standard becomes mandatory for engines when the sources they have to read anyway adopt it. If the top of the web does not move, no vendor is forced to treat the file as authoritative, and growth from below does not change the balance.
So why do I publish one?
Because it costs an hour and the benefit I am after is not retrieval. This site has an `llms.txt`, and the reason it exists is stated here rather than left implied.
The first reason is that writing it forces you to produce a canonical description. To fill that file you have to decide, in one line, what you are and what you are not. It is the same exercise as entity clarity, done in a form that admits no vagueness, and most companies discover while writing it that the line did not exist.
The second is that it is the only natural place to state the negations. A site says what you do; an `llms.txt` can also say what you are not, and that reduces category confusion more than any about page. If you keep getting described as an SEO agency or a software platform, that line is where to put it in writing.
The third is consistency. The file becomes one more artifact that has to say the same thing as the site, the structured data, and the public profiles. That makes it useful indirectly: adding a place to keep aligned forces you to keep the others aligned too.
How do you write one that is worth having?
By treating it as what it is: a readable identity document, not an optimisation file. The quality test is whether someone who does not know you would describe you correctly after reading only that.
- A canonical description in two or three lines, identical to the one used elsewhere.
- The explicit negations: what you are not, which categories do not describe you.
- The addresses of canonical pages, each with a line of context, not a bare list.
- The vocabulary you use, if you work in a field with ambiguous terms.
- Contact, location, and languages, in the same formats as the site.
- The date of last update, and the commitment to honour it.
When is it a waste of time?
When it comes before the basics. If your pages are not legible without JavaScript, or your name is written three different ways across site and profiles, publishing a curated index solves a problem you do not have instead of one you do. The order matters more than the individual choice.
When you expect a ranking effect. There is none, and expecting one produces the wrong conclusion downstream: if citations do not rise you will blame the file instead of looking at where the problem actually is.
And above all when you cannot keep it current. Here the file moves from neutral to harmful: an `llms.txt` describing services you no longer offer or pointing at pages that no longer exist is an inconsistency you declared yourself, about your own identity. An absent file says nothing. A stale file says something false. If you have no moment where you review it, better not to have one.
[ What to take away ]
- Publishing an llms.txt is not the same as consuming it as a signal: the confusion comes from vendors publishing one for their own docs.
- No major vendor has publicly committed to using it as a signal in answer surfaces.
- Across three hundred thousand domains, no citation lift attributable to the file emerges once authority, schema, and recency are controlled.
- It makes sense as an entity-consistency artifact: it forces a canonical description and a statement of what you are not.
- A stale file is worse than no file, because it becomes an inconsistency you declared yourself.
How visibility inside a generative engine gets measured, written out in full. Read the article
[ Related service ]
[ Sources ]
Want the same reading on your case?
The automated preview gives a first signal in seconds. I prepare the useful reading myself, and it arrives within 48 hours.
Send me your case