
Should Your Website Have An llms.txt?
Should your site have an llms.txt? Probably not, and the reason is more useful than the answer.
Systems Ninjas (2026). Should a business publish an llms.txt? Rests on our own measurement of 400 services sampled from a national register, resolving to 109 distinct websites, measured 2026-09-02, resolver precision 65% and published as such. systemsninjas.com/post/should-your-site-have-llms-txt
Take any of it. You do not need to ask. If you find an error in here we would rather hear about it than not, and the correction goes on the page.
Probably not. The reason is more useful than the answer.
We measured 400 businesses taken from a national register of regulated services. Of the sites that carried the file, at least six of the fourteen genuine ones were not written by the business at all. And every single block we found on AI crawlers had been switched on by a supplier, on the owner's behalf.
Two words you will need. An AI crawler is a program that reads websites for an AI company. A CDN is the company that stores a copy of your website around the world and serves it to visitors quickly. Most sites use one. Most owners never think about it.
Nobody in that industry has chosen to keep AI systems out. Nobody has chosen to let them in either. Their suppliers decided for them, and nobody told them.
01 / What it isWhat an llms.txt is supposed to do
It is a plain text file that sits at /llms.txt on your website. Its job is to tell AI systems what your site is about and which pages matter.
Picture a note taped inside your shop door listing what you sell and where to find it. Useful, if anyone reads it.
The idea is a proposal, and that word is doing real work. Somebody suggested it. Nobody has agreed to be bound by it. It is about two years old, and a lot of people who sell services around it have pushed it hard.
The honest status: plenty of websites publish one, and nobody has shown that AI systems read it.
Two different questions get mixed up here, so pull them apart.
Does anything read the file? We have not found evidence that any AI search system people actually use reads it when deciding what to quote.
Does anyone have to do what it says? A proposal binds nobody. Even a system that did read the file would be under no obligation to follow it.
So you would be writing a note that may never be read, and that nobody has to act on if it is. That is a thin thing to build a plan on.
02 / Our withdrawalWe are withdrawing a number we used to repeat
This is the part we would rather not write.
Our own notes used to carry a figure everybody quotes: that 97% of llms.txt files are never requested by anything. Writing this page, we went looking for the original measurement it came from. We could not find one. So we dropped it. It is out of our material and it will not appear in anything we send a client.
That does not make the claim false. It means we cannot stand behind it.
A number you cannot stand behind is not an argument. It is a risk, because anyone can ask where it came from and you have no answer. We would rather say out loud that we dropped it than quietly stop using it and hope nobody noticed.
03 / What we measuredWhat we measured instead
We picked a group we could count and list in full: 400 childcare services drawn from the Australian national register, in 2026-09.
We name the industry on purpose. The field note behind this page names it, and a study you cannot identify is a study you cannot check.
We wrote down exactly what we would measure, and froze it BEFORE we fetched the first page. So nothing could be adjusted afterwards to make the result look tidier.
The first number to ask about is how many businesses we could actually reach, and it is the one most studies hide. It is the number everything else gets divided by. The register holds 98 columns of data and not one website address.
So we had to work out every web address ourselves. That step is where a study like this quietly loses people:
| Services sampled | 400 |
| Rows we found a live matching site for | 138 (34.5%) |
| Distinct websites those collapse to | 109 |
| Our resolver's precision, hand-checked | 26 of 40 = 65% |
Read that as a fact about OUR METHOD, not about those businesses. "Unresolved" means we found nothing. It does not mean the business has no website.
We were matching business names to web addresses the way you would look up a stranger in a phone book. Some of the entries you land on are a different person with a similar name.
04 / The findingWhat we found
21 of the 109 sites served a file at /llms.txt. That is 19.3%. It is also wrong.
It was three times higher than we expected. A result that flatters you is the exact moment to distrust it, not to publish it.
Only 21 sites were involved, which is few enough to check every one. So we opened all 21 by hand.
Six of them were not the business we were looking for at all. Four were online shops, one of them selling records. One was a speech pathology clinic. One was a farm business. One more we could not identify.
The figure that survives checking: 14 of 109. That is 12.8%.
That correction pointed down, and not by accident. This is the most useful thing in the whole study, because it applies to any study you read.
Two big website builders create an llms.txt on their own, without anybody asking. So whenever our method landed on the wrong site, and that wrong site was an online shop, it was MORE likely to carry the file than a real childcare business was.
Our mistakes pushed the number up, never down. Publishing the raw 19.3% would have meant publishing a fault in our own tool and calling it a finding about an industry.
A study built the same way, but checking only a sample of those sites instead of every one, would have published 19.3% with a straight face.
05 / Who wrote themWho actually wrote those files
Of the 14 genuine ones:
- 3 came ready-made with the website builder. Each one names that builder inside the file.
- 3 are the same file with the business name swapped in. Three different brands, all serving one identical marketing sentence. Somebody's agency template, copied three times.
- 1 was produced by an SEO plugin, an add-on that helps a site show up in search, and it says so in the file.
- 1 was clearly written by a person: real sentences, real detail about that actual business, a history that checks out.
- The rest give no clue who wrote them.
So at least 6 of the 14 were not written by the business whose site they sit on. Across a sample of 400 businesses from a national register, we found exactly ONE file that somebody at the business appears to have written themselves.
How good were the files? 16 of 21 started with a proper heading. We followed 63 of the links inside them and 5 were already dead.
Put up once, never looked at again. That tells you what these files really are: a box that got ticked, then forgotten.
06 / The other halfThe other half of the story, and it matters more
While we were there, we read every robots.txt as well. That is the file that tells crawlers which parts of a site they may visit. It is the notice at the entrance rather than the note on the door.
89 of 109 sites had one. Three of those block at least one AI crawler.
Then the part that reframes the whole topic:
All three blocks came off the same CDN's ready-made list. Each one carried that company's own marker, put there automatically. All three blocked the identical eight programs.
Blocks written by a human at the business: zero. Not one business in the sample made this decision itself.
Two details that matter if you are thinking about this seriously:
- None of them blocked the crawler that decides whether you appear in one big AI assistant's answers. They blocked a different crawler from the same company: the one that only gathers material to train the model.
Same company, two crawlers, two entirely different jobs. It is like bolting the back gate and leaving the shop front wide open. A business that believes it has opted out of AI search has most likely opted out of something else.
- The regulator that publishes the register, a government body, does exactly the same thing on its own website. It blocks an AI crawler, inside the same CDN-generated section, and nobody there appears to have chosen it either.
One site was running both files at once, and they flatly disagreed. Its llms.txt invites AI systems in to read the site. Its robots.txt blocks eight AI crawlers at the door.
The owner did not write either side of that argument. One supplier put up the invitation. Another put up the block. Nobody at the business was standing in the middle to notice.
07 / What to doSo what should you actually do?
The answer is not "add an llms.txt". That is the one recommendation our measurement does not support.
The useful version:
Whatever your website currently tells AI systems, you probably did not write it. You can find out in about thirty seconds.
Open two addresses on your own website: yourdomain.com/robots.txt and yourdomain.com/llms.txt. Type them into your browser exactly like that, with your own domain name in front.
If you find a block on AI crawlers that you never chose, your hosting provider made a business decision for you. That is a decision worth making on purpose rather than inheriting.
If you find an llms.txt describing your business in words nobody at your business wrote, the same applies.
If both come back empty, you are with the large majority. Of the 109 sites, 14 carried a genuine llms.txt and only 3 carried any AI-crawler rule at all. Having neither costs you nothing that we can show.
08 / Where it is weakWhere this measurement is weak
We put this as prominently as the findings. Leave this part out and you are not doing research. You are doing marketing in a lab coat.
- Matching businesses to web addresses is the biggest weakness. We got that match right 65% of the time, so about a third of the sites we pulled in were not childcare services at all. We only caught what that was doing to the result because we checked every single file by hand instead of a sample.
- One country, one industry. Childcare in Australia is not "small businesses everywhere", and we will not write it up as though it were.
- We fetched each page once, plainly, from one place in the world. If a file is written onto the page afterwards by code running in the browser, we would not see it. Nor would we see a robots.txt that hands out different rules depending on where the visitor is. We think both are rare. We have not measured that.
- The numbers are smallest exactly where they matter most. Three sites blocking, fourteen genuine files. With counts that small, the true figure could sit some way either side of ours, and that range is printed in the full study rather than rounded away.
And one we like even less. Our first run broke after 325 of 400 sites and saved nothing. Starting again meant fetching sites we had already visited, so we used more of those sites' bandwidth than our published limit implies. Bandwidth is what it costs a site to send its pages out.
We have fixed it. We are telling you here rather than leaving it out.
Measured 2026-09-02. One location, one plain fetch of each page. Our crawler said who it was, obeyed every robots.txt absolutely, and never tried a second time on a page that turned us away with a 403. Full method, raw counts and ranges are in the field note.
Check your own two files before you take anyone's advice, including ours
It takes thirty seconds and needs nobody's permission. If you find something you did not write, send it to us and we will tell you which vendor put it there and what it actually does. If you find nothing at all, that is a cleaner position than most of the sites we measured, and we will tell you that too.
Send us what you found [email protected] · No client of ours is named anywhere on this page, and if you become one, you will not be either.