A long panel of identical switches in a dim, empty corridor, most of them still in the same position.

109 Childcare Websites. 3 Block AI Crawlers. All 3 Blocks Came From Cloudflare.

September 02, 2026

109 childcare websites. 3 block AI crawlers. All 3 blocks came from Cloudflare.

I sampled 400 services from a national register, resolved 109 distinct websites, and read what each one tells AI systems. Three of them block an AI crawler. Not one of the three wrote the rule. This is the method, the numbers, and the correction that nearly cost me the headline.

A long panel of identical switches in a dim, empty corridor, most of them still in the same position.
There is no villain in this one either. The switches were set at the factory, and the corridor has been empty ever since.
Cite this as

Systems Ninjas (2026). llms.txt and AI-crawler posture in Australian centre-based childcare. Measured 2026-09-02. Population: 400 services sampled from a national register of 17,687 eligible centre-based services, resolving to 109 distinct websites. systemsninjas.com/post/australian-childcare-ai-crawler-audit

Every chart on this page carries its own population and this URL burned into the image, so it stays citable after somebody screenshots it. Take them. You do not need to ask.

01 / Why this sectorI needed a population I could actually define

Last week I pointed a scanner at my own website to find out what it tells AI systems. It told me a few things I did not know I had said, and I published the whole list.

That is a small embarrassing story on its own. The interesting question is how common it is.

To answer that I needed a sector where I could define the population properly, instead of picking sites that felt representative and calling it a sample. A real register. Complete, public, and maintained by somebody with a legal reason to keep it accurate.

Australian childcare has one. Every centre-based service in the country sits in a national register kept by the regulator... 18,164 rows on the day I took the snapshot, of which 17,687 were eligible for this study.

And childcare turns out to be close to a perfect test case for the question I actually care about. Which is not is this sector AI ready. It is:

Has anyone here made a decision at all?

These are small local businesses. Nobody running one has sat down of an evening to consider whether an AI company should be allowed to read their fees page. So if a decision exists on their website, somebody else made it. That is the thing worth measuring.

Everything below was pre-registered. I wrote the measurement spec and froze it before the first fetch, so I could not go hunting for whichever number turned out to be interesting.

02 / The denominatorThe first number is the one that limits every other number

The register carries 98 columns. Name, address, suburb, phone, quality ratings, opening hours for every day of the week.

It does not carry a website address. Not one. Neither do the regulator's own service search pages.

So every URL in this study was constructed by us. And that is exactly where a study like this leaks.

Funnel chart: 400 services sampled, 138 matched to a live site (34.5%), 109 distinct websites, and hand-checked resolver precision of 65%.

400 services sampled. 138 matched to a live site, which is 34.5%. Those 138 collapse to 109 distinct websites, because 16 domains serve several of the sampled centres and one domain serves seven of them. Every rate in this article is out of 109.

Then I hand-checked 40 of the matches to see how often the resolver was right.

26 of 40. 65% precision.

That is not a footnote, so it is not going at the bottom. About a third of the resolved set is not the population at all... 8 completely unrelated businesses, 5 parent organisations such as a parish rather than the centre's own site, and 1 I could not determine either way.

I am telling you this before the findings rather than after, because it decides how much weight the findings can carry. And because "unresolved" means MY method found nothing. It does not mean the centre has no website. Recall cannot be estimated this way, so I am not going to report it and neither should anybody quoting this.

03 / llms.txt21 files. Nothing like 21 decisions.

Quick definition, because this one is new enough that most people have not met it. llms.txt is a proposed convention: a plain text file at the root of your site telling AI systems what your business is and which pages matter. Think of robots.txt, but for meaning instead of permission.

21 of the 109 sites served one. That is 19.3%.

Which was about three times higher than I expected. And that is precisely the moment a number deserves to be distrusted rather than published.

So rather than sampling the numerator, I opened all 21 and read them.

Six were not the childcare business at all. Four were online stores, one was a speech-pathology practice, one was a farm. One more I could not identify.

Verified: 14 of 109. 12.8%.

Chart: 21 llms.txt files found, 7 removed as not the childcare business, 14 verified, broken down by who wrote them. Only one was hand-authored.

That correction points the direction it does for a structural reason, and this is the most useful paragraph in the article.

Shopify and Wix generate an llms.txt automatically. So when my resolver picks the wrong domain and lands on an online store, that wrong domain is MORE likely to carry the file than a real childcare site is.

My errors do not cancel out. They push this specific number up.

The part that should worry anyone who publishes numbers

If I had sampled the numerator instead of reading every file in it, I would have published 19.3% with a completely straight face. The contamination would have been invisible in the output, invisible to a reader, and invisible to me.

A study of this shape does not fail loudly. It hands you a clean-looking percentage that happens to be wrong in the direction that makes it interesting.

Now the part that actually matters. Of the 14 real ones:

  1. Three are Wix platform defaults. Each one carries the line "This site is powered by Wix and supports the Model Context Protocol".
  2. Three are the same file with the brand name swapped. Three separately branded childcare groups serve an llms.txt containing an identical sentence about sparking a love of learning. I fetched all three and diffed them.
  3. Two more sit in the same size and link-count class but I did not diff them, so they are counted as suspected and never as confirmed.
  4. One was generated by an SEO plugin, which announces itself inside the file: "Generated by All in One SEO".
  5. Four are unattributed. No generator string, no template match. Recorded as unknown rather than guessed.
  6. One is plainly hand-authored. Bespoke prose, real detail about the centre, correct history. Somebody sat down and wrote it.

So at least 6 of the 14 were demonstrably not written by the business whose site they sit on. And in a 400-service sample of a regulated national register, exactly one person appears to have written this file themselves.

Quality, where a file exists: 16 of the 21 open with a valid Markdown heading. Of 63 links I sampled from inside them, 5 were already dead. Published once, never opened again.

One thing I am deliberately not telling you

Whether any of these files ever get read by an AI system. There is a figure doing the rounds that claims they almost never are. I went looking for where it came from and could not find a primary source, so it is not in this article and it is not hiding in a footnote either.

What survives is the defensible half: these files are being published, and consumption by retrieval systems is unproven. If you have seen that other number quoted at you, ask the person quoting it who measured it.

04 / robots.txtEvery AI block I found was a vendor default

89 of the 109 sites served a readable robots.txt. 20 had none at all.

Three of the 89 block at least one AI crawler. That is 3.4%, and the interval around it is wide: 1.2% to 9.4%. Three is a small number and I am not going to dress it up.

Here is the part I did not expect.

Chart: of 89 readable robots.txt files, 3 block an AI crawler and all three carry Cloudflare's managed content marker, 14 have a permissive rule, and 72 say nothing about AI crawlers at all.

All three of those files carry Cloudflare's own # BEGIN Cloudflare Managed content marker. All three block the identical eight agents, in the identical order.

Hand-written AI-crawler rules in the sample: zero.

Not one site in 89 made this decision itself. And 72 of the 89 contain no rule matching any AI crawler at all, which is not a policy either way. It is silence.

Two more zeros worth having.

Zero sites block OAI-SearchBot. That is the agent governing whether you show up in ChatGPT's search answers, and it is a completely different agent from GPTBot, which is training only. The managed list blocks the training one. So if a centre believed it had opted out of AI, it has opted out of the half that does not affect whether a parent can find it, and left the half that does. Nobody chose that trade. It arrived pre-made.

Zero blanket Disallow: /. Nobody has closed the door.

And the regulator does exactly the same thing. I read its robots.txt myself before writing this sentence: it disallows nine named agents, eight of them AI crawlers, inside the same Cloudflare managed block, and it adds Cloudflare's Content-Signal: ai-train=no policy line on top.

Which means the most sophisticated AI policy anywhere in this study, including the regulator's, was written by a CDN.

05 / The contradictionOne site does both at once

My spec said that under 20 cases I publish observations and no percentage. The set here is one, so here it is as an observation.

One site publishes an llms.txt inviting AI systems to come and read it, while its robots.txt blocks eight AI crawlers from doing that.

It is the cleanest possible illustration of the whole finding, because the owner did not author the contradiction. They published the llms.txt. Cloudflare's managed default published the block.

Two vendors. Opposite instructions. One website. Nobody in the middle.

06 / So whatThe recommendation is not the one you are expecting

Both halves of "AI readiness" in this sector are vendor artefacts rather than business decisions. The llms.txt files are mostly Wix defaults, one agency template replicated across three brands, and an SEO plugin. The crawler blocks are entirely Cloudflare's managed list.

One clearly hand-authored file. Zero hand-written crawler rules.

The sector has not opted out of AI crawling. It has also not opted in. A CDN made the call for the few who have one, and everybody else has no position, because nobody ever asked them for one.

Now, I want to head off the conclusion people will want to draw from this, because it is the wrong one. The useful version is not "go and add an llms.txt". Adding one puts you in exactly the same category as the three Wix defaults... a file on your server that you did not write and will never look at again.

The useful version is smaller than that, and it is free.

Thirty seconds, on your own site, right now

Type your own domain into a browser and put /robots.txt on the end. Then do it again with /llms.txt.

If you see something you did not write, you have just found out who is setting your policy. It is probably your CDN, your website platform, or a plugin somebody installed years ago.

If you see nothing at all, that is also an answer, and it is a more honest one than a file a plugin generated on your behalf.

That is the entire recommendation. It costs nothing, I am not selling it, and it works whether or not you ever speak to anybody like me.

If you run one of these centres, I will send you your own numbers

Tell me the domain and I will run the same checks I ran for this article, then send you the raw output, the method, and what I would look at first. Free, with nothing attached. And if it turns out your llms.txt really was written by somebody on your team, I will tell you that too, and you will be the second one I have found.

Ask for your numbers [email protected] · No institution or business is named anywhere in this article, and yours will not be either.

P.S.Where this method is weak

Stated at the same size as the findings, because the alternative is marketing wearing a lab coat.

  • URL resolution is the dominant weakness, and it is worse than my own spec predicted. The spec anticipated a favourable bias, that name-matching domains belong to more web-invested businesses. The real problem was cruder. At 65% precision, about a third of the resolved set is not the population.
  • The contamination is not random with respect to the outcome, and it biases toward the interesting answer. I caught it only because I verified the numerator exhaustively rather than sampling it.
  • The unit of analysis was wrong in my first spec. 138 resolved service rows are only 109 websites. Corrected, and both denominators are published.
  • One market, one sector. This is Australian centre-based childcare. It is not "small businesses", and it must never be written up as one.
  • Static fetch only. A robots.txt that varies by CDN edge logic, or an llms.txt injected by JavaScript, is invisible to me. I think both are rare. I have not measured that.
  • One vantage point. Everything was fetched from Malaysia. Geo-varied responses would show up as artefacts and I would not know.
  • n is small exactly where it matters most. Three blockers and fourteen verified files. The confidence intervals are printed beside every proportion for that reason, and the 3.4% one runs from 1.2 to 9.4.
  • Generator attribution is an inference, not an observation, although each one rests on a string the file publishes about itself.
  • I wasted these sites' bandwidth. My first run died at 325 of 400 and wrote nothing, because results were only saved at the very end. The restart then re-fetched every host it had already crawled. On a deliberately slow crawl, saving only at the end turns one interruption into a second visit for data I already had and threw away. Fixed. It is here because my real footprint was larger than my published rate limit implies.

And the honest summary of the last one: the failure was mine, the cost was theirs. That is the correct order to say it in.

Crawl posture. One declared user agent naming Systems Ninjas, linking a method page and carrying a contact address. robots.txt was fetched first for every host and honoured absolutely, including before /llms.txt and not only before /. Rate limit as written in the code: never two requests to one host inside 3 seconds, at most 4 hosts in flight, at most 3 requests per site, and a host's own Crawl-delay overrides mine whenever it is larger. No authentication, no cookies carried between hosts, no JavaScript, no forms. A 403 is an answer and was never retried with a different user agent.

Personal data. Stripped at ingestion rather than at publication. The register's phone, fax and provider-legal-name columns are dropped as each row is read, so they never reach a record that could be written out. Email addresses seen on pages are reduced to a boolean and a domain. Nothing in this study needed anybody's name.

Nobody is named. No business, no centre and no brand appears anywhere in this article, and that is a design rule rather than a courtesy. My resolver is 65% precise, so naming a business would mean stating as fact something my own instrument gets wrong about one time in three. Aggregate reporting is the only mode in which that error rate is survivable. The regulator is named because you need it to reproduce the study, and because everything said about it here is one public URL away from being checked.

Reproducing it. The sample was drawn from the national register export with a fixed, published seed, snapshotted on 2026-09-02 and hashed. Websites change, and several of these may already be different. Re-measure before quoting any of it.

Uldis Zalcmanis
Written by

Uldis Zalcmanis aka Systems Rockstar

Founder, Systems Ninjas · Kuala Lumpur

I build the automation and AI systems businesses actually run on, and I am the person who gets called when one of them quietly stops working. Born in Riga, based in Malaysia, father of two, and constitutionally unable to stop reading page source.

The Lab is where the measurements get published, including the ones that do not flatter us.

Uldis Zalcmanis

Uldis Zalcmanis

Founder of Systems Ninjas

Back to Blog