
We Audited Our Own Chatbot Detector. On Our Own Headline's Question, It Was Right Once In Eight.
We audited our own chatbot detector. On the question our own headline asked, it was right once in eight.
Systems Ninjas (2026). Measured error rate of our own chatbot-detection rule. 200 sites measured 2026-09-02, pre-registered and hashed before the first fetch: the 40 institutions from our published scan plus 160 .my domains drawn at random. Strata are never pooled and neither is a probability sample of Malaysian businesses. systemsninjas.com/post/we-audited-our-own-chatbot-detector
Take any of it. You do not need to ask. If you find an error in here we would rather hear about it than not, and the correction goes on the page.
Every number we have ever published about how many businesses have chat rests on one rule. The rule takes the code a website hands to your browser and searches it for the names of 31 chat companies. If any of those names is in there, the rule says that site has chat.
It is a bit like deciding whether a shop has a front door by reading the building plans instead of walking up to it. The plans are usually right. Usually is not the same as always.
That rule is what produced our scan of 40 Malaysian universities. Nobody had ever checked whether it works.
So we checked. We wrote the promise down first, before we fetched a single site: we would publish the result whatever it said.
It said this.
Think of the rule as a metal detector on a beach. Two separate things matter about it. When it beeps, how often is there really a coin down there. And of all the coins buried on that beach, how many does it beep at.
The first row answers both. When our rule says a site has chat, a visitor can really find something roughly six times in ten. And when a site really does have a chat control a visitor can see, our rule notices it roughly six times in ten.
Read the third row as this: for finding AI chatbots, the rule is not merely weak. It is noise. In the university group it pointed at eight sites, and exactly one of those eight had an assistant that tells you it is AI. In the second group of sites it pointed at sixteen, and not one of them did.
That third row is the one that hurts, because "40 Malaysian Universities. 3 AI Chatbots." was our own headline. We counted the thing we were worst at counting, and then we put the count in the title.
01A third of the sites could not be graded at all
On 67 of the 200 sites, 33.5% of them, we could not reach any verdict at all. That number is a finding in its own right. It is not a rounding error.
39 could not be reached. 9 gave us an error page. 4 sat behind a cookie notice or a check for robots. 3 were blank. 2 told us to stay out in robots.txt, the file where a site says which automated visitors may read it, so we never fetched them. 2 were login screens. 1 was an address with no website on it, up for sale.
And 7 had a button floating on the page that we could see but could not name. An unlabelled little face. A grid of dots. A bare speech bubble. The honest answer on those is that a visitor could not tell either, and our own ethics rule stops us clicking to find out.
A measuring tool that cannot answer for a third of what you point it at leaves a third-sized hole in every study built on it. In a normal write-up those sites quietly drop out of the sum. The report then says "we measured 134 sites", and a blind spot has been dressed up as a smaller study.
02Why it fails, and it fails both ways
Nearly all the wrong yeses come from one pattern. A wrong yes is the rule saying a site has chat when the site has none. The pattern is wa.me, the start of every WhatsApp click-to-chat link, and it fired on 27 sites. It was the most common "chatbot" the whole run found.
On 8 of those sites a visitor had no way to start a conversation at all. What the rule had found was a WhatsApp icon sitting in the row of social icons at the bottom of the page. In one case it was a share-this-page-to-WhatsApp button. A phone number printed on a poster is not a receptionist.
Of the 23 matches a human could judge, 16 were a link that opens an app, not a chat window on the site itself. Our own published article says, in its own body, that a WhatsApp link is "a link, not an assistant". The rule that produced that article counts it as one. We wrote the correct sentence and shipped the tool that contradicts it.
The misses split three ways, and only one of the three can be fixed by putting more company names on the list. A miss is a site that really does have chat and that our rule walked straight past:
- Chat companies we had never put on the list. Zoho SalesIQ, RetailCRM, Avada, NewOaks, and two more we had never heard of. Six sites.
- Chat companies that were on our list, on pages that never name them. LiveChat, Freshchat and Intercom, each one added to the page a moment after it loads, by a separate tool a site owner uses to bolt extra code on. Our rule reads the page as it arrives, in the box, before anything has been unpacked. Four sites, and invisible to any rule that does not run the page the way a browser does.
- Chat the business built itself and runs on its own address. It lives at
chatbot.<their-domain>,chat.<their-domain>,app.<their-domain>. Four sites. No list of company names can ever catch this kind, because there is no company to name.
03What we are correcting
Our university article named three AI chatbots. Under a stated definition, that the surface says it is AI before you speak to it, at least four institutions had one. The one we missed was running a self-hosted assistant on its own subdomain, which is the exact class no vendor pattern can ever catch, because there is no vendor in it. That is a positive finding about them and an error by us. They are not named here for the same reason none of the forty were named in the original: the mechanism is the useful part.
We also attributed one institution's widget to the wrong vendor. The rule matched a stale vendor string still sitting in the page source while a completely different product was actually running on the page. A detector that reads the source rather than the running page will believe the leftover string every time.
04The rule we now cannot argue with
Our research plan forbids publishing a named list of businesses that are failing at something. Until today that was a belief we held. Now it is a sum anyone can do.
Our rule misses 42 to 43% of the sites that do have chat a visitor can see. So naming a business and saying "they have no chat" means publishing something we can now work out our own odds of being wrong about. Those odds are close to a coin toss.
A chat window that gets added after the page loads becomes "they have no chat." A site that answers our crawler with a 403, which is a server saying no thank you, becomes "they have no chat." A business that built its own assistant and put it on its own address becomes "they have no chat." In all three cases the business is doing the thing, and we would be telling the world it is not.
Reporting only totals is not us being polite. It is the only way a tool can carry its own error rate and still be worth publishing. A wrong answer inside a percentage costs the percentage a little accuracy. A wrong answer with a company's name beside it is a wrong statement about that company. That is why every coverage number we publish is a percentage and never a list.
05What we are doing about it
We are not going to fix the rule and quietly run it again. The frozen version stays published, with these numbers attached to it.
If we build a better detector, it gets reported beside this one as a separate measurement. It never replaces it. We wrote that rule into the plan before we knew we would need it, and a promise like that is only worth something when it is made early.
If you have published a coverage number, check your own instrument
Ours was wrong in a direction we would never have guessed, and we only found out because we went looking. If you have a scanner, a detector or a scraper producing numbers you publish or sell on, we will tell you how we designed this audit and what we would check first. No charge and no pitch, because the method is the whole point.
How we audited it [email protected] · No client of ours is named anywhere on this page, and if you become one, you will not be either.