How accurate is your AI visibility tool?
Almost every percentage an AI visibility platform shows you rests on a classifier deciding whether an answer named your brand. If that classifier is unreliable you have no way to know how far out your dashboard is. Here is how to check, what a real answer looks like, and what happened when we ran the test on ourselves.
Every AI visibility platform, ours included, sells you a number. Share of voice, visibility score, mention rate, presence across the engines.
Underneath all of them sits one unglamorous piece of software: a classifier that reads an assistant’s answer and decides whether your brand was named in it. Every figure you are shown is downstream of that single decision, repeated a few thousand times.
So the first question to ask any vendor in this category, before pricing or engine coverage or dashboard design, is how often that classifier is right.
Why it is the load-bearing question
Consider what the classifier has to get right.
An assistant says “there are several good options, including Acme”. Straightforward. An assistant says “I am not familiar with Acme, but similar companies include…” and a naive matcher sees the brand name and counts a mention, when the engine has just told the user it has never heard of you. An assistant writes “acme” inside a URL, or mentions Acme Corp when you are Acme Digital, or refers to you only as “the Manchester agency mentioned above”.
Each of those is a judgement call, and each one lands in your score as a one or a zero.
Now multiply. A weekly scan asks dozens of questions across six assistants, sometimes several times over. A classifier that misreads a fair share of those calls is not giving you a slightly fuzzy picture; it is giving you a number whose error you cannot see and cannot bound, because misreadings in opposite directions do not have to cancel out.
Nothing about a wrong classifier looks broken. That is the problem.
The four questions
Ask any vendor these. They take a minute and they separate the field quickly.
1. Have you measured your classifier’s accuracy, and what is the number?
You want a figure with a sample size attached and a date on it. A percentage on its own, with neither, is a marketing claim wearing a measurement’s clothes.
2. Who did the labelling?
This is the one that matters most and the one nobody expects. If the people who built the classifier also decided which answers counted as mentions, the measurement is circular: they will label the edge cases the way their own code does, without meaning to. A real measurement needs an annotator who has never touched the classifier, working blind, with the software’s verdict hidden until after they have chosen.
3. Can I see the labelled sample?
Published labels mean anyone can re-run the measurement against a later version and check the number still holds. Without them you are taking the vendor’s word for the labels and the arithmetic both.
4. Does “mention” mean the same thing as “citation” in your product?
These are different events. A mention is the assistant naming your brand. A citation is the assistant linking to your site as a source. On our own stored observations the strict citation rate runs more than five times lower than the mention rate. Both are worth measuring, and a tool that counts the first while calling it the second is describing a better position than the one you hold.
What happened when we ran it on ourselves
We publish our answer, so here it is in full, including the parts that are unflattering. As of September 2026 it is the only blind, out-of-sample accuracy measurement for a brand-mention classifier we have been able to find published by anyone in this category, and we would rather it were not.
We measured our own mention detector on 19 August 2026, blind and out of sample. 47 stored answer cells, labelled by an annotator who has never edited the classifier, with the software’s verdict hidden until after each choice.
| Result | 95% interval (Wilson) | |
|---|---|---|
| Precision | 7 of 7 correct, no false positives | 64.6% to 100% |
| Recall | 7 of 8 (87.5%) | 52.9% to 97.8% |
| Agreement with the human label, on this stratified sample | 46 of 47 (97.9%) | 88.9% to 99.6% |
Read the precision row carefully, because how it is written is deliberate.
We report it as a count and not as a percentage. The finding is that there were no false positives in 47 blind-labelled cells. We publish the interval beside it, running from 64.6% upwards, because seven positives cannot carry a point estimate and rounding that row into a perfect-sounding percentage would claim far more than 47 cells can support.
Three further limits:
- The sample is stratified, not random. It over-weights the cells where the classifier and a naive matcher disagree. That makes it a harder test than a random draw, and it also means these are not population estimates for your account.
- There is one annotator, so inter-rater agreement is unmeasured.
- The intervals overlap our earlier, pre-fix study, which measured 66.7% precision on 126 cells before we repaired the failure modes it found. So this result is consistent with the fixes having worked. It is not proof that they did.
There is one disagreement in the 47. It is a false negative we had already documented, and a second annotator, blind to the first, had independently called that single cell the same way.
What the fixes were actually worth
Accuracy studies are small by nature, so here is a larger measurement of the same change from a different angle.
Re-scoring all 5,977 stored answer cells with the tightened classifier removes 11 answers that previously counted as mentions and adds none. The stored mention rate moves from 21.78% to 21.60%.
Every one of those 11 was read by hand. Each is an engine saying it had never heard of the brand, and being counted as a mention anyway.
That is a small correction, and it only ever moves in one direction: down. A classifier change that made our own numbers look better would deserve considerably more scepticism than one that quietly removes eleven mentions from our own published rate.
The measurement re-runs against any future version of the code with one command:
node tools/brand-mention-regress.js --fixture tests/fixtures/accuracy/brand-mention-sample-2026-08-18.json --verbose
What a good answer looks like, and what a bad one looks like
A good answer has a date, a sample size, an interval, and a named limitation. It will probably be less flattering than the marketing page, because a measured number usually is.
A weak answer has one of these shapes:
- A round percentage with no sample size behind it. Somebody chose that number. Nobody measured it.
- Accuracy measured by the team that wrote the classifier, which is where the edge cases break.
- No distinction between mention and citation, or the two words used interchangeably in the same document.
- An accuracy figure with no date. Classifiers change, and a figure measured before a fix describes software that no longer exists. That is why we retired our own earlier numbers instead of continuing to quote them.
What we found when we looked
In September 2026 we went looking for a published accuracy figure from every platform in this category: Profound, Peec AI, Otterly, Evertune, Scrunch, AthenaHQ, Ahrefs Brand Radar and the Semrush AI Toolkit.
Two of them publish a methodology page. Ahrefs Brand Radar’s is the more candid document in the category and worth reading: it states plainly that its impressions are modelled rather than measured, and that it does not claim a validated link between Google search volume and how often a question is asked inside an AI tool. Evertune publishes its measurement approach in detail.
Neither states a precision figure, a recall figure, or the size of a labelled sample for deciding whether an answer named a brand. Nor could we find one from the other six.
Two things that finding does not mean. It does not mean these tools are inaccurate; an unpublished measurement is not a failed one, and several of these teams are visibly careful about what their numbers do and do not represent. And it is a search of what is public, in one month, so any vendor holding a figure we missed can send it and we will add it here.
What it does mean is that if you want to know how often your tool reads an answer correctly, right now you mostly have to run the test yourself.
Measuring this properly is genuinely awkward, which is the likeliest explanation: you need labelled data, you need an annotator who is not the author, and you have to publish a number you cannot control in advance.
What to do with this
If you are choosing a platform, ask the four questions before you compare prices. The answers will tell you more about a vendor than any feature table, and the second one alone will thin the field.
If you already use one, run the test yourself. Export thirty or forty classified answers, hide the tool’s verdict, and label them by hand. Look particularly for answers where the assistant said it did not recognise the brand and the tool counted a mention anyway. That was the failure mode in our own software, and it is the cheapest one to check for.
And if you want the whole method, our measurement standards page carries the full working, including the figures on this page, generated from the committed labelled sample rather than typed in by hand.
We would rather publish a small honest number with a wide interval than a large one with no sample size behind it.
Frequently asked questions
Which AI visibility tool is the most accurate?
Ask each vendor four questions: how was accuracy measured and on what sample size, who did the labelling, can you see the labelled sample, and does a mention mean the same thing as a citation in the product. If an answer arrives without a sample size, treat the figure as a claim and not a measurement. The second question does most of the work, because a classifier graded by the people who wrote it is grading its own homework.
What is a mention detector and why does its accuracy matter?
It is the component that reads an AI assistant's answer and decides whether your brand was named. Every other number a platform shows you is built on top of that decision: your visibility score, your share of voice, your trend line, your competitor comparison. If the classifier is unreliable, the figures downstream inherit the problem, and nothing on the dashboard will look broken.
Is a brand mention the same as a citation?
No, and the gap is large. A mention is the assistant naming your brand in its answer. A citation is the assistant linking to your site as a source. On our own stored observations the strict citation rate runs more than five times lower than the mention rate. Both are legitimate metrics, but a tool that counts mentions and labels them citations is describing a better position than the one you hold.
How can I check an AI visibility tool's accuracy myself?
Export a sample of the answers the tool has classified, hide its verdict, and label thirty or forty of them yourself. Then compare. Look for two failure modes: answers where the tool says you were mentioned and you were not, and answers where the assistant said it had never heard of you but the tool counted a mention anyway. The second is the one we found in our own software before we fixed it, so it is worth checking first.