Measurement standards
Two AI visibility tools can give you different numbers for the same brand on the same day, and both can be right, because they measured different things. This page says exactly what we measure, how well, and what our numbers are not yet good enough for.
Where we stand today
In August 2026 the IAB published Measuring Visibility in the AI Era, a framework for judging whether AI visibility data is fit for the decision you want to make. It sets out seven criteria and two quality tiers: directional and decision-grade.
Measured against those criteria, a typical SearchScore Tracker programme today clears the decision-grade bar on four of the seven, cadence, platform reporting, prompt type coverage and methodology disclosure, sits at directional on reproducibility, and falls below the directional bar on two: sample size and query volume. Both of those come down to one decision, which is how many times we ask each question and how many questions we ask.
We publish that rather than round it up, and we show the same tier inside the product next to your score.
Our standing, criterion by criterion
| Criterion | Where we are | Why |
|---|---|---|
| Sample size | Below directional | Each question is asked once per engine per scan. A single response is a sample of one, and the framework treats multiple responses per query as the directional bar. |
| Query volume | Below directional | 25 questions on the plan you can buy today. The framework "treats fewer than 50 queries per measurement program as exploratory rather than directional". |
| Reproducibility | Directional | The directional bar asks a provider to document how much variation is typical, and the ranges below do exactly that. It cannot go higher until we re-run your queries inside the window, which needs the same decision as sample size. |
| Prompt type coverage | Decision-grade | All four of the framework's intent types in every question set, with results segmented by intent in the product. This used to depend on your business type, and for accounts that never set one it was a single intent. |
| Methodology disclosure | Decision-grade | This page, plus per-scan provenance in the product and a full data export. |
| Testing cadence | Decision-grade | Every tracked brand is scanned on a fixed weekly schedule. |
| Platform reporting | Decision-grade | Six engines, reported separately and never blended into one figure: ChatGPT, Gemini, Claude, Perplexity, Grok and DeepSeek. |
How much do the answers actually move?
AI answers are not stable. Ask the same question twice and you may get a different result, so any single reading carries uncertainty. The framework asks providers to document how much variation is typical, so here is ours.
We asked the same eight questions ten times each across all six engines, in one window, and counted how often the verdict changed between consecutive runs. On a heavily cited brand the average flip rate was 4.2%, ranging from 0% on Claude to 8.3% on DeepSeek. Across our stored history, over roughly 525 comparable pairs per engine, the equivalent figure is 3.2%.
Two honest caveats. Variation depends heavily on the brand: we repeated the exercise on a barely cited domain and five of the six engines never mentioned it at all, which produces a flip rate of zero that means "invisible" rather than "stable". And because we currently ask each question once per scan, these ranges describe what we have measured elsewhere, not a confidence interval on your specific number.
Definitions that differ between vendors
A mention is not a citation
We report both, separately. A mention is your brand appearing anywhere in an answer. A citation is the answer relying on you as a source, either by linking to you or by naming you as the source. Tools that report a single "visibility" number usually mean the first and imply the second.
Questions that name you are scored separately
Asking an AI assistant about a brand by name will usually surface that brand. Counting those in a headline score flatters it, so questions that name you feed a separate brand defence figure instead of your main score. Where a buying question names you and a competitor together, we keep it in the score, because it is a real buying question, but we show you how much of your result rests on it.
A failed call is not a zero
If an engine is unreachable we record no observation. We never count an outage as the engine choosing not to mention you, because that would show as a drop you did not cause. If every engine fails, the scan is discarded rather than saved as a collapse to zero.
What we are doing about the two we fail
- Sample size. The mechanism for asking each question several times and reporting the frequency is built and tested. Turning it on multiplies the cost of every scan, so it is a pricing decision as much as an engineering one.
- Query volume. Reaching the 50-query floor means roughly doubling the tracked set, with the same cost consequence. Until recently it was not purchasable at any price: our question generator could not actually produce more than about 45 unique questions, whatever the plan said, so the volume floor was blocked by a defect rather than by a decision. That is fixed, and 50 and 100 are now exact.
Reproducibility has moved to directional because the variation ranges below are now published, which is what that rung asks for. It stays there rather than going higher, because decision-grade needs those ranges measured within the window, and that means re-running your queries, which is the same decision as sample size.
We would rather move these deliberately, and say where we are in the meantime, than quietly describe early signal as decision-grade. The framework's own warning is that the common failure is "treating directional data as decision-grade without recognizing the gap".
Check it yourself
Every Tracker account can export the raw observations behind its score as CSV or JSON: one row per question, engine and scan, with the retrieval mode, the model that actually answered, and the measurement baseline each row belongs to. If you cannot recompute a number, you are taking it on trust, and the whole argument of this page is that you should not have to.
Reviewed August 2026. Framework reference: IAB, Measuring Visibility in the AI Era, August 2026.