Measurement standards

Two AI visibility tools can give you different numbers for the same brand on the same day, and both can be right, because they measured different things. This page says exactly what we measure, how well, and what our numbers are not yet good enough for.

Where we stand today

In August 2026 the IAB published Measuring Visibility in the AI Era, a framework for judging whether AI visibility data is fit for the decision you want to make. It sets out seven criteria and two quality tiers: directional and decision-grade.

Measured against those criteria, a typical SearchScore Tracker programme today clears the decision-grade bar on four of the seven, cadence, platform reporting, prompt type coverage and methodology disclosure, sits at directional on reproducibility, and falls below the directional bar on two: sample size and query volume. Both of those come down to one decision, which is how many times we ask each question and how many questions we ask.

We publish that rather than round it up, and we show the same tier inside the product next to your score.

There is no IAB certification, and we do not claim one. The framework states plainly that it "does not establish certification or evaluate individual providers". It describes itself as groundwork for a possible future certification programme. Any vendor telling you they are "IAB certified" today is describing something that does not exist.

Our standing, criterion by criterion

CriterionWhere we areWhy
Sample size Below directional Each question is asked once per engine per scan. A single response is a sample of one, and the framework treats multiple responses per query as the directional bar.
Query volume Below directional 25 questions on the plan you can buy today. The framework "treats fewer than 50 queries per measurement program as exploratory rather than directional".
Reproducibility Directional The directional bar asks a provider to document how much variation is typical, and the ranges below do exactly that. It cannot go higher until we re-run your queries inside the window, which needs the same decision as sample size.
Prompt type coverage Decision-grade All four of the framework's intent types in every question set, with results segmented by intent in the product. This used to depend on your business type, and for accounts that never set one it was a single intent.
Methodology disclosure Decision-grade This page, plus per-scan provenance in the product and a full data export.
Testing cadence Decision-grade Every tracked brand is scanned on a fixed weekly schedule.
Platform reporting Decision-grade Six engines, reported separately and never blended into one figure: ChatGPT, Gemini, Claude, Perplexity, Grok and DeepSeek.
A note on the word "exploratory". The IAB framework defines two tiers, directional and decision-grade. It uses "exploratory" once, for query volume, to describe a programme that falls below directional. We have borrowed that word as our own label for the same idea across all seven criteria. It is our shorthand, not an IAB grade, and we would rather say so than let it look like a rating somebody gave us.

How much do the answers actually move?

AI answers are not stable. Ask the same question twice and you may get a different result, so any single reading carries uncertainty. The framework asks providers to document how much variation is typical, so here is ours.

We asked the same eight questions ten times each across all six engines, in one window, and counted how often the verdict changed between consecutive runs. On a heavily cited brand the average flip rate was 4.2%, ranging from 0% on Claude to 8.3% on DeepSeek. Across our stored history, over roughly 525 comparable pairs per engine, the equivalent figure is 3.2%.

Two honest caveats. Variation depends heavily on the brand: we repeated the exercise on a barely cited domain and five of the six engines never mentioned it at all, which produces a flip rate of zero that means "invisible" rather than "stable". And because we currently ask each question once per scan, these ranges describe what we have measured elsewhere, not a confidence interval on your specific number.

Definitions that differ between vendors

A mention is not a citation

We report both, separately. A mention is your brand appearing anywhere in an answer. A citation is the answer relying on you as a source, either by linking to you or by naming you as the source. Tools that report a single "visibility" number usually mean the first and imply the second.

Questions that name you are scored separately

Asking an AI assistant about a brand by name will usually surface that brand. Counting those in a headline score flatters it, so questions that name you feed a separate brand defence figure instead of your main score. Where a buying question names you and a competitor together, we keep it in the score, because it is a real buying question, but we show you how much of your result rests on it.

A failed call is not a zero

If an engine is unreachable we record no observation. We never count an outage as the engine choosing not to mention you, because that would show as a drop you did not cause. If every engine fails, the scan is discarded rather than saved as a collapse to zero.

What we are doing about the two we fail

Reproducibility has moved to directional because the variation ranges below are now published, which is what that rung asks for. It stays there rather than going higher, because decision-grade needs those ranges measured within the window, and that means re-running your queries, which is the same decision as sample size.

We would rather move these deliberately, and say where we are in the meantime, than quietly describe early signal as decision-grade. The framework's own warning is that the common failure is "treating directional data as decision-grade without recognizing the gap".

Check it yourself

Every Tracker account can export the raw observations behind its score as CSV or JSON: one row per question, engine and scan, with the retrieval mode, the model that actually answered, and the measurement baseline each row belongs to. If you cannot recompute a number, you are taking it on trust, and the whole argument of this page is that you should not have to.

Reviewed August 2026. Framework reference: IAB, Measuring Visibility in the AI Era, August 2026.