Can AI visibility be measured without pretending to be certain?

We took the seven assumptions underneath our own measurement to the published research and tried to break them. One survived unqualified. The one that broke was ours: a page an AI read is not proof the page shaped what it said.

Of the eight verdicts our own assumptions came back with, one was a clean pass. The survivor indicts the whole market's single-reading number — models set to be deterministic still varied by up to 15% between runs. The one that broke was ours, and it changed the product: we now check whether a cited page supports the claim credited to it, rather than treating a fetch as evidence.

23 July 2026

You are buying a percentage for your brand's AI visibility. The number is precise; nothing about how it was produced is. Rather than defend ours, we took the assumptions underneath our own measurement to the published research and audited them.

The number is precise; the thing it measures moves

Two modes of the same assistant, answering the same questions, shared 25.6% of the sources they cited. Ahrefs states, as a limitation of its own analysis, that its previous research found 45% of AI Overview citations change between generations — a vendor flagging its own surface as unstable, which is what makes it worth quoting.

How much of an answer is even about your brand? A study of 12,933 answers covering 20 brands across eight languages found the language a question was asked in accounted for 26.5% of what moves a single answer, against 1.5% for which brand it was about. That study is a single-author preprint, its brands are Central and Eastern European, and its outcome measure is sentiment, so take it as suggestive — it carries no finding here on its own.

We tried to break our own seven assumptions

7 assumptions, 8 verdicts

One survived unqualified. 3 came back qualified, 3 challenged, 1 untested.

survived
qualified
qualified
qualified
challenged
challenged
challenged
untested

Seven load-bearing assumptions sit underneath our measurement. We took all of them to the published research: 54 sources across eight research angles, 222 claims extracted, 48 surviving a three-vote check, and every load-bearing citation re-checked by two independent models from separate providers, neither of which produced the review. The scoreboard, before any finding: eight verdicts came back, because the first assumption was judged as two separate questions. One survived unqualified. Three survived only with qualification, three were challenged, and one had no evidence either way.

Same settings, same question, different answer

Same settings, different answers

Accuracy spread across runs of models configured to be deterministic: up to 15% between runs, up to 70% between the best and worst possible.

0153570run to runbest to worstaccuracy spread, percentage points

The assumption that survived is the one that indicts the market's single-reading number, not just ours: that measuring once is not enough. Across five models configured to be deterministic, on eight tasks over ten runs, accuracy varied by up to 15% between runs, the gap between the best and worst possible performance reached 70%, and none of the models delivered repeatable accuracy across all tasks.

A second study traces the cause to evaluation batch size, GPU count and GPU version, rooted in the non-associative nature of floating-point arithmetic — the provider's infrastructure, not a setting a caller can reach. There is a real counterweight: a peer-reviewed study finds that picking the single most likely word at each step generally beats sampling on most tasks. Our reply is ours, not theirs — that option is not on offer when you are querying somebody else's interface with live search.

The page it read is not proof it used the page

The assumption that broke was ours: that observing what a model retrieved was evidence about the answer it gave. A study of citation behaviour finds attributed answers often lack genuine reliance on the document they cite — up to 57% of citations were attached after the fact rather than actually used. The version that used the page and the version that did not can produce the same visible output, so the difference cannot be read from outside.

Incomplete support is ordinary, not exceptional: on one long-form question-answering dataset, even the best models lack complete citation support 50% of the time. Where a passage sits changes whether it gets used at all — in multi-document question answering and key-value retrieval, performance is highest when the relevant passage is at the beginning or end of the input and degrades significantly in the middle, even for models built for long inputs.

Unsupported spans, by what the model was doing

Share of responses carrying at least one unsupported span across 17,790 annotated responses: 29.1% answering questions, 68.6% writing from structured data. The 9.8% mark is a separate comparison of the two strongest models in the same corpus.

0%25%50%75%100%question answeringwriting from structured datatwo strongest modelsseparate comparison

In a corpus of 17,790 annotated responses, the share carrying at least one unsupported span was 29.1% for question answering and 68.6% for writing from structured data — the regime a shopping answer most resembles. A separate comparison within the same corpus put the two strongest models in it at 9.8%. Those are different slices, not the two ends of one range.

So we check the pages instead of trusting the fetch

The fix was already standing in the literature: output about the world is to be checked against an independent, provided source rather than assumed from the fact that a model produced it. A source review now runs on delivered work and asks whether the page actually supports the brand claim credited to it. It works where an engine shows which pages it read; where an engine hides them we say so rather than calling it an absence. And it reduces the exposure — it does not turn retrieval back into evidence that the page shaped the answer.

We still can't show the three failure modes are separate

3 brand groups where the test needs 25

The first test of whether the three failure modes separate had 3 brand groups to work with. The estimate needs about 25 before it can return an answer either way.

filled: brand groups the test had · outlined: what it still needs

Our refusal to publish a single combined number only holds if the three failure modes we diagnose are genuinely three things. That has never been shown. The first test could not decide it: the sample held 3 brand groups where the estimate needs about 25. It publishes as undecided — not as passed, and not as failed. Undecided is the true answer here, not a softer word for either verdict.

Search benchmarks and school tests, not shopping answers

The heaviest evidence here comes from classical search-engine test collections and educational measurement, carried to AI answers by analogy. Three findings lean on preprints, one of them single-author. The review surfaced no verified evidence at all on personalisation, memory, geography, or drift over time — an absence in the review, not a clean bill for anyone's measurement. Everything measured here runs against provider interfaces with live search rather than the consumer apps, so no personalisation, no conversation memory and no location enters any of it.

What a diagnosis can honestly claim

What is left standing is narrow. A diagnosis can say which failure mode the observed evidence fits, by rules published before the measurement ran — not why a model did what it did, because cause cannot be read from an output. Reliability comes from spread rather than repetition: in that same variance study, brand-ranking reliability was 0.01 from a single answer and about 0.36 across the full spread of languages and models.

So ask a vendor how many phrasings and how many engines their number spans, and whether anyone checked that a cited page supports the claim credited to it. Not how precise the number is. And ask what happens when the answer is unclear — “undecided” has to be something they are allowed to deliver, or the rest is decoration.

A review, not a measurement

This is a review, not a measurement: no new queries were run and no measurement money was spent. Each assumption was written down first, then tested against the published record, and the verdict recorded whether it went our way or not. Two things belong here and nowhere else: three claims were thrown out by our own check before they reached a finding, and one number we had previously used was corrected downward on re-reading the primary source.

How we ran this

What we measured
Seven load-bearing assumptions underneath our own measurement, tested against 54 published sources across eight research angles. 222 claims were extracted and 48 survived a three-vote check. Every load-bearing citation was fetched from its published source and re-checked by two independent models from separate providers, neither of which produced the review.
When
Review conducted July 2026; every source re-fetched and re-verified on 23 July 2026 before publication.
Which AI
No new measurement was run for this piece — it tests our own assumptions against published research, Citations fetched from the publishers and checked by two independent models from separate providers

The full numbers

Every number in this piece, at full precision. The prose rounds for reading; this table doesn't.

MeasureAs shownExactSource
Assumptions under our own measurement that were stress-tested77Measured
Verdicts returned (assumption 1 was judged as two separate questions)88Derived
Assumptions that survived unqualified11Measured
Assumptions that survived only with qualification33Measured
Assumptions the literature challenged33Measured
Assumptions with no evidence either way11Measured
Verdicts that were not a clean pass77Derived
Brand groups the first separateness test had to work with33Measured
Brand groups that test needs before it can return an answer2525Measured
Published sources reviewed5454Measured
Claims extracted from those sources222222Measured
Claims that survived the review's own three-vote check4848Measured
Independent votes each claim had to pass33Measured
Claims our own check threw out33Measured
Independent checkers each citation passed, from separate providers22Measured
Separate research angles searched88Measured
Vendor's own stated limitation — AI Overview citations that change between generations45%45%Cited
Repeatability study — accuracy variation between runs of the same settings15%15%Cited
Repeatability study — gap between best and worst possible performance70%70%Cited
Repeatability study — models tested55Cited
Repeatability study — runs per task1010Cited
Repeatability study — tasks tested88Cited
Citation benchmark — share of the time the best models lack complete citation support on one question-answering dataset50%50%Cited
Grounding corpus — responses with at least one unsupported span, writing from structured data68.6%68.6%Cited
Grounding corpus — the same measure for the two strongest models in it, on a separate 450-response comparison9.8%9.8%Cited
Grounding corpus — responses with at least one unsupported span, question answering29.1%29.1%Cited
Grounding corpus — responses annotated17,79017,790Cited
Vendor study — cited-source overlap between two modes of the same assistant25.6%25.6%Cited
Citation study — share of citations attached after the fact57%57%Cited
Variance study — share of one answer explained by which brand it is about1.5%1.5%Cited
Variance study — brands covered (Central and Eastern European)2020Cited
Variance study — share of one answer explained by the language the question was asked in26.5%26.5%Cited
Variance study — languages covered88Cited
Variance study — brand-ranking reliability across the full spread of languages and models0.360.36Cited
Variance study — brand-ranking reliability from a single answer0.010.01Cited
Variance study — responses analysed12,93312,933Cited

What this doesn't settle

  • The heaviest evidence comes from classical search-engine test collections and educational measurement, not from AI shopping answers — it bears on this by analogy, not as a direct test.
  • Three findings lean on preprints. The one behind the brand-versus-language split is single-author, covers Central and Eastern European brands across eight languages, and measures sentiment rather than recommendation rates.
  • The review surfaced no verified evidence on personalisation, memory, geography, or drift over time. That is a gap in what was reviewed, not a finding that those threats are small.
  • Each verdict judges an assumption as literally stated. Several challenges soften where the measurement in practice is more careful than the assumption written down.
  • Whether the three failure modes are empirically separate is still undecided — the first test had too few brand groups to answer it either way.

References

  1. Berk Atil, Rebecca J. Passonneau, and colleagues. Non-Determinism of “Deterministic” LLM Settings Accessed 23 July 2026.
  2. Jiayi Yuan and colleagues. Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference Accessed 23 July 2026.
  3. Yifan Song, Guoyin Wang, Sujian Li, and Bill Yuchen Lin. The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism Accessed 23 July 2026.
  4. Jonas Wallat, Maria Heuss, Maarten de Rijke, and Avishek Anand. Correctness is not Faithfulness in RAG Attributions Accessed 23 July 2026.
  5. Nelson F. Liu and colleagues. Lost in the Middle: How Language Models Use Long Contexts Accessed 23 July 2026.
  6. Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. Enabling Large Language Models to Generate Text with Citations Accessed 23 July 2026.
  7. Cheng Niu and colleagues. RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models Accessed 23 July 2026.
  8. Hannah Rashkin and colleagues. Measuring Attribution in Natural Language Generation Models Accessed 23 July 2026.
  9. D. Żatuchin. Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers Accessed 23 July 2026.
  10. Lida Rashidi, Justin Zobel, and Alistair Moffat. Query Variability and Experimental Consistency: A Concerning Case Study Accessed 23 July 2026.
  11. Ahrefs. Are AI Mode and AI Overviews Just Different Versions of the Same Answer? Accessed 23 July 2026.
  12. Semrush. ChatGPT reasoning modes and AI visibility Accessed 23 July 2026.

How many phrasings and engines is your AI visibility number built on?

A diagnosis runs the same questions twice, across phrasings and engines, and shows the pages behind each answer — with “inconclusive” as a result it is allowed to return.

See what a diagnosis finds