Musing23 July 20267 min read

Can AI visibility be measured without pretending to be certain?

We took the seven assumptions underneath our own measurement to the published research and tried to break them. One survived unqualified. The one that broke was ours: a page an AI read is not proof the page shaped what it said.

You are buying a percentage for your brand's AI visibility. The number is precise; nothing about how it was produced is. Rather than defend ours, we took the assumptions underneath our own measurement to the published research and audited them.

The number is precise; the thing it measures moves

Two modes of the same assistant, answering the same questions, shared 25.6% of the sources they cited.[12] Ahrefs states, as a limitation of its own analysis, that its previous research found 45% of AI Overview citations change between generations[11] — a vendor flagging its own surface as unstable, which is what makes it worth quoting.

How much of an answer is even about your brand? A study of 12,933 answers covering 20 brands across eight languages found the language a question was asked in accounted for 26.5% of what moves a single answer, against 1.5% for which brand it was about.[9] That study is a single-author preprint, its brands are Central and Eastern European, and its outcome measure is sentiment, so take it as suggestive — it carries no finding here on its own.

We tried to break our own seven assumptions

Seven load-bearing assumptions sit underneath our measurement. We took all of them to the published research: 54 sources across eight research angles, 222 claims extracted, 48 surviving a three-vote check, and every load-bearing citation re-checked by two independent models from separate providers, neither of which produced the review. The scoreboard, before any finding: eight verdicts came back, because the first assumption was judged as two separate questions. One survived unqualified. Three survived only with qualification, three were challenged, and one had no evidence either way.

Same settings, same question, different answer

The assumption that survived is the one that indicts the market's single-reading number, not just ours: that measuring once is not enough. Across five models configured to be deterministic, on eight tasks over ten runs, accuracy varied by up to 15% between runs, the gap between the best and worst possible performance reached 70%, and none of the models delivered repeatable accuracy across all tasks.[1]

A second study traces the cause to evaluation batch size, GPU count and GPU version, rooted in the non-associative nature of floating-point arithmetic — the provider's infrastructure, not a setting a caller can reach.[2] There is a real counterweight: a peer-reviewed study finds that picking the single most likely word at each step generally beats sampling on most tasks.[3] Our reply is ours, not theirs — that option is not on offer when you are querying somebody else's interface with live search.

The page it read is not proof it used the page

The assumption that broke was ours: that observing what a model retrieved was evidence about the answer it gave. A study of citation behaviour finds attributed answers often lack genuine reliance on the document they cite — up to 57% of citations were attached after the fact rather than actually used.[4] The version that used the page and the version that did not can produce the same visible output, so the difference cannot be read from outside.

Incomplete support is ordinary, not exceptional: on one long-form question-answering dataset, even the best models lack complete citation support 50% of the time.[6] Where a passage sits changes whether it gets used at all — in multi-document question answering and key-value retrieval, performance is highest when the relevant passage is at the beginning or end of the input and degrades significantly in the middle, even for models built for long inputs.[5]

In a corpus of 17,790 annotated responses, the share carrying at least one unsupported span was 29.1% for question answering and 68.6% for writing from structured data — the regime a shopping answer most resembles.[7] A separate comparison within the same corpus put the two strongest models in it at 9.8%.[8] Those are different slices, not the two ends of one range.

So we check the pages instead of trusting the fetch

The fix was already standing in the literature: output about the world is to be checked against an independent, provided source rather than assumed from the fact that a model produced it.[4][8] A source review now runs on delivered work and asks whether the page actually supports the brand claim credited to it. It works where an engine shows which pages it read; where an engine hides them we say so rather than calling it an absence. And it reduces the exposure — it does not turn retrieval back into evidence that the page shaped the answer.

We still can't show the three failure modes are separate

Our refusal to publish a single combined number only holds if the three failure modes we diagnose are genuinely three things. That has never been shown. The first test could not decide it: the sample held 3 brand groups where the estimate needs about 25. It publishes as undecided — not as passed, and not as failed. Undecided is the true answer here, not a softer word for either verdict.

Search benchmarks and school tests, not shopping answers

The heaviest evidence here comes from classical search-engine test collections and educational measurement, carried to AI answers by analogy.[10] Three findings lean on preprints, one of them single-author. The review surfaced no verified evidence at all on personalisation, memory, geography, or drift over time — an absence in the review, not a clean bill for anyone's measurement. Everything measured here runs against provider interfaces with live search rather than the consumer apps, so no personalisation, no conversation memory and no location enters any of it.

What a diagnosis can honestly claim

What is left standing is narrow. A diagnosis can say which failure mode the observed evidence fits, by rules published before the measurement ran — not why a model did what it did, because cause cannot be read from an output. Reliability comes from spread rather than repetition: in that same variance study, brand-ranking reliability was 0.01 from a single answer and about 0.36 across the full spread of languages and models.[9]

So ask a vendor how many phrasings and how many engines their number spans, and whether anyone checked that a cited page supports the claim credited to it. Not how precise the number is. And ask what happens when the answer is unclear — “undecided” has to be something they are allowed to deliver, or the rest is decoration.

A review, not a measurement

This is a review, not a measurement: no new queries were run and no measurement money was spent. Each assumption was written down first, then tested against the published record, and the verdict recorded whether it went our way or not. Two things belong here and nowhere else: three claims were thrown out by our own check before they reached a finding, and one number we had previously used was corrected downward on re-reading the primary source.

How we ran this

What we measured: Seven load-bearing assumptions underneath our own measurement, tested against 54 published sources across eight research angles. 222 claims were extracted and 48 survived a three-vote check. Every load-bearing citation was fetched from its published source and re-checked by two independent models from separate providers, neither of which produced the review.

When: Review conducted July 2026; every source re-fetched and re-verified on 23 July 2026 before publication.

References

  1. Atil, B., Passonneau, R. J., et al. (2024). *Non-determinism of “deterministic” LLM settings*. arXiv. https://arxiv.org/abs/2408.04667
  2. Yuan, J., et al. (2025). *Understanding and mitigating numerical sources of nondeterminism in LLM inference*. arXiv. https://arxiv.org/abs/2506.09501
  3. Song, Y., Wang, G., Li, S., & Lin, B. Y. (2025). *The good, the bad, and the greedy: Evaluation of LLMs should not ignore non-determinism*. NAACL 2025. https://aclanthology.org/2025.naacl-long.211/
  4. Wallat, J., Heuss, M., de Rijke, M., & Anand, A. (2024). *Correctness is not faithfulness in RAG attributions*. arXiv. https://arxiv.org/abs/2412.18004
  5. Liu, N. F., et al. (2023). *Lost in the middle: How language models use long contexts*. arXiv. https://arxiv.org/abs/2307.03172
  6. Gao, T., Yen, H., Yu, J., & Chen, D. (2023). *Enabling large language models to generate text with citations*. EMNLP 2023. https://aclanthology.org/2023.emnlp-main.398/
  7. Niu, C., et al. (2024). *RAGTruth: A hallucination corpus for developing trustworthy retrieval-augmented language models*. arXiv. https://arxiv.org/abs/2401.00396
  8. Rashkin, H., et al. (2023). *Measuring attribution in natural language generation models*. Computational Linguistics. https://aclanthology.org/2023.cl-4.2/
  9. Żatuchin, D. (2026). *Where does the noise come from? A variance-components decomposition of non-determinism in LLM brand answers*. arXiv. https://arxiv.org/abs/2607.13304
  10. Rashidi, L., Zobel, J., & Moffat, A. (2024). *Query variability and experimental consistency: A concerning case study*. ACM. https://dl.acm.org/doi/10.1145/3664190.3672519
  11. Ahrefs. (2026). *Are AI Mode and AI Overviews just different versions of the same answer?*. Ahrefs. https://ahrefs.com/blog/ai-overviews-vs-ai-mode/
  12. Semrush. (2026). *ChatGPT reasoning modes and AI visibility*. Semrush. https://www.semrush.com/blog/chatgpt-reasoning-ai-visibility/