Musing23 July 20267 min read

Can AI visibility be measured without pretending to be certain?

I took the seven assumptions underneath my own measurement to the published research and tried to break them. One survived unqualified. The one that broke was mine: a page an AI read is not proof the page shaped what it said.

In shortOf the eight verdicts my own assumptions came back with, one was a clean pass. The survivor indicts the whole market's single-reading number — models set to be deterministic still varied by up to 15% between runs. The one that broke was mine, and it changed the product: I now check whether a cited page supports the claim credited to it, rather than treating a fetch as evidence.

I have never been able to defend, cleanly, the percentage a brand buys for its AI visibility. The figure is precise; nothing about how it was produced is. Rather than defend mine, I took the assumptions underneath my own measurement to the published research and audited them.

I still don't know how much of an AI answer is about your brand.

On the days I sell a single number, I feel like I have failed the honesty the work demands.

Is the number stable? Is it real? Is it yours? A percentage is a translation of a moving thing into a still one, and the translation is where the fragility hides.

The number is precise; the thing it measures moves

Two modes of the same assistant, answering the same questions, shared 25.6% of the sources they cited.[12] Ahrefs states, as a limitation of its own analysis, that its previous research found 45% of AI Overview citations change between generations[11] — a vendor flagging its own surface as unstable, which is what makes it worth quoting.

How much of an answer is even about your brand? A study of 12,933 answers covering 20 brands across eight languages found the language a question was asked in accounted for 26.5% of what moves a single answer, against 1.5% for which brand it was about.[9] That study is a single-author preprint, its brands are Central and Eastern European, and its outcome measure is sentiment, so take it as suggestive — it carries no finding here on its own.

I tried to break my own seven assumptions

Seven load-bearing assumptions sit underneath my measurement. I took all of them to the published research: 54 sources across eight research angles, 222 claims extracted, 48 surviving a three-vote check, and every load-bearing citation re-checked by two independent models from separate providers, neither of which produced the review. The scoreboard, before any finding: eight verdicts came back, because the first assumption was judged as two separate questions. One survived unqualified. Three survived only with qualification, three were challenged, and one had no evidence either way.

8 verdicts — 7 assumptions
PASS
Measuring once is not enough
QUALIFIED
Results are stable to the day
CHALLENGED
A retrieved page shaped the answer
CHALLENGED
Presence and pick rate are independent
CHALLENGED
Position in context does not matter
NO EVIDENCE
The three failure modes are separate
QUALIFIED
Wording does not change the number
QUALIFIED
Geography and language are neutral

1 clean pass · 3 qualified · 3 challenged · 1 no evidence. The first assumption split into two separate questions.

Same settings, same question, different answer

The assumption that survived is the one that indicts the market's single-reading number, not just mine: that measuring once is not enough. Across five models configured to be deterministic, on eight tasks over ten runs, accuracy varied by up to 15% between runs, the gap between the best and worst possible performance reached 70%, and none of the models delivered repeatable accuracy across all tasks.[1]

Run-to-run variance — deterministic models, same task, same settings
Accuracy variation between runsup to 0%
0%100%

Models configured to be deterministic still varied by up to 15%. The gap between best and worst performance reached 70%. This indicts the market's single-reading number, not just this measurement.

A second study traces the cause to evaluation batch size, GPU count and GPU version, rooted in the non-associative nature of floating-point arithmetic — the provider's infrastructure, not a setting a caller can reach.[2] There is a real counterweight: a peer-reviewed study finds that picking the single most likely word at each step generally beats sampling on most tasks.[3] My reply is mine, not theirs — that option is not on offer when you are querying somebody else's interface with live search.

The page it read is not proof it used the page

The assumption that broke was mine: that observing what a model retrieved was evidence about the answer it gave. A study of citation behaviour finds attributed answers often lack genuine reliance on the document they cite — up to 57% of citations were attached after the fact rather than actually used.[4] The version that used the page and the version that did not can produce the same visible output, so the difference cannot be read from outside.

The assumption that broke — citations attached after the fact
Citations with no genuine reliance on the cited pageup to 0%

A retrieved page is not evidence the page shaped the answer. The output is the same either way.

Incomplete support is ordinary, not exceptional: on one long-form question-answering dataset, even the best models lack complete citation support 50% of the time.[6] Where a passage sits changes whether it gets used at all — in multi-document question answering and key-value retrieval, performance is highest when the relevant passage is at the beginning or end of the input and degrades significantly in the middle, even for models built for long inputs.[5]

In a corpus of 17,790 annotated responses, the share carrying at least one unsupported span was 29.1% for question answering and 68.6% for writing from structured data — the regime a shopping answer most resembles.[7] A separate comparison within the same corpus put the two strongest models in it at 9.8%.[8] Those are different slices, not the two ends of one range.

So I check the pages instead of trusting the fetch

The fix was already standing in the literature: output about the world is to be checked against an independent, provided source rather than assumed from the fact that a model produced it.[4][8] A source review now runs on delivered work and asks whether the page actually supports the brand claim credited to it. It works where an engine shows which pages it read; where an engine hides them I say so rather than calling it an absence. And it reduces the exposure — it does not turn retrieval back into evidence that the page shaped the answer.

I still can't show the three failure modes are separate

My refusal to publish a single combined number only holds if the three failure modes I diagnose are genuinely three things. That has never been shown. The first test could not decide it: the sample held 3 brand groups where the estimate needs about 25. It publishes as undecided — not as passed, and not as failed. Undecided is the true answer here, not a softer word for either verdict.

Search benchmarks and school tests, not shopping answers

The heaviest evidence here comes from classical search-engine test collections and educational measurement, carried to AI answers by analogy.[10] Three findings lean on preprints, one of them single-author. The review surfaced no verified evidence at all on personalisation, memory, geography, or drift over time — an absence in the review, not a clean bill for anyone's measurement. Everything measured here runs against provider interfaces with live search rather than the consumer apps, so no personalisation, no conversation memory and no location enters any of it.

What a diagnosis can honestly claim

What is left standing is narrow. A diagnosis can say which failure mode the observed evidence fits, by rules published before the measurement ran — not why a model did what it did, because cause cannot be read from an output. Reliability comes from spread rather than repetition: in that same variance study, brand-ranking reliability was 0.01 from a single answer and about 0.36 across the full spread of languages and models.[9]

So ask a vendor how many phrasings and how many engines their number spans, and whether anyone checked that a cited page supports the claim credited to it. Not how precise the number is. And ask what happens when the answer is unclear — “undecided” has to be something they are allowed to deliver, or the rest is decoration.

A review, not a measurement

This is a review, not a measurement: no new queries were run and no measurement money was spent. Each assumption was written down first, then tested against the published record, and the verdict recorded whether it went my way or not. Two things belong here and nowhere else: three claims were thrown out by my own check before they reached a finding, and one number I had previously used was corrected downward on re-reading the primary source.

How we ran this

What I measured: Seven load-bearing assumptions underneath my own measurement, tested against 54 published sources across eight research angles. 222 claims were extracted and 48 survived a three-vote check. Every load-bearing citation was fetched from its published source and re-checked by two independent models from separate providers, neither of which produced the review.

When: Review conducted July 2026; every source re-fetched and re-verified on 23 July 2026 before publication.

References

  1. Atil, B., Passonneau, R. J., et al. (2024). *Non-determinism of “deterministic” LLM settings*. arXiv. https://arxiv.org/abs/2408.04667
  2. Yuan, J., et al. (2025). *Understanding and mitigating numerical sources of nondeterminism in LLM inference*. arXiv. https://arxiv.org/abs/2506.09501
  3. Song, Y., Wang, G., Li, S., & Lin, B. Y. (2025). *The good, the bad, and the greedy: Evaluation of LLMs should not ignore non-determinism*. NAACL 2025. https://aclanthology.org/2025.naacl-long.211/
  4. Wallat, J., Heuss, M., de Rijke, M., & Anand, A. (2024). *Correctness is not faithfulness in RAG attributions*. arXiv. https://arxiv.org/abs/2412.18004
  5. Liu, N. F., et al. (2023). *Lost in the middle: How language models use long contexts*. arXiv. https://arxiv.org/abs/2307.03172
  6. Gao, T., Yen, H., Yu, J., & Chen, D. (2023). *Enabling large language models to generate text with citations*. EMNLP 2023. https://aclanthology.org/2023.emnlp-main.398/
  7. Niu, C., et al. (2024). *RAGTruth: A hallucination corpus for developing trustworthy retrieval-augmented language models*. arXiv. https://arxiv.org/abs/2401.00396
  8. Rashkin, H., et al. (2023). *Measuring attribution in natural language generation models*. Computational Linguistics. https://aclanthology.org/2023.cl-4.2/
  9. Żatuchin, D. (2026). *Where does the noise come from? A variance-components decomposition of non-determinism in LLM brand answers*. arXiv. https://arxiv.org/abs/2607.13304
  10. Rashidi, L., Zobel, J., & Moffat, A. (2024). *Query variability and experimental consistency: A concerning case study*. ACM. https://dl.acm.org/doi/10.1145/3664190.3672519
  11. Ahrefs. (2026). *Are AI Mode and AI Overviews just different versions of the same answer?*. Ahrefs. https://ahrefs.com/blog/ai-overviews-vs-ai-mode/
  12. Semrush. (2026). *ChatGPT reasoning modes and AI visibility*. Semrush. https://www.semrush.com/blog/chatgpt-reasoning-ai-visibility/

Coda

When a model cites a page it never actually used, what exactly has your visibility number counted?