Musing23 July 20269 min read

Would ChatGPT still recommend your brand if you asked it a different way?

A re-analysis of 8,160 measurements we already held, turned on our own two-run reliability guarantee. Holding the wording still across both runs left a whole class of error untested, and the answer forces a change to the instrument.

Two runs agreed — about the day, not the wording

Most AI-visibility numbers a brand buys rest on one reassurance: the shop ran the measurement twice, on two different days, and the two runs agreed. Two runs agreeing shows the answer is stable to when you ask — to the day. It shows nothing about how you ask, because the wording of the question was held fixed across both runs.

The question a buyer types into ChatGPT is not one question. It is a distribution of phrasings — "best skincare brands," "top skincare brands," "skincare brands worth buying," "best skincare brands under $30" — and a number measured on one phrasing is a sample of size one from that distribution. When you ask is the day axis: two runs on different days, everything else held still, tests whether day-to-day sampling moves the answer. How you ask is the wording axis: the same intent phrased several ways, tested against itself. The two-run check holds the second axis still and varies the first.

That is a design choice — the wordings are picked once and reused — not a discovered property of the measurement. For a wording bias shared by both runs, the bias is a constant that cancels in the difference between them; a constant does not appear as a disagreement. The check is therefore blind to the wording design's contribution — not to every disagreement near a line; day-to-day noise can still flip a brand there, and that disagreement says the day moved it, not the wording. That blindness is structural: true whenever both runs share a wording set, independent of this data. It does not need this category.

Any brand the check ever called stable on wording-decided ground was stable to the day, not to the wording — the diagnosis would have read "reliable" but meant only "reliable to the day." A reliability guarantee has a hole exactly the shape of the thing it claims to guard. What the wording turns out to be worth is the question this piece answers, with measurements we already held — 8,160 of them, from our own archive. Two runs agreeing was never evidence about wording.

The one bet we placed before looking

Before looking at any outcome, we fixed one pass line: a wording-induced spread above 0.05 would mean eight wordings isn't enough — locked in before the result, along with the category itself, so the verdict below isn't a story fitted after the fact.

At eight wordings, the measured spread came in at 0.078. The verdict: insufficient. That is a disclosed test outcome, not a claim that eight is universally insufficient — what eight fails to cover is the rest of this piece.

One brand read 0.63 asked one way, 0.16 asked another

Rephrasing the question moves the number. In this one skincare category, three cosmetic variants of the same intent — wordings that should be equivalent — move a brand-engine cell's measured rate by 0.20 on average and by 0.47 at worst. One cell, Cetaphil on OpenAI, read 0.63 asked one equivalent way and 0.16 asked another, across three ways of asking the same thing. Wording is a first-order error source in this category's measurement, not a rounding detail. "First-order" is a judgment about size — this archive's spread against this archive's band width — and it reads as structural when it is not; it is fenced to this one category.

The population the finding is stated on is narrow on purpose: six mid-tier skincare brands on two engines, one buyer profile, one collection window, one archive of 8,160 measurements. Every magnitude number here is one category's measurement. The structural claim — that a reliability number measured on one wording set is a function of that set — does not need this data and arrives unfenced.

A rival reading has to be named where it would explain the result away: maybe the 0.20 average is just noise from too few repeats of each wording. It is ruled out not by a model but by one registered fact — each rate rests on 51 asks of that exact wording per engine, across three batches, not a handful.

The wordings that move a rate most are not the equivalent ones. The three cosmetic variants move it by 0.20 on average; all eight together move it by 0.585 on average, so the equivalent-rewording spread is just over a third of the all-eight spread. Narrower wordings — a price ceiling, a channel constraint — differ more, because they change what the question asks for. This data contains no analysis of how a wording routes through the model; if a reader asks why the wordings move the rate, the honest answer is that the question is not tested here. What is tested is that they do.

Six cells the model is sure about, six it isn't

The twelve cells do not sit on a sliding scale of wording-sensitivity. They split clean, and the split is the structural claim. Six cells land on the same result for every subset of two or more of the eight wordings — zero flips, at every subset size from two to eight. The other six change result on a third to over half of the 70 four-of-eight subsets. "A third to over half" is the four-of-eight figure and travels with that denominator; it is not a constant across subset sizes, and at seven-of-eight one of those same six cells reads zero. The split is two clusters with little between.

A load-bearing exception: the one-wording case is the single outlier. Two of the six proof cells — Kiehl's on Anthropic and La Roche-Posay on OpenAI — flip on one of their eight one-wording subsets and on zero at every size from two to eight. "Zero flips at every subset size from two to eight" is the verified claim, not "every possible subset," because the one-wording case falsifies that.

Being wording-proof is not a virtue of the brand — it is a position property. The proof cells sit far from any decision line: La Roche-Posay sits near certainty on both engines, 0.919 on Anthropic and 0.944 on OpenAI, so a wording move has nowhere to push it across. The decided cells sit near a line: Cetaphil 0.355 on OpenAI, just below 0.40, and 0.417 on Anthropic, just above — within a wording move of the boundary, so wording can push it across. Whether a wording move changes the verdict is decided entirely by whether the rate started near a line. The proof cells are not better-measured; they are the ones the model is sure about either way. The decided cells are the ones the model is ambivalent about, so every wording can. That instability concentrates near a decision line is a structural form and publishes without a category fence.

What is fenced to this category is the shape of the split: that it is clean and roughly half-and-half is this archive's configuration, not a universal property. The archive is filtered to mid-tier brands on purpose — the saturating brands the model names near every time and the near-zero brands it almost never names were removed because they would flatten every stability number into "never moves." Part of the clean split is that population cut plus the band geometry: a rate has to sit near a line to be movable across it, and the filter removed the brands at the rails. That does not break the finding — a position property is still a position property — it bounds it. The 0.40 and 0.60 lines are our parameters, our decision lines, not natural boundaries; the split's cleanliness depends on where we drew them. The line is a choice we own.

Twenty wordings moves the average, not the worst case

The registered test says eight wordings is not enough. The smallest count that fixes the average is twenty — but even twenty does not fix the worst case: the worst cell's projected spread stays above the 0.05 line at every count we projected, up to thirty-two. Buying more wordings gets expensive fast, and even then doesn't solve the whole problem. Within a single kind of question, the eight-wording budget comes closer to enough — but that is a hypothesis, not a settled result, since three cosmetic phrasings alone cannot confirm it.

One category, one day, one design space

This is a measurement-sensitivity finding, not a causal one: different wordings move the measured rate; nothing here says why, and nothing here is a lever a brand can pull. It covers one category, one buyer profile, one archive, two engines (OpenAI and Anthropic) kept separate throughout — and one limit never closes: no one outside the model providers knows the real distribution of how buyers actually phrase these questions, so this piece cannot say how representative its eight wordings are, and does not promise to.

A reliability number is a function of the wording it was measured on

The method claims here are settled: a two-run check's blindness to wording is structural, true whenever both runs share a wording set, and instability concentrates near a decision line. None of these needs a second category. The magnitude numbers — the 0.20 average, the 0.47 worst, the 0.63 to 0.16 cell, the twenty-wording projection, the flip rates — are this category's measurement, fenced as such.

What a reader can do with this is narrow, and it is not to reword anything: wording routes the measurement, not the model, so tuning your own phrasings is the wrong move. When someone sells you an AI-visibility number, ask:

  1. How many different phrasings was it measured across? A number measured on one phrasing is a single draw from a distribution this piece shows is wide.
  2. Did their reliability check vary the phrasing, or only the day? Two runs that agree is a claim about the day unless the wordings differed.
  3. Where does the brand sit? A brand the model names nearly always or nearly never has a number wording cannot move much, while a brand in the middle has one wording can move a long way, and that is the position most brands selling into a competitive category are actually in. Whether your own number is fragile is a question about your position, not the vendor.

The program's next question is the confirmation run: one new category, the same eight wordings, the same two engines, one buyer profile, at the archive's research depth of 816 measured calls — about AUD 419, an estimate from the file's per-call economics, not a quote. What it would settle is whether the magnitude generalises — whether the split reproduces, whether the mean equivalent-rewording move is still about 0.20, whether twenty wordings is this category's answer or a property of the instrument. The method claims it would not touch; they are already settled. A reproduction lifts the magnitude out of the one-category fence; until then, the fence holds.

How we ran this

What we measured: A re-analysis of one stored archive: 8,160 call-and-brand rows across the skincare category (not chosen for this question — it was fixed by the Stability Floor study's pre-registered pilot gate, which returned PROCEED on skincare before this re-analysis was framed), eight wordings, two engines (OpenAI and Anthropic), collected in July 2026. Each wording was asked 51 times per engine across three batches; the archive holds 816 measured calls in total. The analysis population is the six mid-tier brands — twelve brand-engine cells; the saturating and near-zero brands were filtered out before the analysis on purpose, and the two invented brand-controls are excluded from every population.

When: The archive was collected in July 2026; this piece is a re-analysis and ran no new queries and incurred no new measurement spend. Deviations: none.