Does rephrasing the question change which brand AI recommends?

A re-analysis of 8,160 measurements we already held, turned on our own two-run reliability guarantee. Holding the wording still across both runs left a whole class of error untested, and the answer forces a change to the instrument.

Rephrasing the question moves which brand an AI names for about half of the twelve brand-engine cells we measured in one skincare category. Six of the cells stay put — draw any two or more of the eight wordings and they land the same way. The other six sit near a decision line, and there the answer changed on a third to over half of the seventy ways to pick four of the eight wordings. The two-run reliability check we sell could not have caught it: both runs use the same wordings, so a shared wording effect is baked into both numbers identically and cancels — the runs can agree and still be wrong the same way.

23 July 2026Results registered before data was seen

Two runs agreed — about the day, not the wording

A marketer types "best skincare brands" into ChatGPT, reads the answer, and rephrases it several different ways over a week — the way a real buyer would. The brand the assistant names moves. The question they eventually ask is whether rephrasing the question changes which brand AI recommends, and, if it does, how anyone would know.

Most AI-visibility numbers a brand buys rest on a quiet reassurance: the shop ran the measurement twice, on two different days, and the two runs agreed. Two runs agreeing is supposed to mean the number is solid. It means less than that. Two runs that agree have shown the answer is stable to when you ask — to the day. They have shown nothing about how you ask, because the wording of the question was held fixed across both runs. A whole class of error was never exercised.

The question a buyer types into ChatGPT is not one question. It is a distribution of phrasings — "best skincare brands," "top skincare brands," "skincare brands worth buying," "best skincare brands under $30" — and a number measured on one phrasing is a sample of size one from that distribution. A two-run check that reuses the same phrasing on both days tells you how stable that one sample is to the day. It tells you nothing about how wide the distribution it was drawn from is, because it never drew a second sample from it. The wording axis is the distribution; holding it still is what makes the number repeatable, and it is also what makes the repeatability silent on the one thing that moves the number most.

When you ask is the day axis: a measurement run on one day and again on another, with everything else held still, tests whether day-to-day sampling moves the answer. How you ask is the wording axis: the same intent phrased several different ways, tested against itself. The two-run check holds the second axis still and varies the first. That is a design choice — the wordings are picked once and reused — not a discovered property of the measurement, and a design choice cannot validate the thing it never varied. A reliability claim that never varies the wording is incomplete whatever the wording turns out to be worth; that much is structural and does not need this data.

The question is what the wording turns out to be worth. This piece answers it with measurements we already held. An earlier piece in this program audited our assumptions against published research and changed what we deliver when one of our own bets broke; this one turns the same posture on the instrument itself — no external sources, just the 8,160 measurements already in our archive, asked a question we had not asked of it. What breaks here is one of our guarantees. Two runs agreeing was never evidence about wording.

The one bet we placed before looking

Before looking at any outcome, we placed one bet and fixed one pass line. The bet: eight wordings is enough to average wording luck out of a brand's measured rate. The pass line sat at 0.05 on the wording-induced spread — the size of the swing the wordings induce in a brand-engine cell's measured rate. That line is a quarter of the 0.20 middle band — a swing of that size is a quarter of the distance from the 0.40 line to the 0.60 line, so a brand sitting mid-band could absorb it without crossing. A brand sitting in the middle of the range should not move across a decision line on wording luck alone. That is the motivation for the line, not a guarantee against a class change: a spread-based bar sets a ceiling on how large the wording swing can be and still pass; it does not make a class change impossible, and at this archive's current spread the drop-one flip rate is still 11.5%. We state the criterion as the proxy it is — a bar on spread size — not a proof.

The category was not chosen for this question. Skincare was fixed in advance by a different study's pilot gate: the Stability Floor study pre-registered a rule that its pilot verdict would select the headline category, that verdict came back PROCEED on skincare, and the archive this piece re-analyses was collected under that rule. The category was settled before this question was asked, by a rule written before either was answered, and we could not have chosen it to fit the result. The line and the verdict it produces are the load-bearing pre-commitment.

The interesting part of this data is exactly the part that invites story-fitting after the fact: the split between cells that move and cells that do not, the size of the spread, the ladder of what it would cost to average it out. Any of those could be reshaped to tell a cleaner story once the answer is visible. Naming the bet and the line before the analysis is what keeps that temptation honest. The metric, its population, and the pass line were all specified before any outcome was seen; the per-brand rates, the split, and the cost ladder were computed after collection with no threshold set, and are observations, not bets.

The line returned a verdict. At eight wordings the wording-induced spread came in at 0.078, above the 0.05 line. The registered verdict was insufficient — eight wordings is not enough. That is a disclosed test outcome; the honesty is the claim, and it is structural. It says our own test failed and we are saying so, not that eight is universally insufficient. What eight is insufficient for is the question the rest of the piece unpacks.

One brand read 0.63 asked one way, 0.16 asked another

Rephrasing the question moves the number. In this one skincare category, three cosmetic variants of the same intent — wordings that should be equivalent — move a brand-engine cell's measured rate by 0.20 on average and by 0.47 at worst. One cell, Cetaphil on OpenAI, read 0.63 asked one equivalent way and 0.16 asked another, across three ways of asking the same thing. That is the finding: wording is a first-order error source in this category's measurement, not a rounding detail. It is fenced to this one category, because "first-order" is a judgment about size — this archive's spread against this archive's band width — and it reads as structural when it is not.

The twelve brand-engine cells here are six mid-tier skincare brands — Cetaphil, Paula's Choice, The Ordinary, Kiehl's, Drunk Elephant, La Roche-Posay — each measured on two engines, OpenAI and Anthropic. That is the population the finding is stated on, and it is narrow on purpose: a mid-tier filter, one buyer profile, one collection window, one archive of 8,160 measurements. Every magnitude number in this section is one category's measurement; the structural claim — that a reliability number measured on one wording set is a function of that set — does not need this data and arrives unfenced.

The average equivalent-rewording move, 0.20, is as wide as the entire middle band — the gap from the 0.40 line to the 0.60 line. The worst cell's move, 0.47, is more than double that gap. A measured rate is the share of the 51 asks that named a given brand, and an equivalent rewording moving that share by a band-width is not a rounding detail; it is enough to carry a brand from one side of a line to the other. Drop one of the eight wordings and a cell's classification changes 11.5% of the time; at five wordings, 19%; at one wording, 30%. That is measured on the archive, not projected from it. The middle band — the region between our 0.40 and 0.60 decision lines — is only 0.20 wide, so a cell whose rate sits near a line is movable across it by wording, and a cell whose rate sits far from a line is not. Where the movable cells sit, and whether that is a property of the brand or of the position, is the next section's argument.

A rival reading has to be named at the point where it would explain the result away: maybe the 0.20 average is just noise from too few repeats of each wording. It is not ruled out by a model; it is ruled out by one registered fact. Each rate rests on 51 asks of that exact wording per engine, across three batches — not a handful. That is what the registered test was built on, and it is where the answer to the noise rival stops; this piece does not add a second one.

The wordings that move a rate most are not the equivalent ones. The three cosmetic variants move a rate by 0.20 on average; all eight wordings together move it by 0.585 on average, and the equivalent-rewording spread is just over a third of the all-eight spread. The wordings that narrow the question — a price ceiling, a channel constraint — differ more than the equivalent rewordings do, because they change what the question asks for. That is a description of the measured pattern, not a mechanism. This data contains no analysis of how a wording routes through the model's retrieval to the brand it names; if a reader asks why the wordings move the rate, the honest answer is that the question is not tested here. What is tested here is that they do.

Six cells the model is sure about, six it isn't

Where you sit decides whether wording moves you

Each dot is one brand-engine cell. Across, its measured rate; up, how often the result changes across the four-of-eight wording subsets. The wording-proof cells sit flat on zero and spread across the scale; the wording-decided cells cluster between the two decision lines.

0.00.20.40.60.81.00%20%40%60%0.40 line0.60 linemeasured rate (share of asks that named the brand)how often the result changeswording-decided · between the lineswording-proof · flat on zero

The twelve cells do not sit on a sliding scale of wording-sensitivity. They split clean, and the split is the structural claim of this piece. Six cells land on the same result for every subset of two or more of the eight wordings — zero flips, at every subset size from two to eight. The other six change result on a third to over half of the 70 four-of-eight subsets — 33% to 57% across them. "A third to over half" is the four-of-eight figure and it travels with that denominator: it is not a constant across subset sizes, and at seven-of-eight one of those same six cells reads zero. The split is not a gradient; it is two clusters with little between.

A load-bearing exception: the one-wording case is the single outlier. Two of the six proof cells — Kiehl's on Anthropic and La Roche-Posay on OpenAI — flip on one of their eight one-wording subsets and on zero at every size from two to eight. The other four proof cells are zero at every size including one. "Zero flips at every subset size from two to eight" is the verified claim; it is not "every possible subset," because the one-wording case falsifies that. Being wording-proof is not a virtue of the brand or of its measurement — it is a position property.

The six wording-proof cells sit far from any decision line — their full-eight rates run from 0.03 to 0.94, and they cluster at the rails, not in the middle. La Roche-Posay sits near certainty on both engines, at 0.919 on Anthropic and 0.944 on OpenAI. Kiehl's and Drunk Elephant sit near zero — Kiehl's on OpenAI at 0.029, Drunk Elephant on OpenAI at 0.091. None of the six is anywhere near a line, so a wording-induced move has nowhere to push them across.

The six wording-decided cells sit near a line — just under or just over the 0.40 and 0.60 boundaries. Cetaphil on OpenAI reads 0.355, just below 0.40; Cetaphil on Anthropic reads 0.417, just above it. Paula's Choice reads 0.404 on Anthropic and 0.461 on OpenAI, both between the lines. The Ordinary reads 0.630 on OpenAI and 0.640 on Anthropic, both just above 0.60. Every one of the six is within a wording move of a boundary, and the flip rates show it: at four of the eight wordings these six cells change result between 33% and 57% of the 70 half-sets — Paula's Choice on OpenAI flips on 0.571 of them, Cetaphil on Anthropic on 0.529, The Ordinary on OpenAI on 0.400, The Ordinary on Anthropic on 0.329. Whether a wording move changes the verdict is decided entirely by whether the rate started near a line: a cell out at the extremes has nowhere to be pushed across, while a cell near a line has a boundary within reach.

The proof cells are not better-measured or more stable in any virtuous sense — they are the ones the model is sure about either way, so no wording can push them across. The decided cells are the ones the model is ambivalent about, so every wording can. That claim is a structural form — instability concentrates near a decision line — and it publishes without a category fence. What is fenced to this category is the shape of the split: that it is clean, and roughly half-and-half, is this archive's configuration, not a universal property of AI measurement.

The archive is filtered to mid-tier brands on purpose. The saturating brands — the ones the model names near every time — and the near-zero brands — the ones it almost never names — were removed before this analysis because they would flatten every stability number into "never moves." Part of what looks like a clean proof-versus-decided split is that population cut plus the band geometry: a rate has to sit near the 0.40 or 0.60 line to be movable across it, and the filter removed the brands that sit at the rails. This does not break the finding — a position property is still a position property — it bounds it. The clean split is what you see once the population has been narrowed to the brands the instrument is ambivalent about.

The 0.40 and 0.60 lines are our parameters. They are our decision lines, like the price and the cost ceiling, not natural boundaries of the scale. The split's cleanliness depends on where we drew them; a different pair of lines would move cells between the clusters. The line is a choice we own.

A check that holds the wording still cannot see the wording

We sell a two-run check. A failure mode is named for a brand only when two day-separated runs agree on it. It rests on a premise worth stating outright, because everything that follows depends on it: both runs use the same wordings, by design. The wording set is picked once and reused across both runs. That is not an accident we missed; it is how the check was built. For a stable wording bias — the kind induced by a shared wording set, common to both runs and cancelling in the difference between them — that bias can never show up as the two runs disagreeing. A constant does not appear as a difference.

The two-run check is therefore blind to detecting, attributing, and bounding the wording design's contribution. It is not blind to every disagreement near a line. Day-to-day noise can still flip a brand sitting near a decision line in one run and not the other, and there can be day-by-wording interaction; when the two runs do disagree near a line, that disagreement neither measures nor bounds the wording error — it says the day moved it. The precise claim is the blindness to the wording design's contribution, not the absence of any disagreement.

We built this re-analysis to ask whether eight wordings is enough, and the same data that answered "no" also showed us that our existing two-run guarantee could not have answered it. The two-run check varies the day and holds the wording still, so a wording effect that is stable across runs is a constant that cancels in any run-versus-run comparison. The guarantee was built to catch day drift, and it does; it was never built to catch wording drift, and it does not. The blindness is structural — true whenever both runs share a wording set, independent of this category — and it is the finding that forces the fix.

What the two runs agreeing certified, then, is narrow. It certified that the day did not move the brand across a line between run 1 and run 2. It did not certify that the wording set was innocent — the wording set was the same on both days, so its effect was baked into both numbers identically and subtracted out. A brand the check called stable could have been sitting on a wording-induced ledge the whole time: stable between the two days, because the ledge was there on both, and movable the moment a buyer phrased the question differently. The check's agreement and the wording's effect are orthogonal here, and that is not a defect of the check — it is what holding the wording still was for.

Any brand the two-run check ever called stable on wording-decided ground was stable to the day, not to the wording. We would have named a failure mode for it without the wording axis ever being exercised — the diagnosis would have read "reliable," and the word would have meant "reliable to the day," nothing wider. A reliability guarantee a buyer is paying for has a hole exactly the shape of the thing it claims to guard. This is not a limitation appended to the finding; it is the finding, and it is what the next two sections answer.

Twenty wordings moves the average, not the worst case

The registered test says eight wordings is not enough. The smallest number at which the average spread meets the pass line is twenty — but twenty is a projection that assumes every wording is interchangeable, not an observed result. The analysis treats wordings as interchangeable draws and extrapolates; the archive holds eight wordings, not twenty. The projection is optimistic in one direction: it treats the eight wordings as draws from one pool, when the measured pattern shows they are not — the narrower wordings move a rate more than the equivalent ones, and a projection that assumes they are interchangeable smooths that unevenness away. And no projected number of wordings clears the line for every brand. The worst cell's projected spread is 0.144 at eight wordings, 0.091 at twenty, and 0.072 at thirty-two — still above the 0.05 line at every projected count. Twenty fixes the average; it does not fix the worst case.

Buying twenty wordings would cost. A product snapshot at eight wordings runs about 144 measured calls and costs about AUD 74 — roughly a tenth of the AUD 750 price. At twelve wordings the share is 14.8%, just under our cost line; at sixteen it is 19.7%, over; at twenty the call count rises to 360 and the share is 24.7%, a quarter — about AUD 185. The cost line is 15%. These are our numbers — our price, our cost line, a parameter we set, not a law of physics.

The point of saying so is to keep a reader from reading "twenty wordings needed" as a measured fact about the world. The measured fact is the spread; the ceiling is a choice we own and could change. Holding both is the honest framing: the spread is what the data shows, and the line we will not cross to average it out is what we set.

The same data opens a different door. Within a single kind of question — the three cosmetic variants that should be equivalent — the eight-wording budget spent inside that one kind is marginally adequate: the within-kind spread at eight wordings is 0.042, just under the 0.05 line. That margin is thin enough that calling eight "enough" within one kind already overstates it, and the figure is exploratory, not pre-registered. It rests on three cosmetic phrasings, and the results write-up itself calls it indicative, not precise. The pre-registered verdict is the pooled one — eight is insufficient on the pooled test — and within-kind adequacy is a hypothesis the fix leans on, not a result it rests on. Whether eight is enough within a single kind of question is not something three cosmetic phrasings can settle; the pooled verdict stands, and the pooled-versus-within-kind distinction is the lever the fix turns on.

Where wording decides, ship the variation as the finding

The direction under review is to design the instrument so that a failure mode is named for a brand only when wording cannot move it across the line — and where wording does decide the answer, the wording-sensitivity itself becomes the finding the buyer receives, with the per-wording evidence shown. This is in-flight direction, not a landed change. The mechanics are not settled yet, so this piece describes the direction the fix takes, not how the gate is built. No resampling procedure, no threshold value, no gate arithmetic, and no timeline belong here, because none is settled. The fix belongs inside a single run, on the wording axis, from data every run already collects.

The gate asks one question: would an equally valid set of wordings have given the same label? When the answer is yes, the label stands. When the answer is no, the instrument says nothing instead of guessing — it does not issue a diagnosis it cannot support. That is the wording axis the two-run check never exercised, now exercised within one run.

The rejected alternative is to vary the wordings between run 1 and run 2 — to convert the hidden error into a visible disagreement. It does not work. Varying wordings between the runs would convert hidden wording error into an unsellable "inconclusive" for nearly half of this category's mid-tier brand-engine cells — six of twelve, about half — which is hiding the problem behind a shrug, not fixing it. The buyer would pay for a diagnosis and receive "we cannot say," on the brands where the question matters most, because the two runs were made to disagree on purpose. The fix belongs inside the run, where the instrument can say what its own evidence supports, not between the runs, where it can only refuse to speak.

What the buyer would see is the difference. Where wording cannot move a brand's label, they get the diagnosis — the same label the two-run check would have issued, now earned on the wording axis as well as the day axis. Where wording can, they get the variation as evidence — the per-wording rates and the spread — not a shrug, and not a single number pretending the wording axis was never there. The product's claim tightens because of its own audit: a measurement shop that finds a hole names it, narrows its own assertions, and ships the hole as the result. Softening the gate to protect margin is ruled out, and already-delivered diagnoses are not re-computed or withdrawn; the limitation is disclosed going forward, not retro-applied.

One category, one day, one design space

One category, one buyer profile, one collection window, one archive. The finding is a measurement-sensitivity finding, not a causal one: different wordings produce different rates. It is not a claim that wording causes the model to recommend differently, and not a lever a brand can pull on a consumer surface. Wording routes the measurement; it does not route the model. The pooled-versus-within-kind distinction is in the cost section; the two-run blindness is structural and in the self-audit; the split's one-category scope and the population cut are in the split section. Each limit is argued where it bites and only indexed here.

The archive carries no buyer-profile dimension at all, while a product snapshot runs three — the finding does not cover buyer profiles, and the piece does not imply it does. The two engines are not collapsed: the twelve cells are six brands seen through two different AI systems, OpenAI and Anthropic, and the same brand can behave differently on each. The title says "AI" as reader language; the body reports OpenAI and Anthropic as separate cells throughout.

One limit is permanent rather than pending. We cannot know whether our wording design space resembles how buyers actually address an AI assistant. Only the model providers hold real prompt distributions, and they do not release them; survey data is too thin to substitute. An owner decision in July 2026 makes this a disclosed limit of the instrument, not an open research task. The piece says so and does not promise to close it.

A reliability number is a function of the wording it was measured on

A brand's AI-visibility number is a function of the wording it was measured on. The piece has traced the chain: a brand's position near a decision line decides whether wording can move it, and for about half of the twelve brand-engine cells in this one category it can; the two-run reliability check we sell holds the wording still across both runs, so the wording effect is a constant that cancels and the check cannot detect, attribute, or bound it; the fix runs inside a single run on the wording axis, where the instrument says what its own evidence supports and stays silent where it cannot. The only honest instrument is one that says so where it matters and stays silent where it cannot. The method claims are settled. The two-run check's blindness to wording is structural and true whenever both runs share a wording set; a reliability claim that never varies the wording is incomplete whatever the effect is worth; our registered test failed and we are saying so; instability concentrates near a decision line. None of these needs a second category.

The magnitude numbers — the 0.20 average, the 0.47 worst, the 0.63 to 0.16 cell, the twenty-wording projection, the flip rates — are this category's measurement, fenced as such, and they are in the piece because a reader has to see "0.63 to 0.16" to feel how much wording moved things here. What a reader can do with this is narrow, and it is not to reword anything: wording routes the measurement, not the model, so tuning your own phrasings is the wrong move. When someone sells you an AI-visibility number, ask how many different phrasings it was measured across — a number measured on one phrasing is a single draw from a distribution this piece shows is wide. Ask whether their reliability check varied the phrasing or only the day; two runs that agree is a claim about the day unless the wordings differed. And ask where the brand sits — a brand the model names nearly always or nearly never has a number wording cannot move much, while a brand in the middle has one wording can move a long way, and that is the position most brands selling into a competitive category are actually in. Whether your own number is fragile is a question about your position, not about the vendor.

The question the program asks next is the confirmation run. One new category, the same eight wordings, the same two engines, one buyer profile, at the archive's research depth of 816 measured calls — about AUD 419, an estimate derived from the file's per-call economics, not a quote. What it would settle is whether the magnitude generalises: whether the split reproduces, whether the mean equivalent-rewording move is still about 0.20, whether twenty wordings is this category's answer or a property of the instrument. The method claims it would not touch; they are already settled. A reproduction lifts the magnitude out of the one-category fence; until then, the fence holds.

How we ran this

What we measured
A re-analysis of one stored archive: 8,160 call-and-brand rows across the skincare category (not chosen for this question — it was fixed by the Stability Floor study's pre-registered pilot gate, which returned PROCEED on skincare before this re-analysis was framed), eight wordings, two engines (OpenAI and Anthropic), collected in July 2026. Each wording was asked 51 times per engine across three batches; the archive holds 816 measured calls in total. The analysis population is the six mid-tier brands — twelve brand-engine cells; the saturating and near-zero brands were filtered out before the analysis on purpose, and the two invented brand-controls are excluded from every population.
When
The archive was collected in July 2026; this piece is a re-analysis and ran no new queries and incurred no new measurement spend. Deviations: none.
Which AI
OpenAI and Anthropic, via their APIs with live web search, no personalization, memory, or geo., The title says "AI" as reader language; the body reports OpenAI and Anthropic as separate cells throughout.

The full numbers

Every number in this piece, at full precision. The prose rounds for reading; this table doesn't.

MeasureAs shownExactSource
How far equivalent rewordings move a brand's rate, on average0.200.204248Derived
How far equivalent rewordings move a brand's rate, worst case0.470.470588Derived
The brand's rate, asked one equivalent way0.630.627451Derived
The same brand's rate, asked another equivalent way0.160.156863Derived
How far all eight wordings move a brand's rate, on average0.5850.584967Derived
Average rewording spread as a fraction of the all-eight-wordings spread35%34.9162%Derived
Brand-engine cells in the headlinetwelve12Measured
Ways of askingeight8Measured
Possible four-of-eight subsets7070Derived
Brand-engine cells that never flip, of twelvesix6Derived
Brand-engine cells that flip, of twelvesix6Derived
Lowest flip rate among the decided six33%32.8571%Derived
Highest flip rate among the decided six57%57.1429%Derived
Cetaphil, Anthropic, flip rate at four wordings0.5290.528571Measured
Paula's Choice, Anthropic, flip rate at four wordings0.5290.528571Measured
The Ordinary, Anthropic, flip rate at four wordings0.3290.328571Measured
Cetaphil, OpenAI, flip rate at four wordings0.4140.414286Measured
Paula's Choice, OpenAI, flip rate at four wordings0.5710.571429Measured
The Ordinary, OpenAI, flip rate at four wordings0.4000.400000Measured
Lowest full-eight rate among the wording-proof six0.030.029412Derived
Highest full-eight rate among the wording-proof six0.940.943627Derived
Drunk Elephant, Anthropic, full-eight rate (proof)0.0390.039216Measured
Kiehl's, Anthropic, full-eight rate (proof)0.1000.100490Measured
La Roche-Posay, Anthropic, full-eight rate (proof)0.9190.919118Measured
Drunk Elephant, OpenAI, full-eight rate (proof)0.0910.090686Measured
Kiehl's, OpenAI, full-eight rate (proof)0.0290.029412Measured
La Roche-Posay, OpenAI, full-eight rate (proof)0.9440.943627Measured
Cetaphil, Anthropic, full-eight rate (decided)0.4170.416667Measured
Paula's Choice, Anthropic, full-eight rate (decided)0.4040.404412Measured
The Ordinary, Anthropic, full-eight rate (decided)0.6400.639706Measured
Cetaphil, OpenAI, full-eight rate (decided)0.3550.355392Measured
Paula's Choice, OpenAI, full-eight rate (decided)0.4610.460784Measured
The Ordinary, OpenAI, full-eight rate (decided)0.6300.629902Measured
How often dropping to one wording changes a brand's result30%30.2083%Measured
…at five wordings19%18.8988%Measured
…at seven wordings (drop one of eight)11.5%11.4583%Measured
Share of one-wording subsets on which the two exception cells change result12.5%12.5000%Derived
Wording-induced spread of the estimate at eight wordings0.0780.077735Measured
Lower bound of the spread estimate0.0610.060707Measured
Upper bound of the spread estimate0.0810.081339Measured
The pass line we pre-registered0.050.050000Measured
Registered verdictinsufficientinsufficientDerived
Spread within a single kind of question, at eight wordings0.0420.042028Derived
Smallest number of wordings at which the average spread meets the criterion2020Derived
Smallest number of wordings at which every cell meets the criterionnonenullDerived
Worst cell's projected spread at eight wordings0.1440.144239Derived
Worst cell's projected spread at twenty wordings0.0910.091224Derived
Worst cell's projected spread at thirty-two wordings0.0720.072119Derived
Wording-induced spread of the estimate at twenty wordings0.0490.049164Measured
Lower decision line0.400.400000Measured
Upper decision line0.600.600000Measured
Width of the middle band0.200.200000Derived
The snapshot price the cost shares are againstAUD 750750Measured
Our self-imposed cost ceiling15%15.0000%Measured
Measured calls per run at eight wordings144144Derived
Run cost at eight wordings~AUD 7474Measured
Share of the AUD 750 price at eight wordings9.9%9.8667%Derived
Share at twelve wordings14.8%14.8000%Derived
Share at sixteen wordings19.7%19.7333%Derived
Measured calls per run at twenty wordings360360Derived
Run cost at twenty wordings~AUD 185185Derived
Share at twenty wordings24.7%24.6667%Derived
Cost per measured call (derived estimate)~AUD 0.51 per call0.513889Derived
Confirmation run, full research depth — measured calls816816Derived
Confirmation run, full depth — cost (estimate)~AUD 419419.3333Derived
Confirmation run, one-third depth — measured calls272272Derived
Confirmation run, one-third depth — cost (estimate)~AUD 140139.7778Derived
Confirmation run, product depth — measured calls4848Derived
Confirmation run, product depth — cost (estimate)~AUD 2524.6667Derived
Call-and-brand rows re-analyzed8,1608,160Measured
Times each wording was asked, per engine (3 batches of 17)5151Derived
Measured calls behind the archive816816Derived

What this doesn't settle

  • Whether the magnitude generalises beyond this one skincare category — the mean spread, the split, the twenty-wording projection — is unsettled until a confirmation run reproduces it in a second category.
  • Whether our wording design space resembles how real buyers address an AI assistant. Only the model providers hold real prompt distributions, and they do not release them; this is a disclosed permanent limit of the instrument, not an open research task (owner decision, July 2026).
  • Whether a within-kind wording budget is adequate. The within-kind spread of 0.042 is exploratory, rests on three cosmetic phrasings, and is not the pre-registered verdict; it is a hypothesis, not a result.
  • Whether wording routes through the model's retrieval to the brand it names. This data contains no such analysis; the piece states that wordings move rates and does not claim a mechanism.

Was your AI-visibility number measured on one wording or many?

A diagnosis runs your questions across several wordings and shows whether wording can move your brand's result. Not a score: the evidence.

See what a diagnosis finds