If you ask AI again two weeks later, does the answer change?
We locked in a bet that repeat AI answers would show more spread than pure chance, before collecting any data. They didn't — and two weeks later, most brand-recommendation rates were indistinguishable from that same chance-sized band, under a test that could only catch a swing bigger than about 30 points.
Asking again two weeks later mostly landed back inside the same wide swing that asking the identical question three times on one single day already produces — true for 38 of 42 tracked combinations — and the test could only have caught a swing bigger than about 30 percentage points to begin with.
A brand watches an AI-recommendation number move from one week to the next. The move reads as a signal. The model likes the brand less, a competitor took the spot, something is wrong.
That reading assumes a change between two readings reflects a real change worth acting on. The same swing shows up between two readings of the identical question, on the identical day, before any time passes. The rest of this piece measures that swing directly.
Two refuted, one bounded
Three bets were locked in before any data came in, published regardless of how each one landed.
- Coverage — refuted. Ranges built from 17-call batches covered the fuller same-day rate 96% of the time — the check did not detect anything extra beyond plain chance.
- Extra clustering — refuted. A second, more direct same-day test also did not detect clustering beyond plain chance.
- Bounded drift — confirmed. 38 of 42 combinations were indistinguishable from that same noise two weeks later — bounded, because at 50% odds this design could only have caught a change bigger than about 30 points.
96 percent, not the 90 we bet against
All 126 interval checks across the 42 tracked combinations — which ones covered the pooled rate, and which five didn't.
Before the first dispatch, one bet was locked in. The bet said repeat answers would show more spread than plain chance produces. Ranges built from 17 calls would miss the fuller same-day rate, built from 51 calls, more than 10% of the time. That bet was refuted. Those ranges covered the fuller rate 96.0% of the time, across 126 checks. Plain chance explains the same-day spread. The check did not detect anything extra beyond it.
98.47 against a chance ceiling of 112.69
The observed same-day spread against the simulated chance-only ceiling.
A second same-day test asks the question more directly. Does the spread between three same-day repeat sets exceed what plain chance alone produces? We simulated what the spread would look like if every repeat were pure chance at the measured rate, 2,000 times, and compared the real spread to that simulated range. The chance ceiling sat at 112.69 — higher than all but the top 2.5% of those 2,000 simulated chance-only outcomes. The real spread measured 98.47. The test did not detect extra clustering beyond ordinary variation from chance (p=.159).
38 of 42 stayed inside that same noise, two weeks later
How many day-0-to-day-14 comparisons landed inside same-day noise, out of 42 tracked.
The main comparison runs day 0 against day 14. Day 0 pooled 51 calls per brand-question-provider combination across three same-day sets; day 14 used 17 calls. Of 42 combinations, 38 (90.5%) landed inside the same range that same-day chance alone produces. This comparison could only have caught a change of about 30pp at 50% odds of catching it, 40pp at 80% odds. Smaller changes had lower odds of being caught either way.
Four combinations moved beyond that same-day range, all four on the Anthropic provider. The Ordinary, one phrasing: 25.5pp (p=.028). Kiehl's, one phrasing: 35.3pp (p=.015). The Ordinary, a second phrasing: 33.3pp (p=.023). Cetaphil, one phrasing: 25.5pp (p=.043).
If all 42 comparisons were truly unchanged, about 2.1 would cross this threshold by chance alone. Four were observed. No correction was applied because none was pre-registered. These four are reported as-is, not classified as real change or as noise.
Same-day chance already swings twenty points
The typical range a repeat measurement falls in, from chance alone, as the number of calls grows — median across the 42 tracked combinations, and the single widest one.
The bands above are wide because chance alone swings a lot at these call counts. At 51 calls the typical band is 12.2pp. Drop to 34 calls and it widens to 14.8pp. At 24 calls it is 17.4pp. At 17 calls — one full set — it is 19.9pp. At 12 calls it is 23.6pp. At 8 calls it is 26.0pp. At 3 calls it is 36.5pp. A band is the range a repeated measurement typically falls in from chance alone, at that many calls.
How often each pair's order flips at 24 calls, against how far apart their rates sit.
The same chance drives the order of two brands. Cetaphil and Paula's Choice on Anthropic rated 41.7% against 40.4%, a 1.2pp gap, nearly tied; their order flips 46.7% of the time, close to a coin toss. Paula's Choice against Cetaphil on OpenAI rated 46.1% against 35.5%, a 10.5pp gap, and the order flips 21.7% of the time. The Ordinary against Paula's Choice on OpenAI rated 63.0% against 46.1%, a 16.9pp gap, and flips 11.4% of the time. The Ordinary against Cetaphil on Anthropic rated 64.0% against 41.7%, a 22.3pp gap, and flips 4.8% of the time. The closer two rates sit, the more often their order flips from which 24 calls you happened to draw. A wide-enough gap holds its order almost every time.
The three bets above were locked in before any data was collected. The half-width curve and the rank-flip numbers here were computed after, from the same measured rates, as description rather than a pre-registered test.
Ten skincare brands, two APIs, fourteen days
This study measured two provider APIs, OpenAI and Anthropic, both with live web search, not a consumer chat app. The APIs carry no personalization, memory, or location. Consumer surfaces may differ; this is a controlled proxy.
One category was measured, skincare, with ten configured brands: two familiar reference brands, six candidates, and two made-up names as a control. Each provider ran 544 calls across the whole study — eight ways of asking, seventeen repeats, four separate batches (three on day 0, one on day 14).
The made-up names never appeared. Across 544 calls per provider, a made-up brand was named 0 times on both providers. That shows the measurement is not inventing brand names wholesale. It says nothing about accuracy for real brands.
This design has only even odds of catching a change smaller than about 30pp, and roughly 80% odds at 40pp. That detection floor applies to every day-0-to-day-14 comparison in this piece, not a caveat on one.
Day 7, a third planned checkpoint, never ran, and day 0 and day 14 were measured at different times of day — both disclosed in full below, alongside what each costs the comparison above.
The design cannot attribute any detected or non-detected change to a specific cause — not a provider update, not a change in which pages get surfaced, not anything else. This design cannot tell those apart.
This study does not test how much the wording of the question matters, whether a persona changes the answer, or whether providers agree with each other. Those are separate studies.
A single weekly reading mostly counts how many times you asked
A single AI-recommendation reading, taken once, carries the wide band from above with it. Two single readings, taken close together and each built from a handful of calls, differ by an amount this study cannot separate from ordinary same-day noise. Only a move past roughly 30pp clears that noise, and even that cannot be pinned to a specific cause.
The same 30-point floor, read three ways
For an SEO or a marketer watching a weekly AI-visibility dashboard, the practical read is the one above: a number moving from one week to the next, by less than roughly 30 points and built from a handful of calls, is not, on its own, evidence that anything changed.
For someone running or selling Generative Engine Optimization (GEO) work, the same logic applies to a before-and-after claim. A tracked recommendation rate that moves after a content change or any other tactic has to clear the detection floor of the measurement making the claim — under a design like this one, about 30 percentage points. Below that floor, the before-and-after numbers are not distinguishable from two draws of the same noisy process, and are not proof the tactic did anything.
For an AI researcher, the two refuted bets are the more interesting result: neither the coverage check nor the second same-day test detected anything beyond plain chance in same-day repeats, at this scale, on real brand questions. Same-day batches of 17 calls already carry a typical band of roughly 20 points; the day-0-to-day-14 comparison, separately, could only have caught drift above roughly 30 points. Another study measuring AI-visibility stability could use both figures when sizing its own sample, within the limits of one category and two providers.
Two APIs, eight phrasings, three bets locked first
The measurement ran on two provider APIs: OpenAI gpt-5.4 and Anthropic claude-sonnet-5. The category was skincare, with ten configured brands, eight phrasings, and one persona. Day 0 ran three same-day batches of 17 calls per combination, pooled to 51. Day 14 ran one batch of 17 calls. Three bets were registered before the first dispatch: the two same-day chance-comparison bets and the day-0-to-day-14 bounded-drift bet. All three were set to publish regardless of outcome, by design. The bets, the thresholds, and the analysis code were locked in before the first dispatch.
Five things ran differently from the plan.
First, on day 0 an account billing outage delayed the OpenAI set by about 40 minutes relative to the Anthropic set. The set was funded and completed inside the same window. The cost is a same-day timing asymmetry within that first set. No correction was applied, because the set still completed inside its registered window.
Second, the original fixed measurement window — 10:00 to 14:00 AEST each day — was dropped partway through the study, after day 0 had already run inside it. Day 14 ran at 06:50 to 07:42 AEST, well outside it. The cost is that day 0 and day 14 differ in time of day, so a difference, or its absence, between them cannot be cleanly separated from that. No correction was applied; none was pre-registered.
Third, day 7 never ran. Its window passed with no dispatch, and the decision was made to skip it rather than run it late. The cost is that the drift comparison covers day 0 to day 14 only, not the fuller three-point design. No correction was applied; the missing calls cannot be reconstructed.
Fourth, day 14 dispatched about 16 hours later than scheduled within its allowed window, after two automated provider health checks flagged and were confirmed as false alarms. The planned confirmation re-checks ran early, 17 to 19 hours apart instead of 24, by a direct decision. The cost is none to the measured content; this is an operations note, not a data change. No correction was needed for that reason.
Fifth, the figures in this piece were built after the numbers above were final. The frozen analysis step that produced them emits numbers only, no charts. The cost is none; the figures present the same locked numbers, built afterward. No correction was needed for that reason.
Earlier published research made the same point in general. Sielinski's "Quantifying Uncertainty in AI Visibility" (arXiv:2603.08924) argued that a single-run reading is a sample from a noisy process, not a fixed truth. This study tests that question directly, on real brand questions, with a bet locked in before looking.
How we ran this
- What we measured
- 42 tracked brand-question-provider combinations, out of the full set, chosen because their day-0 rate wasn't already stuck near 0% or 100%. Ten configured skincare brands (two familiar reference brands, six candidates, two made-up controls), eight phrasings, two provider APIs (OpenAI, Anthropic). Day 0: three same-day batches of 17 calls per combination, pooled to 51. Day 14: one batch, 17 calls.
- When
- Day 0: 20 July 2026. Day 14: 3 August 2026 (a planned day 7 was skipped — see the piece).
- Which AI
- OpenAI — gpt-5.4, via API, live search, Anthropic — claude-sonnet-5, via API, live search
The full numbers
Every number in this piece, at full precision. The prose rounds for reading; this table doesn't.
| Measure | As shown | Exact | Source |
|---|---|---|---|
| Same-day intervals that covered the pooled rate | 96.0% | 96.0317% | Measured |
| Same-day interval checks (coverage test) | 126 | 126 | Derived |
| Day-0-to-day-14 pairs indistinguishable from same-day noise | 90.5% | 90.4762% | Measured |
| Pairs indistinguishable from same-day noise, of 42 | 38 | 38 | Measured |
| Day-0-to-day-14 pairs tested | 42 | 42 | Measured |
| Smallest two-week change detectable at 50% odds | 30pp | 30.0000% | Measured |
| Smallest two-week change detectable at 80% odds | 40pp | 40.0000% | Measured |
| Calls per cell, day 0 (pooled) | 51 | 51 | Measured |
| Calls per cell, day 14 | 17 | 17 | Measured |
| Pairs that moved beyond same-day noise | 4 | 4 | Measured |
| Pairs expected to move by chance alone, of 42 | 2.1 | 2.1 | Derived |
| The Ordinary, one phrasing — day-0/day-14 change | 25.5pp | 25.4902% | Measured |
| The Ordinary, one phrasing — test result | p=.028 | 0.0284 | Measured |
| Kiehl's, one phrasing — day-0/day-14 change | 35.3pp | 35.2941% | Measured |
| Kiehl's, one phrasing — test result | p=.015 | 0.0146 | Measured |
| The Ordinary, a second phrasing — day-0/day-14 change | 33.3pp | 33.3333% | Measured |
| The Ordinary, a second phrasing — test result | p=.023 | 0.0228 | Measured |
| Cetaphil, one phrasing — day-0/day-14 change | 25.5pp | 25.4902% | Measured |
| Cetaphil, one phrasing — test result | p=.043 | 0.0425 | Measured |
| Same-day batch-to-batch spread, measured | 98.47 | 98.469367 | Measured |
| Same-day spread expected from chance alone (97.5th pct) | 112.69 | 112.691921 | Measured |
| Test result, same-day spread vs. chance | p=.159 | 0.1585 | Measured |
| Typical band width at 3 calls | 36.5pp | 36.5424% | Measured |
| Typical band width at 8 calls | 26.0pp | 25.9623% | Measured |
| Typical band width at 12 calls | 23.6pp | 23.5629% | Measured |
| Typical band width at 17 calls | 19.9pp | 19.9266% | Measured |
| Typical band width at 24 calls | 17.4pp | 17.3936% | Measured |
| Typical band width at 34 calls | 14.8pp | 14.8399% | Measured |
| Typical band width at 51 calls | 12.2pp | 12.2474% | Measured |
| Cetaphil vs. Paula's Choice, Anthropic — higher rate | 41.7% | 41.6667% | Measured |
| Cetaphil vs. Paula's Choice, Anthropic — lower rate | 40.4% | 40.4412% | Measured |
| Cetaphil vs. Paula's Choice, Anthropic — order flips | 46.7% | 46.6750% | Measured |
| Cetaphil vs. Paula's Choice, Anthropic — rate gap | 1.2pp | 1.2255% | Derived |
| The Ordinary vs. Cetaphil, Anthropic — higher rate | 64.0% | 63.9706% | Measured |
| The Ordinary vs. Cetaphil, Anthropic — lower rate | 41.7% | 41.6667% | Measured |
| The Ordinary vs. Cetaphil, Anthropic — order flips | 4.8% | 4.7750% | Measured |
| The Ordinary vs. Cetaphil, Anthropic — rate gap | 22.3pp | 22.3039% | Derived |
| The Ordinary vs. Paula's Choice, OpenAI — higher rate | 63.0% | 62.9902% | Measured |
| The Ordinary vs. Paula's Choice, OpenAI — lower rate | 46.1% | 46.0784% | Measured |
| The Ordinary vs. Paula's Choice, OpenAI — order flips | 11.4% | 11.3750% | Measured |
| The Ordinary vs. Paula's Choice, OpenAI — rate gap | 16.9pp | 16.9118% | Derived |
| Paula's Choice vs. Cetaphil, OpenAI — higher rate | 46.1% | 46.0784% | Measured |
| Paula's Choice vs. Cetaphil, OpenAI — lower rate | 35.5% | 35.5392% | Measured |
| Paula's Choice vs. Cetaphil, OpenAI — order flips | 21.7% | 21.6750% | Measured |
| Paula's Choice vs. Cetaphil, OpenAI — rate gap | 10.5pp | 10.5392% | Derived |
| Calls per provider testing the two made-up brands, whole study | 544 | 544 | Measured |
| Times a made-up brand was named | 0 | 0 | Measured |
| Brand-question-provider combinations analyzed | 42 | 42 | Measured |
| Total measurement spend, this study | US$92.48 | 92.48 | Measured |
| Pre-approved spend ceiling | US$160 | 160.00 | Measured |
What this doesn't settle
- Provider APIs with live search only — not a consumer chat product.
- One category, ten configured brands, two providers, this measurement window — does not generalize beyond it.
- Cannot detect a change smaller than about 30 percentage points at this depth.
- Day 0 and day 14 differ in time of day; the two cannot be cleanly separated.
- Cannot attribute any change, or the absence of one, to a specific cause.
References
How many times would you need to ask, to trust the number?
A diagnosis runs each question multiple times, across two day-separated runs, and shows the band behind every rate — not a single reading dressed up as a fact.