If you ask AI again two weeks later, does the answer change?

We locked in a bet that repeat AI answers would show more spread than pure chance, before collecting any data. They didn't — and two weeks later, most brand-recommendation rates were indistinguishable from that same chance-sized band, under a test that could only catch a swing bigger than about 30 points.

Asking again two weeks later mostly landed back inside the same wide swing that asking the identical question three times on one single day already produces — true for 38 of 42 tracked combinations — and the test could only have caught a swing bigger than about 30 percentage points to begin with.

3 August 2026Results registered before data was seen

A brand watches an AI-recommendation number move from one week to the next. The move reads as a signal. The model likes the brand less, a competitor took the spot, something is wrong.

That reading assumes a change between two readings reflects a real change worth acting on. The same swing shows up between two readings of the identical question, on the identical day, before any time passes. The rest of this piece measures that swing directly.

Two refuted, one bounded

Three bets were locked in before any data came in, published regardless of how each one landed.

  1. Coverage — refuted. Ranges built from 17-call batches covered the fuller same-day rate 96% of the time — the check did not detect anything extra beyond plain chance.
  2. Extra clustering — refuted. A second, more direct same-day test also did not detect clustering beyond plain chance.
  3. Bounded drift — confirmed. 38 of 42 combinations were indistinguishable from that same noise two weeks later — bounded, because at 50% odds this design could only have caught a change bigger than about 30 points.

96 percent, not the 90 we bet against

Every same-day check

All 126 interval checks across the 42 tracked combinations — which ones covered the pooled rate, and which five didn't.

each column one combination · each row one same-day batch · 5 of 126 missed

Before the first dispatch, one bet was locked in. The bet said repeat answers would show more spread than plain chance produces. Ranges built from 17 calls would miss the fuller same-day rate, built from 51 calls, more than 10% of the time. That bet was refuted. Those ranges covered the fuller rate 96.0% of the time, across 126 checks. Plain chance explains the same-day spread. The check did not detect anything extra beyond it.

98.47 against a chance ceiling of 112.69

Below the chance ceiling

The observed same-day spread against the simulated chance-only ceiling.

050100150chance ceiling 112.69observed 98.47same-day batch-to-batch spread

A second same-day test asks the question more directly. Does the spread between three same-day repeat sets exceed what plain chance alone produces? We simulated what the spread would look like if every repeat were pure chance at the measured rate, 2,000 times, and compared the real spread to that simulated range. The chance ceiling sat at 112.69 — higher than all but the top 2.5% of those 2,000 simulated chance-only outcomes. The real spread measured 98.47. The test did not detect extra clustering beyond ordinary variation from chance (p=.159).

38 of 42 stayed inside that same noise, two weeks later

Thirty-eight of forty-two

How many day-0-to-day-14 comparisons landed inside same-day noise, out of 42 tracked.

38 of 42
inside same-day noise — day 0 to day 14

The main comparison runs day 0 against day 14. Day 0 pooled 51 calls per brand-question-provider combination across three same-day sets; day 14 used 17 calls. Of 42 combinations, 38 (90.5%) landed inside the same range that same-day chance alone produces. This comparison could only have caught a change of about 30pp at 50% odds of catching it, 40pp at 80% odds. Smaller changes had lower odds of being caught either way.

Four combinations moved beyond that same-day range, all four on the Anthropic provider. The Ordinary, one phrasing: 25.5pp (p=.028). Kiehl's, one phrasing: 35.3pp (p=.015). The Ordinary, a second phrasing: 33.3pp (p=.023). Cetaphil, one phrasing: 25.5pp (p=.043).

If all 42 comparisons were truly unchanged, about 2.1 would cross this threshold by chance alone. Four were observed. No correction was applied because none was pre-registered. These four are reported as-is, not classified as real change or as noise.

Same-day chance already swings twenty points

The band narrows slowly

The typical range a repeat measurement falls in, from chance alone, as the number of calls grows — median across the 42 tracked combinations, and the single widest one.

0%10%20%30%40%381217243451medianwidest combinationcalls per brand-question-provider combination

The bands above are wide because chance alone swings a lot at these call counts. At 51 calls the typical band is 12.2pp. Drop to 34 calls and it widens to 14.8pp. At 24 calls it is 17.4pp. At 17 calls — one full set — it is 19.9pp. At 12 calls it is 23.6pp. At 8 calls it is 26.0pp. At 3 calls it is 36.5pp. A band is the range a repeated measurement typically falls in from chance alone, at that many calls.

The closer the gap, the more it flips

How often each pair's order flips at 24 calls, against how far apart their rates sit.

0%10%20%30%40%50%0510152025nearly tiedwidest gap measuredrate gap, percentage pointsorder flips

The same chance drives the order of two brands. Cetaphil and Paula's Choice on Anthropic rated 41.7% against 40.4%, a 1.2pp gap, nearly tied; their order flips 46.7% of the time, close to a coin toss. Paula's Choice against Cetaphil on OpenAI rated 46.1% against 35.5%, a 10.5pp gap, and the order flips 21.7% of the time. The Ordinary against Paula's Choice on OpenAI rated 63.0% against 46.1%, a 16.9pp gap, and flips 11.4% of the time. The Ordinary against Cetaphil on Anthropic rated 64.0% against 41.7%, a 22.3pp gap, and flips 4.8% of the time. The closer two rates sit, the more often their order flips from which 24 calls you happened to draw. A wide-enough gap holds its order almost every time.

The three bets above were locked in before any data was collected. The half-width curve and the rank-flip numbers here were computed after, from the same measured rates, as description rather than a pre-registered test.

Ten skincare brands, two APIs, fourteen days

This study measured two provider APIs, OpenAI and Anthropic, both with live web search, not a consumer chat app. The APIs carry no personalization, memory, or location. Consumer surfaces may differ; this is a controlled proxy.

One category was measured, skincare, with ten configured brands: two familiar reference brands, six candidates, and two made-up names as a control. Each provider ran 544 calls across the whole study — eight ways of asking, seventeen repeats, four separate batches (three on day 0, one on day 14).

The made-up names never appeared. Across 544 calls per provider, a made-up brand was named 0 times on both providers. That shows the measurement is not inventing brand names wholesale. It says nothing about accuracy for real brands.

This design has only even odds of catching a change smaller than about 30pp, and roughly 80% odds at 40pp. That detection floor applies to every day-0-to-day-14 comparison in this piece, not a caveat on one.

Day 7, a third planned checkpoint, never ran, and day 0 and day 14 were measured at different times of day — both disclosed in full below, alongside what each costs the comparison above.

The design cannot attribute any detected or non-detected change to a specific cause — not a provider update, not a change in which pages get surfaced, not anything else. This design cannot tell those apart.

This study does not test how much the wording of the question matters, whether a persona changes the answer, or whether providers agree with each other. Those are separate studies.

A single weekly reading mostly counts how many times you asked

A single AI-recommendation reading, taken once, carries the wide band from above with it. Two single readings, taken close together and each built from a handful of calls, differ by an amount this study cannot separate from ordinary same-day noise. Only a move past roughly 30pp clears that noise, and even that cannot be pinned to a specific cause.

The same 30-point floor, read three ways

For an SEO or a marketer watching a weekly AI-visibility dashboard, the practical read is the one above: a number moving from one week to the next, by less than roughly 30 points and built from a handful of calls, is not, on its own, evidence that anything changed.

For someone running or selling Generative Engine Optimization (GEO) work, the same logic applies to a before-and-after claim. A tracked recommendation rate that moves after a content change or any other tactic has to clear the detection floor of the measurement making the claim — under a design like this one, about 30 percentage points. Below that floor, the before-and-after numbers are not distinguishable from two draws of the same noisy process, and are not proof the tactic did anything.

For an AI researcher, the two refuted bets are the more interesting result: neither the coverage check nor the second same-day test detected anything beyond plain chance in same-day repeats, at this scale, on real brand questions. Same-day batches of 17 calls already carry a typical band of roughly 20 points; the day-0-to-day-14 comparison, separately, could only have caught drift above roughly 30 points. Another study measuring AI-visibility stability could use both figures when sizing its own sample, within the limits of one category and two providers.

Two APIs, eight phrasings, three bets locked first

The measurement ran on two provider APIs: OpenAI gpt-5.4 and Anthropic claude-sonnet-5. The category was skincare, with ten configured brands, eight phrasings, and one persona. Day 0 ran three same-day batches of 17 calls per combination, pooled to 51. Day 14 ran one batch of 17 calls. Three bets were registered before the first dispatch: the two same-day chance-comparison bets and the day-0-to-day-14 bounded-drift bet. All three were set to publish regardless of outcome, by design. The bets, the thresholds, and the analysis code were locked in before the first dispatch.

Five things ran differently from the plan.

First, on day 0 an account billing outage delayed the OpenAI set by about 40 minutes relative to the Anthropic set. The set was funded and completed inside the same window. The cost is a same-day timing asymmetry within that first set. No correction was applied, because the set still completed inside its registered window.

Second, the original fixed measurement window — 10:00 to 14:00 AEST each day — was dropped partway through the study, after day 0 had already run inside it. Day 14 ran at 06:50 to 07:42 AEST, well outside it. The cost is that day 0 and day 14 differ in time of day, so a difference, or its absence, between them cannot be cleanly separated from that. No correction was applied; none was pre-registered.

Third, day 7 never ran. Its window passed with no dispatch, and the decision was made to skip it rather than run it late. The cost is that the drift comparison covers day 0 to day 14 only, not the fuller three-point design. No correction was applied; the missing calls cannot be reconstructed.

Fourth, day 14 dispatched about 16 hours later than scheduled within its allowed window, after two automated provider health checks flagged and were confirmed as false alarms. The planned confirmation re-checks ran early, 17 to 19 hours apart instead of 24, by a direct decision. The cost is none to the measured content; this is an operations note, not a data change. No correction was needed for that reason.

Fifth, the figures in this piece were built after the numbers above were final. The frozen analysis step that produced them emits numbers only, no charts. The cost is none; the figures present the same locked numbers, built afterward. No correction was needed for that reason.

Earlier published research made the same point in general. Sielinski's "Quantifying Uncertainty in AI Visibility" (arXiv:2603.08924) argued that a single-run reading is a sample from a noisy process, not a fixed truth. This study tests that question directly, on real brand questions, with a bet locked in before looking.

How we ran this

What we measured
42 tracked brand-question-provider combinations, out of the full set, chosen because their day-0 rate wasn't already stuck near 0% or 100%. Ten configured skincare brands (two familiar reference brands, six candidates, two made-up controls), eight phrasings, two provider APIs (OpenAI, Anthropic). Day 0: three same-day batches of 17 calls per combination, pooled to 51. Day 14: one batch, 17 calls.
When
Day 0: 20 July 2026. Day 14: 3 August 2026 (a planned day 7 was skipped — see the piece).
Which AI
OpenAI — gpt-5.4, via API, live search, Anthropic — claude-sonnet-5, via API, live search

The full numbers

Every number in this piece, at full precision. The prose rounds for reading; this table doesn't.

MeasureAs shownExactSource
Same-day intervals that covered the pooled rate96.0%96.0317%Measured
Same-day interval checks (coverage test)126126Derived
Day-0-to-day-14 pairs indistinguishable from same-day noise90.5%90.4762%Measured
Pairs indistinguishable from same-day noise, of 423838Measured
Day-0-to-day-14 pairs tested4242Measured
Smallest two-week change detectable at 50% odds30pp30.0000%Measured
Smallest two-week change detectable at 80% odds40pp40.0000%Measured
Calls per cell, day 0 (pooled)5151Measured
Calls per cell, day 141717Measured
Pairs that moved beyond same-day noise44Measured
Pairs expected to move by chance alone, of 422.12.1Derived
The Ordinary, one phrasing — day-0/day-14 change25.5pp25.4902%Measured
The Ordinary, one phrasing — test resultp=.0280.0284Measured
Kiehl's, one phrasing — day-0/day-14 change35.3pp35.2941%Measured
Kiehl's, one phrasing — test resultp=.0150.0146Measured
The Ordinary, a second phrasing — day-0/day-14 change33.3pp33.3333%Measured
The Ordinary, a second phrasing — test resultp=.0230.0228Measured
Cetaphil, one phrasing — day-0/day-14 change25.5pp25.4902%Measured
Cetaphil, one phrasing — test resultp=.0430.0425Measured
Same-day batch-to-batch spread, measured98.4798.469367Measured
Same-day spread expected from chance alone (97.5th pct)112.69112.691921Measured
Test result, same-day spread vs. chancep=.1590.1585Measured
Typical band width at 3 calls36.5pp36.5424%Measured
Typical band width at 8 calls26.0pp25.9623%Measured
Typical band width at 12 calls23.6pp23.5629%Measured
Typical band width at 17 calls19.9pp19.9266%Measured
Typical band width at 24 calls17.4pp17.3936%Measured
Typical band width at 34 calls14.8pp14.8399%Measured
Typical band width at 51 calls12.2pp12.2474%Measured
Cetaphil vs. Paula's Choice, Anthropic — higher rate41.7%41.6667%Measured
Cetaphil vs. Paula's Choice, Anthropic — lower rate40.4%40.4412%Measured
Cetaphil vs. Paula's Choice, Anthropic — order flips46.7%46.6750%Measured
Cetaphil vs. Paula's Choice, Anthropic — rate gap1.2pp1.2255%Derived
The Ordinary vs. Cetaphil, Anthropic — higher rate64.0%63.9706%Measured
The Ordinary vs. Cetaphil, Anthropic — lower rate41.7%41.6667%Measured
The Ordinary vs. Cetaphil, Anthropic — order flips4.8%4.7750%Measured
The Ordinary vs. Cetaphil, Anthropic — rate gap22.3pp22.3039%Derived
The Ordinary vs. Paula's Choice, OpenAI — higher rate63.0%62.9902%Measured
The Ordinary vs. Paula's Choice, OpenAI — lower rate46.1%46.0784%Measured
The Ordinary vs. Paula's Choice, OpenAI — order flips11.4%11.3750%Measured
The Ordinary vs. Paula's Choice, OpenAI — rate gap16.9pp16.9118%Derived
Paula's Choice vs. Cetaphil, OpenAI — higher rate46.1%46.0784%Measured
Paula's Choice vs. Cetaphil, OpenAI — lower rate35.5%35.5392%Measured
Paula's Choice vs. Cetaphil, OpenAI — order flips21.7%21.6750%Measured
Paula's Choice vs. Cetaphil, OpenAI — rate gap10.5pp10.5392%Derived
Calls per provider testing the two made-up brands, whole study544544Measured
Times a made-up brand was named00Measured
Brand-question-provider combinations analyzed4242Measured
Total measurement spend, this studyUS$92.4892.48Measured
Pre-approved spend ceilingUS$160160.00Measured

What this doesn't settle

  • Provider APIs with live search only — not a consumer chat product.
  • One category, ten configured brands, two providers, this measurement window — does not generalize beyond it.
  • Cannot detect a change smaller than about 30 percentage points at this depth.
  • Day 0 and day 14 differ in time of day; the two cannot be cleanly separated.
  • Cannot attribute any change, or the absence of one, to a specific cause.

How many times would you need to ask, to trust the number?

A diagnosis runs each question multiple times, across two day-separated runs, and shows the band behind every rate — not a single reading dressed up as a fact.

See what a diagnosis finds