Do ChatGPT, Claude, and Gemini recommend the same brands?

We asked all three engines' own APIs the same skincare questions, twice, on separate days, and registered one number before looking: how much their real recommendations overlap. The answer is yes, almost entirely — and the one place they didn't agree turns out to be a coin flip for all three, not a real difference of opinion.

Yes, almost entirely: ChatGPT, Claude, and Gemini named the same two skincare brands as sure bets, and split on only one — a brand that landed within a coin flip of the halfway line for every single engine, which is not the same as the engines actually disagreeing.

24 July 2026Results registered before data was seen

Every AI engine feels like a separate black box to reverse-engineer — win ChatGPT, and you still don't know whether Claude or Gemini will name you, so the visibility work triples with no guarantee one result means anything for the others. The short answer, from 816 real calls: the three engines named nearly the same skincare shortlist, and the gap between them was thinner than it looks.

Two comparisons already exist — and neither is this one

Ahrefs compared Google's AI Overviews against Google's own AI Mode on the same queries and found the two surfaces cited the same URL only 13.7% of the time (16.3% restricted to just the top 3 citations each), across 730,000 response pairs. Semrush ran the same 100 prompts through ChatGPT's minimal-reasoning and high-reasoning modes and found only 25.6% of the cited domains overlapped. Both are one company comparing its own two surfaces against itself — Google versus Google, OpenAI versus OpenAI — and both are about which source URLs got cited, not which brand got recommended. Nobody has measured whether different AI companies agree on which brand to actually name.

The open question this closes: if a brand lands on ChatGPT's shortlist, does that mean anything for whether Claude or Gemini also name it?

One number, locked in before the first call

The same skincare questions went to all three engines' own APIs, run twice on two separate days so a one-day fluke can't pass as a finding; skincare was the category because a companion study had already validated it works cleanly for this kind of measurement. A brand counted as a real recommendation for an engine only if that engine recommended it in more than 50% of its calls, pooled across both days — not a simple majority once, the full two-day figure. The one number locked in before any call ran: if the three engines' shortlists overlapped less than 40% on average, that would mean the engines split; anything higher means they largely agree. 816 real calls total across the three engines and both days; 6 real brands were in the comparison, plus reference brands that every engine recommends almost automatically, excluded because they'd make any comparison look like agreement.

Two of the three gave the identical answer, twice

Where the three engines' shortlists actually differ

Six real skincare brands, whether each engine recommended the brand in more than half its calls, pooled across two separate days. Filled = yes.

ChatGPT
Claude
Gemini
La Roche-Posay
The Ordinary
Paula's Choice
Cetaphil
Kiehl's
Drunk Elephant
Filled = recommended in more than half of that engine’s calls, pooled across both days. Five of six brands are unanimous; Paula’s Choice is the one row that splits — and it sits inside each engine’s own margin of uncertainty (see below).

The three engines' shortlists overlapped 78% on average — well clear of the 40% line that would have meant they split, nearly double it. Claude and Gemini's shortlists were 100% identical. ChatGPT overlapped with each of the other two at 67% — the same figure in both directions. The most memorable fact: ChatGPT gave the exact same shortlist on both separate days, 100% agreement a full day apart, and so did Gemini, 100%, independently, with no coordination between the two runs. “The engines split the shortlist” is the interpretation this evidence rules out; “the engines largely agree on who to recommend” is what happened.

The one brand that's a coin flip for everyone

Only 1 brand out of 6 ever caused a split between any pair of engines: Paula's Choice. ChatGPT recommended it in 47% of its calls, Claude in 51%, Gemini in 55% — landing just under, just over, and further over the 50% line respectively. ChatGPT's likely range for that rate is 41.2% to 53.0%; Claude's is 44.8% to 56.6%; Gemini's is 48.8% to 60.6%. These three ranges overlap each other almost completely, so this evidence cannot rule out that all three engines actually treat this brand identically, and it was sampling noise alone that decided which side of the 50% line each one's count happened to land on. These are the pre-registered per-brand ranges, not a new procedure invented after the fact. One more sign this is a toss-up rather than a stable per-engine opinion: Claude's own shortlist only matched itself across its own two separate days 67% of the time, versus 100% for the other two — the same brand wobbled in and out for Claude across its own two independent days. The apparent disagreement isn't the engines disagreeing with each other; it's all three circling the same genuinely uncertain answer.

Skincare, two days, and one blind spot

This is one category (skincare), a two-day window, not months or years. Gemini's own API does not reveal which web pages it read before answering, so — unlike the other two engines — there's no way to check whether Gemini's choices trace back to specific retrieved pages; it can only be compared on its final answers. This does not generalize to every product category, every pair of AI companies, or across time.

Your own borderline brands are the risk, not which engine you court

If a brand clears the halfway point solidly on one engine, this evidence says it's likely to clear it on the others too — so spending energy asking whether Gemini favors you as much as ChatGPT does is often solving the wrong problem. The real exposure this data points to is being a brand that sits genuinely near that halfway line on any one engine: that instability is real and worth attention, rather than picking a favorite platform to chase.

Real calls, registered before any of them landed

This was pre-registered before the first call was dispatched — not a re-analysis of old data — so the finding was committed to before any data collection, not just before analysis. Real calls went against the three companies' own APIs with live web search, not the consumer ChatGPT, Claude, and Gemini apps people use day to day, run twice on two separate calendar days specifically so a one-day fluke couldn't pass as a finding. Total cost: $103.36.

Deviations

Partway through data collection, a shared safety limit on how many calls could run that month briefly blocked one engine's first day of calls. The limit protected a completely separate, unrelated study's budget, not this one's; rather than wait, the monthly limit was raised, and all three engines' first-day calls then completed normally. This did not touch the $103.36 budget for this study or change any measurement — it was a scheduling limit, not a money limit.

How we ran this

What we measured
816 real API calls — three engines' own APIs, eight real skincare brands (six in the registered comparison, two saturating reference brands excluded) plus two invented controls, eight phrasings run 17 times each, twice, two calendar days apart.
When
23 and 24 July 2026, inside the same daily time window both days.
Which AI
ChatGPT (OpenAI API) with live web search, Claude (Anthropic API) with live web search, Gemini (Google API) with live web search — its API does not reveal which pages it read, so it is excluded from the retrieval-vs-ranking comparison available for the other two

The full numbers

Every number in this piece, at full precision. The prose rounds for reading; this table doesn't.

MeasureAs shownExactSource
Average shortlist overlap across all three engines78%77.7778%Measured
Line set in advance for 'the engines split'40%40.0000%Measured
Shortlist overlap — Claude vs Gemini100%100.0000%Measured
Shortlist overlap — ChatGPT vs Claude67%66.6667%Measured
Shortlist overlap — ChatGPT vs Gemini67%66.6667%Measured
Line for 'this counts as a real recommendation'50%50.0000%Measured
Paula's Choice — ChatGPT recommended rate47%47.0588%Measured
Paula's Choice — ChatGPT likely range, low end41.2%41.2094%Measured
Paula's Choice — ChatGPT likely range, high end53.0%52.9902%Measured
Paula's Choice — Claude recommended rate51%50.7353%Measured
Paula's Choice — Claude likely range, low end44.8%44.8252%Measured
Paula's Choice — Claude likely range, high end56.6%56.6249%Measured
Paula's Choice — Gemini recommended rate55%54.7794%Measured
Paula's Choice — Gemini likely range, low end48.8%48.8390%Measured
Paula's Choice — Gemini likely range, high end60.6%60.5867%Measured
ChatGPT — same shortlist both separate days100%100.0000%Measured
Gemini — same shortlist both separate days100%100.0000%Measured
Claude — shortlist overlap between the two days67%66.6667%Measured
Brands the three engines disagreed on11Measured
Brands in the registered comparison66Measured
Fake planted brands recommended, out of 816 calls0%0.0000%Measured
Real API calls run for this study816816Derived
What the study cost$103.36103.36Measured
Ahrefs — response pairs examined730,000730,000Cited
Ahrefs — Google's own two AI surfaces, same-URL overlap13.7%13.7%Cited
Ahrefs — Google's own two AI surfaces, top-3 citation overlap16.3%16.3%Cited
Semrush — ChatGPT's own two modes, cited-domain overlap25.6%25.6%Cited
Semrush — prompts run per mode100100Cited

What this doesn't settle

  • One category (skincare) and a two-day window — this is a finding for skincare on these three APIs on these dates, not a number that generalizes across categories, engine versions, or time.
  • Gemini's API hides which pages it read, so it's excluded from the retrieval-vs-ranking comparison available between ChatGPT and Claude.
  • These queries ran against the three companies' own APIs with live web search — not the consumer ChatGPT, Claude, or Gemini apps. The APIs carry no personalization, no conversation memory, and no geo-targeting.
  • The one brand where the engines' shortlists differed sits within each engine's own uncertainty range — this evidence cannot confirm the engines actually see that brand differently, only that their pooled counts landed on different sides of the 50% line.

Does every AI system your customers use actually agree on your brand?

A diagnosis runs the same evidence-first check across the AI systems that matter for your category — and shows you exactly where they agree and where one of them is the outlier. Not a score: the evidence, run twice.

See what a diagnosis finds