Do ChatGPT, Claude, and Gemini recommend the same brands?
We asked all three engines' own APIs the same skincare questions, twice, on separate days, and registered one number before looking: how much their real recommendations overlap. The answer is yes, almost entirely — and the one place they didn't agree turns out to be a coin flip for all three, not a real difference of opinion.
Yes, almost entirely: ChatGPT, Claude, and Gemini named the same two skincare brands as sure bets, and split on only one — a brand that landed within a coin flip of the halfway line for every single engine, which is not the same as the engines actually disagreeing.
Every AI engine feels like a separate black box to reverse-engineer — win ChatGPT, and you still don't know whether Claude or Gemini will name you, so the visibility work triples with no guarantee one result means anything for the others. The short answer, from 816 real calls: the three engines named nearly the same skincare shortlist, and the gap between them was thinner than it looks.
Two comparisons already exist — and neither is this one
Ahrefs compared Google's AI Overviews against Google's own AI Mode on the same queries and found the two surfaces cited the same URL only 13.7% of the time (16.3% restricted to just the top 3 citations each), across 730,000 response pairs. Semrush ran the same 100 prompts through ChatGPT's minimal-reasoning and high-reasoning modes and found only 25.6% of the cited domains overlapped. Both are one company comparing its own two surfaces against itself — Google versus Google, OpenAI versus OpenAI — and both are about which source URLs got cited, not which brand got recommended. Nobody has measured whether different AI companies agree on which brand to actually name.
The open question this closes: if a brand lands on ChatGPT's shortlist, does that mean anything for whether Claude or Gemini also name it?
One number, locked in before the first call
The same skincare questions went to all three engines' own APIs, run twice on two separate days so a one-day fluke can't pass as a finding; skincare was the category because a companion study had already validated it works cleanly for this kind of measurement. A brand counted as a real recommendation for an engine only if that engine recommended it in more than 50% of its calls, pooled across both days — not a simple majority once, the full two-day figure. The one number locked in before any call ran: if the three engines' shortlists overlapped less than 40% on average, that would mean the engines split; anything higher means they largely agree. 816 real calls total across the three engines and both days; 6 real brands were in the comparison, plus reference brands that every engine recommends almost automatically, excluded because they'd make any comparison look like agreement.
Two of the three gave the identical answer, twice
Six real skincare brands, whether each engine recommended the brand in more than half its calls, pooled across two separate days. Filled = yes.
The three engines' shortlists overlapped 78% on average — well clear of the 40% line that would have meant they split, nearly double it. Claude and Gemini's shortlists were 100% identical. ChatGPT overlapped with each of the other two at 67% — the same figure in both directions. The most memorable fact: ChatGPT gave the exact same shortlist on both separate days, 100% agreement a full day apart, and so did Gemini, 100%, independently, with no coordination between the two runs. “The engines split the shortlist” is the interpretation this evidence rules out; “the engines largely agree on who to recommend” is what happened.
The one brand that's a coin flip for everyone
Only 1 brand out of 6 ever caused a split between any pair of engines: Paula's Choice. ChatGPT recommended it in 47% of its calls, Claude in 51%, Gemini in 55% — landing just under, just over, and further over the 50% line respectively. ChatGPT's likely range for that rate is 41.2% to 53.0%; Claude's is 44.8% to 56.6%; Gemini's is 48.8% to 60.6%. These three ranges overlap each other almost completely, so this evidence cannot rule out that all three engines actually treat this brand identically, and it was sampling noise alone that decided which side of the 50% line each one's count happened to land on. These are the pre-registered per-brand ranges, not a new procedure invented after the fact. One more sign this is a toss-up rather than a stable per-engine opinion: Claude's own shortlist only matched itself across its own two separate days 67% of the time, versus 100% for the other two — the same brand wobbled in and out for Claude across its own two independent days. The apparent disagreement isn't the engines disagreeing with each other; it's all three circling the same genuinely uncertain answer.
Skincare, two days, and one blind spot
This is one category (skincare), a two-day window, not months or years. Gemini's own API does not reveal which web pages it read before answering, so — unlike the other two engines — there's no way to check whether Gemini's choices trace back to specific retrieved pages; it can only be compared on its final answers. This does not generalize to every product category, every pair of AI companies, or across time.
Your own borderline brands are the risk, not which engine you court
If a brand clears the halfway point solidly on one engine, this evidence says it's likely to clear it on the others too — so spending energy asking whether Gemini favors you as much as ChatGPT does is often solving the wrong problem. The real exposure this data points to is being a brand that sits genuinely near that halfway line on any one engine: that instability is real and worth attention, rather than picking a favorite platform to chase.
Real calls, registered before any of them landed
This was pre-registered before the first call was dispatched — not a re-analysis of old data — so the finding was committed to before any data collection, not just before analysis. Real calls went against the three companies' own APIs with live web search, not the consumer ChatGPT, Claude, and Gemini apps people use day to day, run twice on two separate calendar days specifically so a one-day fluke couldn't pass as a finding. Total cost: $103.36.
Deviations
Partway through data collection, a shared safety limit on how many calls could run that month briefly blocked one engine's first day of calls. The limit protected a completely separate, unrelated study's budget, not this one's; rather than wait, the monthly limit was raised, and all three engines' first-day calls then completed normally. This did not touch the $103.36 budget for this study or change any measurement — it was a scheduling limit, not a money limit.
How we ran this
- What we measured
- 816 real API calls — three engines' own APIs, eight real skincare brands (six in the registered comparison, two saturating reference brands excluded) plus two invented controls, eight phrasings run 17 times each, twice, two calendar days apart.
- When
- 23 and 24 July 2026, inside the same daily time window both days.
- Which AI
- ChatGPT (OpenAI API) with live web search, Claude (Anthropic API) with live web search, Gemini (Google API) with live web search — its API does not reveal which pages it read, so it is excluded from the retrieval-vs-ranking comparison available for the other two
The full numbers
Every number in this piece, at full precision. The prose rounds for reading; this table doesn't.
| Measure | As shown | Exact | Source |
|---|---|---|---|
| Average shortlist overlap across all three engines | 78% | 77.7778% | Measured |
| Line set in advance for 'the engines split' | 40% | 40.0000% | Measured |
| Shortlist overlap — Claude vs Gemini | 100% | 100.0000% | Measured |
| Shortlist overlap — ChatGPT vs Claude | 67% | 66.6667% | Measured |
| Shortlist overlap — ChatGPT vs Gemini | 67% | 66.6667% | Measured |
| Line for 'this counts as a real recommendation' | 50% | 50.0000% | Measured |
| Paula's Choice — ChatGPT recommended rate | 47% | 47.0588% | Measured |
| Paula's Choice — ChatGPT likely range, low end | 41.2% | 41.2094% | Measured |
| Paula's Choice — ChatGPT likely range, high end | 53.0% | 52.9902% | Measured |
| Paula's Choice — Claude recommended rate | 51% | 50.7353% | Measured |
| Paula's Choice — Claude likely range, low end | 44.8% | 44.8252% | Measured |
| Paula's Choice — Claude likely range, high end | 56.6% | 56.6249% | Measured |
| Paula's Choice — Gemini recommended rate | 55% | 54.7794% | Measured |
| Paula's Choice — Gemini likely range, low end | 48.8% | 48.8390% | Measured |
| Paula's Choice — Gemini likely range, high end | 60.6% | 60.5867% | Measured |
| ChatGPT — same shortlist both separate days | 100% | 100.0000% | Measured |
| Gemini — same shortlist both separate days | 100% | 100.0000% | Measured |
| Claude — shortlist overlap between the two days | 67% | 66.6667% | Measured |
| Brands the three engines disagreed on | 1 | 1 | Measured |
| Brands in the registered comparison | 6 | 6 | Measured |
| Fake planted brands recommended, out of 816 calls | 0% | 0.0000% | Measured |
| Real API calls run for this study | 816 | 816 | Derived |
| What the study cost | $103.36 | 103.36 | Measured |
| Ahrefs — response pairs examined | 730,000 | 730,000 | Cited |
| Ahrefs — Google's own two AI surfaces, same-URL overlap | 13.7% | 13.7% | Cited |
| Ahrefs — Google's own two AI surfaces, top-3 citation overlap | 16.3% | 16.3% | Cited |
| Semrush — ChatGPT's own two modes, cited-domain overlap | 25.6% | 25.6% | Cited |
| Semrush — prompts run per mode | 100 | 100 | Cited |
What this doesn't settle
- One category (skincare) and a two-day window — this is a finding for skincare on these three APIs on these dates, not a number that generalizes across categories, engine versions, or time.
- Gemini's API hides which pages it read, so it's excluded from the retrieval-vs-ranking comparison available between ChatGPT and Claude.
- These queries ran against the three companies' own APIs with live web search — not the consumer ChatGPT, Claude, or Gemini apps. The APIs carry no personalization, no conversation memory, and no geo-targeting.
- The one brand where the engines' shortlists differed sits within each engine's own uncertainty range — this evidence cannot confirm the engines actually see that brand differently, only that their pooled counts landed on different sides of the 50% line.
References
- Are AI Mode and AI Overviews Just Different Versions of the Same Answer? (730K Responses Studied) Accessed 24 July 2026.
- Only 25% of Cited Sources Overlap Between ChatGPT's Different Reasoning Modes Accessed 24 July 2026.
Does every AI system your customers use actually agree on your brand?
A diagnosis runs the same evidence-first check across the AI systems that matter for your category — and shows you exactly where they agree and where one of them is the outlier. Not a score: the evidence, run twice.