Do ChatGPT, Claude, and Gemini recommend the same brands?
We asked all three engines' own APIs the same skincare questions, twice, on separate days, and registered one number before looking: how much their real recommendations overlap. The answer is yes, almost entirely — and the one place they didn't agree turns out to be a coin flip for all three, not a real difference of opinion.
Every AI engine feels like a separate black box to reverse-engineer — win ChatGPT, and you still don't know whether Claude or Gemini will name you, so the visibility work triples with no guarantee one result means anything for the others. The short answer, from 816 real calls: the three engines named nearly the same skincare shortlist, and the gap between them was thinner than it looks.
Two comparisons already exist — and neither is this one
Ahrefs compared Google's AI Overviews against Google's own AI Mode on the same queries and found the two surfaces cited the same URL only 13.7% of the time (16.3% restricted to just the top 3 citations each), across 730,000 response pairs.[1] Semrush ran the same 100 prompts through ChatGPT's minimal-reasoning and high-reasoning modes and found only 25.6% of the cited domains overlapped.[2] Both are one company comparing its own two surfaces against itself — Google versus Google, OpenAI versus OpenAI — and both are about which source URLs got cited, not which brand got recommended. Nobody has measured whether different AI companies agree on which brand to actually name.
The open question this closes: if a brand lands on ChatGPT's shortlist, does that mean anything for whether Claude or Gemini also name it?
One number, locked in before the first call
The same skincare questions went to all three engines' own APIs, run twice on two separate days so a one-day fluke can't pass as a finding; skincare was the category because a companion study had already validated it works cleanly for this kind of measurement. A brand counted as a real recommendation for an engine only if that engine recommended it in more than 50% of its calls, pooled across both days — not a simple majority once, the full two-day figure. The one number locked in before any call ran: if the three engines' shortlists overlapped less than 40% on average, that would mean the engines split; anything higher means they largely agree. 816 real calls total across the three engines and both days; 6 real brands were in the comparison, plus reference brands that every engine recommends almost automatically, excluded because they'd make any comparison look like agreement.
Two of the three gave the identical answer, twice
The three engines' shortlists overlapped 78% on average — well clear of the 40% line that would have meant they split, nearly double it. Claude and Gemini's shortlists were 100% identical. ChatGPT overlapped with each of the other two at 67% — the same figure in both directions. The most memorable fact: ChatGPT gave the exact same shortlist on both separate days, 100% agreement a full day apart, and so did Gemini, 100%, independently, with no coordination between the two runs. “The engines split the shortlist” is the interpretation this evidence rules out; “the engines largely agree on who to recommend” is what happened.
The one brand that's a coin flip for everyone
Only 1 brand out of 6 ever caused a split between any pair of engines: Paula's Choice. ChatGPT recommended it in 47% of its calls, Claude in 51%, Gemini in 55% — landing just under, just over, and further over the 50% line respectively. ChatGPT's likely range for that rate is 41.2% to 53.0%; Claude's is 44.8% to 56.6%; Gemini's is 48.8% to 60.6%. These three ranges overlap each other almost completely, so this evidence cannot rule out that all three engines actually treat this brand identically, and it was sampling noise alone that decided which side of the 50% line each one's count happened to land on. These are the pre-registered per-brand ranges, not a new procedure invented after the fact. One more sign this is a toss-up rather than a stable per-engine opinion: Claude's own shortlist only matched itself across its own two separate days 67% of the time, versus 100% for the other two — the same brand wobbled in and out for Claude across its own two independent days. The apparent disagreement isn't the engines disagreeing with each other; it's all three circling the same genuinely uncertain answer.
Skincare, two days, and one blind spot
This is one category (skincare), a two-day window, not months or years. Gemini's own API does not reveal which web pages it read before answering, so — unlike the other two engines — there's no way to check whether Gemini's choices trace back to specific retrieved pages; it can only be compared on its final answers. This does not generalize to every product category, every pair of AI companies, or across time.
Your own close-to-the-line brands are the risk, not which engine you court
If a brand clears the halfway point solidly on one engine, this evidence says it's likely to clear it on the others too — so spending energy asking whether Gemini favors you as much as ChatGPT does is often solving the wrong problem. The real exposure this data points to is being a brand that sits genuinely near that halfway line on any one engine: that instability is real and worth attention, rather than picking a favorite platform to chase.
Real calls, registered before any of them landed
This was pre-registered before the first call was dispatched — not a re-analysis of old data — so the finding was committed to before any data collection, not just before analysis. Real calls went against the three companies' own APIs with live web search, not the consumer ChatGPT, Claude, and Gemini apps people use day to day, run twice on two separate calendar days specifically so a one-day fluke couldn't pass as a finding. Total cost: $103.36.
Deviations
Partway through data collection, a shared safety limit on how many calls could run that month briefly blocked one engine's first day of calls. The limit protected a completely separate, unrelated study's budget, not this one's; rather than wait, the monthly limit was raised, and all three engines' first-day calls then completed normally. This did not touch the $103.36 budget for this study or change any measurement — it was a scheduling limit, not a money limit.
How we ran this
What we measured: 816 real API calls — three engines' own APIs, eight real skincare brands (six in the registered comparison, two saturating reference brands excluded) plus two invented controls, eight phrasings run 17 times each, twice, two calendar days apart.
When: 23 and 24 July 2026, inside the same daily time window both days.
References
- Ahrefs. (2026). *Are AI Mode and AI Overviews just different versions of the same answer? (730K responses studied)*. Ahrefs. https://ahrefs.com/blog/ai-overviews-vs-ai-mode/
- Semrush. (2026). *Only 25% of cited sources overlap between ChatGPT's different reasoning modes*. Semrush. https://www.semrush.com/blog/chatgpt-reasoning-ai-visibility/