AI Newsway

ChatGPT Picks a Favourite in 79% of Shopping Answers. Gemini Does It 7% of the Time.

An audit of 1,536 responses finds the two assistants barely cite the same sources β€” and that API-based research misses what shoppers actually see

|5 min read0
AI Summary
An academic audit of 1,536 AI chatbot responses found ChatGPT expresses a first-person product preference in 79% of shopping answers, compared with 7% for Gemini and 2% for Google AI Overviews. The interfaces shared only 5.4% of cited source domains on average. Provider APIs overlapped their own consumer interfaces on just 12 to 15% of domains, meaning API-based audits do not capture the commercial advice shoppers actually receive.
Online shopping illustration β€” the audit covered 117 real product-recommendation queries put to ChatGPT, Gemini and Google AI Overviews.
Online shopping illustration β€” the audit covered 117 real product-recommendation queries put to ChatGPT, Gemini and Google AI Overviews.

Asked which phone to buy, Google's AI Overviews reports that a particular model is widely rated best. ChatGPT says: my pick right now. That difference in voice is not a stylistic quirk, according to an audit published this week β€” it is a measurable and systematic split in how AI assistants deliver commercial advice, arriving just as both companies build advertising into the same surfaces.

Key takeaways

  • Across 1,536 evaluated responses, ChatGPT expressed a first-person product preference in 79% of product-recommending answers, against 7% for Gemini and 2% for Google AI Overviews.
  • For the same query, the ChatGPT and Gemini interfaces displayed an average of just 5.4% of the same source domains, and shared no domain at all in 76.7% of comparisons.
  • Provider APIs did not reproduce what consumers see, overlapping their own interfaces on only 12.0% of domains for ChatGPT and 14.8% for Gemini β€” which undercuts the standard method for auditing these systems.

How the audit was built

The study, posted to arXiv by a team from Maastricht University, Utrecht University and the University of Zurich, starts from a dataset the authors call ConsumerQ: 2,528 commercial-advice queries written by real users rather than by researchers. From that pool they drew 117 genuine product-recommendation questions and ran each through five surfaces β€” the ChatGPT and Gemini consumer interfaces, both of their APIs, and Google Search's AI Overviews β€” for 1,755 observations collected from a fixed location, with every interface request paired to an API request.

Pairing matters because the paper's sharpest finding is about method rather than about any one assistant. Most published audits of generative search run through provider APIs using researcher-written template queries, for the obvious reasons of cost and reproducibility. The authors argue that for commercial advice specifically, that shortcut measures something consumers never encounter.

Three assistants, three different kinds of answer

The framing gap is the most visible result. ChatGPT states a preference in the first person in 79% of responses, against 7% for Gemini and 2% for AI Overviews. It also labels some product the best in 73% of responses, where Gemini does so 43% of the time and AI Overviews 48%. The same question therefore arrives as a personal recommendation, an undisputed winner, or a menu of options, depending only on which product the shopper opened.

None of it is stable. The recommended products change across repeated identical requests, and even the framing flips β€” a system presents a product as the best in 27 to 40% of repeat cases where it had not done so before. A shopper who asks twice is not asking the same question twice.

The citation layer barely overlaps

Sources diverge further than the answers do. For an identical query, the ChatGPT and Gemini interfaces share on average 5.4% of displayed domains, and in 76.7% of comparisons they have no domain in common whatsoever. Repeated requests keep surfacing new domains rather than converging on a stable set.

That spread has a commercial explanation. Which publishers an assistant can draw on depends on licensing: OpenAI and Google both hold agreements for real-time Reddit data, and OpenAI's deal with Axel Springer puts Politico and Business Insider into OpenAI's pool. The retrieval layer is not a neutral index β€” it is a negotiated one.

Why regulators should care about the API gap

The API result is the one with teeth. Mean domain overlap between an interface and its own API runs 12.0% for ChatGPT and 14.8% for Gemini, and the two expose different types and layers of source information entirely. An auditor querying the API is looking at a different system from the one a consumer uses.

The timing is pointed. On 31 August 2026 the European Commission designated ChatGPT a Very Large Online Search Engine under the Digital Services Act, which brings obligations on systemic risk β€” including consumer protection in online purchasing β€” plus independent audits and researcher data access. The paper's conclusion is that such audits cannot be run through the API and still claim to describe consumer experience.

Commercial pressure is building on the same surface. OpenAI expanded advertising in ChatGPT to 31 European markets in August, and product-recommendation queries already account for roughly 2% of ChatGPT conversations by the largest published usage measurement. We have covered OpenAI's experiments with conversational ad formats; this audit supplies the baseline against which any later change in recommendation behaviour can be measured.

The authors' recommendation to fellow researchers is procedural rather than alarmist: audit the consumer interface, repeat every query, and state which source layer you observed. Single responses collected through an API cannot stand in for the advice people actually receive from a large language model at the point of purchase.

FAQ

Does the study show ChatGPT recommends products for money?

No. The audit measures how recommendations are framed, how consistent they are, and which sources appear β€” not whether any placement was paid. Its contribution is a documented pre-advertising baseline and a method that would detect a shift, which the authors argue is what regulators and researchers currently lack.

Why can't researchers just use the ChatGPT and Gemini APIs?

Because the APIs return a materially different view. Mean domain overlap with the corresponding consumer interface was 12.0% for ChatGPT and 14.8% for Gemini, and the source information the two layers expose differs in kind, so API observations cannot be assumed to represent what shoppers see.

Is the query dataset available?

The authors curated ConsumerQ, a set of 2,528 real commercial-advice queries written by users, and drew the 117 product-recommendation questions analysed in the paper from it. Full details of the dataset and the collection protocol are in the arXiv preprint.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

OpenAI Drops Message Caps on ChatGPT's Free Tier and Hands It a Think Button
LLM & Chatbots

OpenAI Drops Message Caps on ChatGPT's Free Tier and Hands It a Think Button

OpenAI is lifting message caps on ChatGPT's free tier and adding a Think button, with GPT-5.6 Luna becoming the default for Free and Go users.

Seung Jung41 days ago
ChatGPT Ads Reach the UK, Japan, Korea, Brazil and Mexico
LLM & Chatbots

ChatGPT Ads Reach the UK, Japan, Korea, Brazil and Mexico

OpenAI turned on ChatGPT sponsored placements in five more countries on 11 August, limited to logged-in adults on the Free and Go tiers.

Seung Jung33 days ago
OpenAI Ships a Separate ChatGPT for Teens, Betting on Age Assurance Over ID Checks
LLM & Chatbots

OpenAI Ships a Separate ChatGPT for Teens, Betting on Age Assurance Over ID Checks

OpenAI's ChatGPT for Teens launches for ages 13-17 with content limits, 90-minute break reminders, opt-in parental quiet hours and age assurance over ID checks.

Seung Jung29 days ago
ChatGPT Images 2.5 Halves Generation Latency and Adds a Sketch Canvas
LLM & Chatbots

ChatGPT Images 2.5 Halves Generation Latency and Adds a Sketch Canvas

OpenAI shipped ChatGPT Images 2.5 with up to 50% lower latency, a Sketch drawing canvas, templates, pinned comments and two new API models.

Seung Jung6 days ago
OpenAI Ships a Data Agent and a Wall Street ChatGPT on the Same Day
LLM & Chatbots

OpenAI Ships a Data Agent and a Wall Street ChatGPT on the Same Day

OpenAI launched a Data agent for ChatGPT Work and ChatGPT for Financial Services, built with Morgan Stanley and Evercore, on the same Thursday.

Seung Jung6 days ago
iOS 27 Code Shows Siri's Server Model Can Be Swapped for Claude or GPT-5.6
LLM & Chatbots

iOS 27 Code Shows Siri's Server Model Can Be Swapped for Claude or GPT-5.6

Private frameworks in iOS 27 and macOS Golden Gate expose a Model Delegation API and an Inference Provider protocol that can hand Siri's planner to Claude.

Seung Jungyesterday