Give three AI models two near-identical skincare products, one from a brand people already recognize and one they've never heard of, and the familiar name wins the recommendation in every single trial. That is the headline result from a June 2026 study by researchers Xi Chu and Yupeng Hou, who tested GPT-4o-mini, Claude Sonnet, and Gemini Flash against each other and measured what they call an Incumbent Advantage Index of 10.0, the maximum possible score. The advantage is not permanent, and it collapses the moment a real signal appears. That matters for any brand trying to build LLM brand visibility strong enough to get recommended by an AI system, not just mentioned by one.
What the incumbent advantage study actually measured
Chu and Hou ran three experiments using skincare, a category chosen because buyers usually cannot judge quality before they buy and have to lean on brand reputation instead. Under identical product specifications, all three models recommended the established brand over the unfamiliar one in every trial, producing the maximum possible Incumbent Advantage Index of 10.0. A robustness check on search goods, products whose quality is easy to verify before purchase, found the same pattern held outside skincare specifically.
The researchers describe this as a Conditional Monopoly: the advantage is not a fixed property of the brand itself, it is a product of the conditions the model is reasoning under. Strip away any signal that distinguishes the two options and the model defaults to familiarity, much like a shopper grabbing the name they recognize off a crowded shelf. It is a different bias from the one covered in VizibleAI's look at how AI models judge unfamiliar brands by the literal meaning of their name, but it compounds the same risk for new entrants: a model will pick the brand it already knows even when there is no real reason to.
A Conditional Monopoly, in Chu and Hou's own term: a well-known brand's recommendation advantage holds completely when no other signal distinguishes the options, and disappears the moment one does.
The advantage disappears with less than a tenth of a star
A rating difference smaller than 0.1 stars was enough to end the incumbent's dominant position entirely in Chu and Hou's tests. That makes AI recommendation bias toward familiar brands real but shallow: it only wins when the model has nothing else to reason from.
Once the unfamiliar brand had almost any objective edge, even a fractional one, the model stopped defaulting to the household name and started reasoning from the actual data in front of it. That is a sharply different picture from backlink-driven search rankings, where an incumbent's advantage can take years to erode. Here it took less than a tenth of a star.
The practical read for a challenger brand is that the fight is not to out-market the incumbent on recognition, which is a losing game by definition. It is to make sure an actual comparative signal, a rating or a verified spec, is visible and attached to the product in the first place. A category-creating startup with no incumbent to dislodge and no track record to point to faces exactly this problem from day one.
Marketing language that breaks the monopoly, and the risk hiding inside it
Authority-style claims, including clinical-evidence language, were enough to overcome the incumbent's advantage at what the researchers call a Bias Surplus Value of 0.17 rating points, roughly the same persuasive weight as a real 0.17-star quality improvement, even when the clinical evidence behind the claim was fabricated.
That finding should make brand and legal teams uneasy rather than excited. The tactic worked in the study specifically because the models could not verify whether the clinical claim was real, which means it would work just as well for a brand making an honest, substantiated claim as for one making one up. VizibleAI has covered separately how often AI engines get brand facts wrong, and who ends up liable when they do; a brand that leans on an unverifiable claim to win a recommendation today is building its visibility on the same shaky ground those cases sit on. The version of this tactic that actually holds up once anyone checks is the boring one: publish a real, checkable claim with the data behind it.
When every brand runs the same GEO playbook, the edge evaporates for everyone
Coordinated optimization erases its own payoff. When every brand in Chu and Hou's simulation adopted the same generative engine optimization strategy, each brand's individual advantage collapsed from 0.802 to 0.007 on the researchers' payoff measure, while brands that did not participate at all received zero recommendations.
This is the part of the incumbent-advantage research that GEO vendors, VizibleAI included, have less incentive to advertise. Early movers who adopt real differentiation signals before competitors catch on capture a large, measurable advantage. The playing field re-levels once an entire category publishes the same comparative claims, since the signal that used to separate brands has simply become table stakes. Standing entirely outside that shift is the one strategy the data rules out completely: a brand that skips the game is not neutral, it scores zero recommendations in the same tests. Knowing your current AI Share of Voice relative to competitors is what tells you whether you are early to this shift or already behind it.
A separate study put a real price on the same bias
Brand-steering bias is not confined to a lab recommendation score. A 2026 experiment reported in ProMarket by Amit Zac and Michal Gal gave 265 participants a 15-euro grocery budget and a customized AI shopping assistant, and the version of the assistant that steered users toward pricier, brand-name products raised average spending by more than one euro, even though Amazon's own ratings showed no meaningful quality difference between the options.
Participants in that experiment reported lower trust in the AI and said the task felt harder, yet they kept spending more anyway, and none of them detected that they had been steered. That combination of lower trust, higher spending, and undetected manipulation is what should concern brand teams more than the lab percentage from the incumbent study. A research paper's recommendation score is one thing; a live shopping assistant already nudging real purchases, with the buyer unaware it happened, is another. Zac and Gal's proposed fixes, among them diversity-enhancing rankings and limits on price steering, are aimed at the platforms running these assistants, not at the brands being recommended.
Why AI citations don't show the same bias AI recommendations do
Recommendation bias and citation bias point in opposite directions. The Princeton-led GEO research identified an Equalizer Effect, in which a source sitting well outside the top search results can still gain outsized visibility inside an AI-generated answer, the reverse of the incumbent dominance Chu and Hou measured in direct product recommendations.
The difference comes down to what the AI system is doing. A citation task pulls together sources to answer a question, where a well-argued page from an unfamiliar domain can still get quoted if it answers the query precisely. A recommendation task picks a single winner between options, where the model falls back on brand familiarity exactly when it has nothing else to go on. Building Entity recognition, the thing that lets a model treat a brand as a known, disambiguated thing rather than an unfamiliar string of text, is what closes the gap between the two outcomes. VizibleAI's separate research on what actually predicts AI citations and its piece on why entity data carries so much weight with AI engines both point at the same lever: become recognizable to the model before you need it to pick you.
Frequently Asked Questions
Does AI recommendation bias toward incumbent brands actually hold up across different AI models?
Yes, in the specific test Chu and Hou ran, GPT-4o-mini, Claude Sonnet, and Gemini Flash all defaulted to the established brand under identical specifications, producing the maximum Incumbent Advantage Index of 10.0. The researchers note the three models respond differently once a marketing claim is introduced, so the bias is consistent at baseline but not identical in how each model can be moved off it.
How big a competitive edge does a challenger brand need to overcome incumbent bias in AI recommendations?
In the study, a rating advantage smaller than 0.1 stars was enough to end the incumbent's dominant position entirely. That is a far smaller edge than what typically moves organic search rankings, which suggests the bias is shallow rather than structural: it holds only when a model genuinely has no other signal to reason from.
Is it worth making bold marketing claims to win an AI recommendation?
The study found authority-style and clinical-evidence language can overcome incumbent bias at a Bias Surplus Value of 0.17 rating points, but the researchers tested this using fabricated claims the models could not verify. Making an unsubstantiated claim risks the same hallucination and liability exposure covered elsewhere on this site; a real, checkable claim achieves the same lift without the risk.
Does this incumbent bias apply to AI citations as well as AI recommendations?
No, and the direction reverses. Citation tasks show an Equalizer Effect, where a lower-ranked or unfamiliar source can still get quoted in an AI-generated answer if it answers the query precisely. Recommendation tasks, where a model picks one winner between options, are where incumbent brand familiarity takes over by default.
What should a challenger brand do first in response to this research?
Measure where the brand actually stands today rather than assuming the bias applies equally everywhere; AI Share of Voice benchmarking shows the gap varies a lot by category and engine. Then prioritize getting one verifiable, specific differentiator, a rating or a certification, attached to the brand wherever an AI system might encounter it.



