AI engines rarely disagree about whether to recommend a brand. They disagree about which one. BrightEdge's April 2026 analysis of five major AI engines, ChatGPT, Perplexity, Gemini, Google AI Mode, and Google AI Overviews, found that any two of them agree on the same top-100 brand list only 36% to 55% of the time. In healthcare, a separate BrightEdge breakdown found that gap widens further still, down to an average of 60% agreement and as low as 40% between some engine pairs. Track your brand's AI visibility on a single engine, and you are measuring a fraction of a much messier picture.
How much do AI engines actually agree on brand recommendations?
Less than most brand teams assume, and the gap holds up across a large sample. BrightEdge's AI Catalyst platform compared top-100 brand mention lists across ten engine pairs spanning ChatGPT, Perplexity, Gemini, Google AI Mode, and Google AI Overviews, covering nine industries from B2B technology to insurance. Pairwise overlap landed between 36% and 55%, so even the two most aligned engines disagree on close to half of their respective top-100 lists.
That range is an average across industries, not one fixed number, and it moves with the category being queried. A brand doing well in one engine's rankings has no guarantee of showing up at all on another, no matter how strong its AI Share of Voice looks when measured on just that one engine. The overlap figure describes a relationship between two engines and a topic. It is not a stable trait that belongs to the brand.
Any two AI engines agree on the same top-100 brand list only 36% to 55% of the time, a gap wide enough that a brand's strongest engine and its weakest can look like they are describing two different companies.
Why healthcare brands see far less AI agreement than retail ones
Industry changes the agreement rate more than engine choice does. BrightEdge's May 2026 breakdown found retail brands agreeing across AI engines 97% of the time on average, ranging from 92% to 100%. Travel brands averaged 94% agreement and tech brands 88%, while finance dropped to 71% and healthcare fell to just 60%, with individual engine pairs as low as 40%.
Individual brands inside those categories show the pattern at its sharpest, per that same BrightEdge dataset. Mayo Clinic appears in 13.1% of Gemini's healthcare mentions but only 1.5% of Google AI Overviews', a nine-fold gap. Goldman Sachs ranks in the top 10 on Google AI Mode and does not appear in the top 100 at all on ChatGPT, Gemini, or Perplexity. The engines are not disagreeing at random. They are weighing the same evidence differently, and regulated, high-stakes categories expose that difference the most.
The sources engines cite diverge even more than the brands they recommend
Brand overlap across engines is inconsistent. Source overlap is worse. In BrightEdge's citation-source analysis, the pages backing a brand mention overlapped only 16% to 59% of the time between engine pairs, a 43-point spread against the 19-point spread in brand overlap itself. Two engines can name the same company while supporting that mention with almost no shared evidence.
The mix of authority sources versus community content varies by engine, too, according to that same data. Gemini leans on established authority domains for 26% of its citations and on user-generated content for just 0.2%, roughly a 130-to-1 ratio. Google AI Overviews runs much closer to even, at 10% authority versus 18% UGC. That split decides where a brand's Citation Share actually sits. A company that has invested in third-party coverage on Reddit, LinkedIn, and G2 may look strong in one engine's citation mix and nearly invisible in another's.
Rewording a prompt barely moves the needle. Switching engines does.
Brand teams often worry that phrasing a question differently will change which brands an AI engine surfaces. Peec AI's analysis of 37,804 AI responses across five engines and 1,754 prompts, reported by Search Engine Journal in 2026, found that rewording matters far less than the underlying intent. When a reworded prompt kept a similar meaning, Brand Mention rates stayed close to baseline.
The effect only shows up once wording drifts enough to change the actual question being asked. At the lowest semantic-similarity range the study measured, average brand-mention probability fell from 4.9% to 2.5%, a 50% relative drop. List and ranking-style prompts also outperformed conversational phrasing by roughly 20%. None of that comes close to the 36%-to-55% gap between engines. Phrasing is a minor lever. Which engine answers the question is the major one.
What this means for tracking AI Share of Voice
A single-engine visibility check tells you about that one engine and almost nothing reliable about the others. With brand overlap between any two engines running 36% to 55%, and healthcare and finance showing even less agreement, sampling only ChatGPT or only Gemini means missing most of the picture, not a small slice of it.
Treating AI Share of Voice as one blended number compounds the problem, because an aggregate hides exactly the engine-level swings this research documents. Weighting each engine by where your buyers actually ask questions, and tracking brand mentions across ChatGPT, Gemini, and Perplexity separately rather than blended, turns this kind of research into something a brand can act on instead of something to worry about.
Frequently Asked Questions
Why do AI engines recommend different brands for the same query?
Each engine retrieves from its own index and weighs authority, recency, and user-generated content differently. BrightEdge found that citation-source overlap between engine pairs runs as low as 16%, so two engines can lean on almost entirely different evidence while reaching a different brand verdict for the same query.
Which industries show the most disagreement between AI engines on brand recommendations?
Healthcare and finance. BrightEdge found healthcare brands average only 60% agreement across engine pairs, compared with 97% for retail and 94% for travel. Individual healthcare brands can vary by more than nine-fold between engines, while finance brands average 71% agreement.
Does rewording a prompt change which brands an AI engine recommends?
Rarely, as long as the rewording keeps the same meaning. Peec AI's study of 37,804 AI responses found brand-mention rates stayed close to baseline across reworded prompts, and only dropped meaningfully, from 4.9% to 2.5%, once phrasing drifted into the lowest semantic-similarity range the researchers measured.
Should I track AI visibility across every engine or just the most popular one?
Every engine your buyers actually use. With brand overlap between any two engines running 36% to 55%, a single-engine check systematically misses how your brand performs everywhere else, and that gap is widest in regulated categories like healthcare and finance.
What's the difference between brand overlap and citation-source overlap across AI engines?
Brand overlap measures whether two engines name the same companies in their top rankings. Citation-source overlap measures whether they cite the same pages to support those mentions. BrightEdge found source overlap varies more, 16% to 59%, than brand overlap at 36% to 55%, so engines agree on winners more often than on evidence.



