Featured on SaaSBison Featured on Toolfio Listed on Bowora Featured on Uneed Featured on ToolPilot

Google's AI Overviews criticize brands 44% more than ChatGPT, new data shows

A March 2026 BrightEdge analysis found Google's AI Overviews are 44% more likely to say something negative about a brand than ChatGPT. The more useful finding is where that negativity lands: ChatGPT surfaces it far closer to the purchase decision, even though it happens less often overall.

Google's AI Overviews are 44% more likely to say something critical about a brand than ChatGPT is, according to a March 2026 analysis from BrightEdge covering prompts across apparel, electronics, and education. The headline gap sounds small in isolation: negative sentiment showed up in about 2.3% of Google's AI Overview mentions versus 1.6% for ChatGPT. The study's more useful finding sits underneath that number. ChatGPT was roughly 13 times more likely to surface negative information right before a purchase decision, which is the moment a brand can least afford it.

Why Google's AI Overviews criticize brands more than ChatGPT does

Google's AI Overviews pull from a wider, less curated set of web sources than ChatGPT does by default, and that width is most of what drives the criticism gap. BrightEdge found Google surfaced negative brand sentiment in 2.3% of mentions against ChatGPT's 1.6%, a 44% relative difference across the three industries it measured.

BrightEdge, an enterprise SEO platform, pulled the figures from its AI Catalyst monitoring product, which tracks brand mentions across search and AI answer engines for its customers. Comparing the same queries across both engines turned up a second problem: Google and ChatGPT named different brands as category leaders on 73% of overlapping queries. A brand that already shows up inconsistently between engines is effectively managing two separate reputations, and if one of them skews negative, the brand may never find out without checking each engine on its own.

BrightEdge CEO Jim Yu framed the finding bluntly:

For better or worse, AI is your brand's new editorialist. Each engine characterizes your brand differently.

Google pushed back on the framing, telling Fortune that a roughly one-percentage-point gap between AI Overviews and ChatGPT was negligible. That's a reasonable objection to the headline number: 2.3% versus 1.6% is a small absolute difference even if the relative comparison sounds dramatic. What the rebuttal doesn't address is where in the buying journey that gap actually lands, which turns out to matter more than the rate itself.

The bigger risk isn't the rate, it's where negativity shows up in the funnel

The rate of negative AI mentions matters less than the stage of the buying journey where they appear. BrightEdge found Google's negative sentiment clustered heavily at the informational stage, 85% of the time, while ChatGPT pushed negative material much closer to purchase: 19.4% of its negative mentions showed up during consideration-to-purchase queries, and it was about 13 times more likely than Google to do so.

Informational-stage criticism, someone asking a general question about a category or comparing options early on, is friction most brands can absorb. Purchase-stage criticism, someone asking which vendor to pick or whether a specific product holds up, is the moment a monitoring gap turns into a lost deal. A team that only tracks its overall AI Share of Voice number, without breaking mentions down by query intent, would see ChatGPT's lower aggregate negative rate and assume it's the safer engine to watch less closely. The funnel-stage data says close to the opposite for any brand that converts through ChatGPT.

Why the engines disagree: grounding, not brand reality

Much of the disagreement between engines traces back to grounding, meaning whether an answer is built from live web sources retrieved at the moment of the query or drawn from the model's static training data. Engines that ground more heavily pull in more recent reviews, news, and forum posts, which shifts sentiment results independent of anything the brand itself has done.

An independent measurement from cloro tracked grounding rates across engines twice during 2026, once in July and again in August. ChatGPT's grounding rate, the share of answers citing at least one live source, fell from 98.4% to 78.9% over that single month, while Gemini's climbed from 41.1% to 70.7% as Google expanded when its consumer product triggers a live search. Perplexity was the most consistently grounded engine across both measurements, staying above 94% each time and citing between 9.5 and 13 sources per answer.

None of these numbers are fixed. Vendors tune grounding behavior on a rolling basis, which is exactly why a brand's AI sentiment profile can shift from one month to the next with no change in what the brand actually did. An engine that grounds less relies more on whatever pattern of brand mentions sat in its training data to begin with, a separate and slower-moving problem covered in more detail in our look at AI knowledge cutoff lag.

What this means for AI sentiment monitoring

A single blended sentiment score across engines hides the two risks that matter most: which engine is being harsher for structural reasons, and which specific query types carry a brand's negative mentions. Monitoring by engine and by funnel stage, instead of as one aggregate number, is what turns this data into something a team can act on.

Third-party sources are usually the actual origin of negative material, not something the AI engine invented on its own. Reddit threads, G2 reviews, and LinkedIn discussion already account for most of what AI engines cite about a brand, so a sentiment spike is often a reason to check what's being said on those platforms directly before assuming the engine is at fault.

Some of what reads as unfair criticism is a factual error rather than a founded complaint, and the fix for that is different: correcting the record, not managing sentiment. Brands that already track mentions across ChatGPT, Gemini, and Perplexity have the shortest path to adding sentiment-by-engine and sentiment-by-funnel-stage as a second layer on existing tracking, rather than standing up a separate system. AI brand sentiment deserves the same regular review as any other reputation channel, not a one-time audit after a bad news cycle.

Frequently Asked Questions

Why do Google's AI Overviews criticize brands more than ChatGPT?

BrightEdge's March 2026 analysis found Google's AI Overviews showed negative brand sentiment in about 2.3% of mentions versus 1.6% for ChatGPT, a 44% relative gap. The difference traces back to how each engine grounds its answers: Overviews pull from a wider, less curated set of live web sources, which surfaces more of the negative reviews and news already circulating about a brand.

Is a difference like 2.3% versus 1.6% actually significant?

Google has disputed the framing, calling a roughly one-percentage-point gap negligible, which is fair on the raw numbers. The more concrete risk BrightEdge found is timing, not rate: ChatGPT was about 13 times more likely than Google to surface negative information right before a purchase decision, which matters more to revenue than the overall percentage.

What does grounding mean for AI brand sentiment?

Grounding is whether an AI answer is built from live web sources fetched at the moment of the query or from the model's static training data. Engines that ground more heavily, like Perplexity, pull in more recent reviews and news, which shifts sentiment results as those sources change, independent of anything the brand itself has done.

Should brands monitor sentiment separately for each AI engine?

Yes. A blended sentiment score across engines hides which one is driving negative mentions and why, and BrightEdge's data shows the two major engines disagreed on which brands led a category on 73% of overlapping queries. Monitoring by engine, and by whether a query is informational or purchase-stage, surfaces risks an aggregate score misses entirely.

Where does negative AI sentiment about a brand usually come from?

Most of it traces back to third-party sources the AI engine is citing, not something the engine invented. Reddit threads, G2 reviews, and LinkedIn discussion already make up most of what AI engines cite about brands, so a sentiment spike is usually a reason to check what's being said on those platforms directly.