Featured on SaaSBison Featured on Toolfio Listed on Bowora Featured on Uneed Featured on ToolPilot

How often AI engines get brand facts wrong, and who's liable when they do

AI search engines get citations wrong more than 60% of the time, according to Columbia's Tow Center for Digital Journalism, and a Canadian tribunal has already ruled a company liable for what its chatbot invented. Here's what the data shows about brand-specific AI hallucinations, and how to catch one before a customer does.

ChatGPT misattributed article sources in 67% of test queries, according to a widely cited Tow Center for Digital Journalism study from Columbia University. Grok 3 got it wrong 94% of the time and sent users to broken links in more than three-quarters of its citations. Those numbers describe how AI engines handle news articles, not brand facts specifically, but the mechanism is identical: when a model doesn't have a clean answer on hand, it fills the gap with something plausible. For a brand, that gap can become a fabricated return policy, an outdated price, or a discontinued feature described as current, and a British Columbia tribunal has already ruled a company liable for exactly that kind of error.

How often do AI engines actually get brand facts wrong?

Columbia's Tow Center tested eight AI search engines with 1,600 queries total and found an overall citation failure rate above 60%. Perplexity performed best at a 37% error rate. Grok 3 performed worst at 94%. Every engine in the study either misattributed a source, invented one, or linked to a page that no longer existed.

Researchers Klaudia Jaźwińska and Aisvarya Chandrasekar gave each engine a direct quote from a news article and asked it to identify the title, publisher, date, and URL. ChatGPT got 134 of 200 wrong but flagged its own uncertainty only 15 times. Gemini and Grok 3 fabricated more broken links than working ones. Perplexity Pro, despite a licensing partnership with the Texas Tribune, cited syndicated copies of Tribune stories instead of the original in three out of ten tests, a gap between what a model can access and what it actually surfaces that runs through how ChatGPT chooses its sources in the first place.

Brand facts follow the same failure pattern with a narrower blast radius. A pricing page, a return policy, a spec sheet: these are exactly the single-source, rarely updated pages a model has to reconstruct from partial training data or a rushed live fetch, and reconstruction is where citation accuracy breaks down.

Why hallucination rates are falling while brand-specific errors keep happening

General hallucination rates have dropped sharply. Vectara's hallucination leaderboard, an ongoing benchmark tracking model performance since 2023, put several current small models below 3.5% and most flagship models under 11% as of its May 2026 update. That number measures something narrower than most people assume, though: whether a model stays faithful to a document it was just handed, not whether it recalls a brand's actual current policy from memory.

That distinction matters. Vectara's benchmark feeds each model a real document and grades it on factual consistency with that specific document. A brand fact check is the opposite scenario: an AI engine answering a question with no source document in front of it, drawing instead on training data that might be eighteen months stale or a live fetch that missed the one page that actually matters.

A hallucination isn't a lie. It's a model filling a gap in what it knows with something that sounds plausible.

Courts have already ruled that companies own what their chatbot says

Yes, and the precedent is specific. In February 2024, a British Columbia Civil Resolution Tribunal ordered Air Canada to pay a passenger after its support chatbot invented a bereavement fare discount that didn't exist. The airline argued the chatbot was a separate legal entity responsible for its own statements. Tribunal member Christopher Rivers rejected that outright, calling it “a remarkable submission.”

Moffatt v. Air Canada is now the reference case for chatbot liability, and it set a plain standard: a company answers for what its AI support tools tell customers, whether the model made an honest error or filled a knowledge gap with something invented. A similar pattern played out at Cursor in April 2025, when its support chatbot told a user the company enforced a single-device login policy that didn't exist. Co-founder Michael Truell confirmed on Reddit that no such policy had ever been written, but the fabricated rule had already spread through developer forums before the correction caught up. That reputational fallout compounds the legal one; we've written before about how AI-driven brand risk shows up long before a lawsuit does.

Third-party sources make brand misinformation harder to catch

Most AI citations about a brand don't come from that brand's own website. Third-party sources such as Reddit threads, LinkedIn posts, and G2 reviews make up the large majority of what AI engines cite about a company, as we've covered before, which means outdated or wrong information sitting on a forum can outrank a brand's current, accurate page as a source.

That echo effect compounds the hallucination problem. If an engine already leans on third-party threads instead of a brand's own pages, a stale complaint from two product versions ago can keep surfacing as the answer to a pricing or policy question long after the brand actually fixed it.

How to catch a hallucinated brand claim before a customer does

Run the exact questions a real customer would ask, across ChatGPT, Claude, Gemini, and Perplexity, on a fixed schedule, and compare each answer against a brand's actual current policy, pricing, and specs. A single spot check misses drift. A hallucinated claim tends to appear, get repeated across a few sessions, and then either self-correct as models retrain or persist for months if nothing challenges it.

A brand that only reacts when a customer forwards a screenshot is always working from a stale signal. Tracking mention accuracy alongside share of voice and sentiment, the way our guide to tracking brand mentions across ChatGPT, Gemini, and Perplexity lays out, catches a fabricated policy while it's still an isolated answer rather than a version three engines already agree on. VizibleAI checks brand mentions across those four engines on a recurring schedule and flags when an engine's version of a policy, price, or spec no longer matches what the brand has actually published.

Frequently Asked Questions

What counts as an AI hallucination about a brand?

A brand hallucination is any answer where an AI engine states a policy, price, feature, or fact about a company that isn't true, whether the model invented it outright or filled a gap in outdated training data with something plausible. The Air Canada and Cursor cases are both examples: neither company had the fabricated policy on record anywhere.

How often do AI search engines cite sources correctly?

Not often, according to the most detailed public study available. Columbia's Tow Center for Digital Journalism found AI search engines failed to cite sources correctly in more than 60% of 1,600 test queries across eight engines, ranging from a 37% error rate at Perplexity to 94% at Grok 3.

Can a company be held legally liable for what its chatbot says?

Yes. A British Columbia tribunal ruled in Moffatt v. Air Canada that a company is responsible for its chatbot's statements even when the chatbot invented the claim itself, rejecting Air Canada's argument that the bot was a separate legal entity. The ruling is now cited as the reference case for AI chatbot liability.

Does a lower hallucination-benchmark score mean an AI engine gets brand facts right?

Not necessarily. Benchmarks like Vectara's leaderboard measure whether a model stays faithful to a document it was just given, which is a different task from recalling a brand's current policy from memory or a live web fetch. A model can score well on document-summarization fidelity and still misstate a company's return policy.

How can a brand tell if an AI engine is giving wrong information about it?

Run the exact questions a customer would ask, such as pricing, refund policy, or product specs, across ChatGPT, Claude, Gemini, and Perplexity on a recurring schedule, then compare each answer to what the brand has actually published. A one-time check misses drift; hallucinated claims tend to persist until something corrects them.

Do third-party sites make brand hallucinations more likely?

Yes. Most AI citations about a brand pull from third-party sources like Reddit, G2, and LinkedIn rather than the brand's own site, so outdated or inaccurate information sitting on those platforms can keep surfacing as an AI engine's answer long after a company has corrected the record on its own pages.

Finding out what AI engines say about your brand before your customers do

Catching a hallucinated policy or price after a customer has already acted on it is expensive, and per the Air Canada ruling, it can be a legal problem as much as a reputational one. Pairing that kind of check with a full GEO audit shows whether an error is a one-off answer or a pattern repeating across engines. See what ChatGPT, Claude, Gemini, and Perplexity are already saying about your brand with VizibleAI.