Mephistopheles
HomeDoes AI hallucinate?Grok hallucinations
Model hallucination profile

Does Grok hallucinate? What the data shows

Last updated: July 26, 2026

Yes. Grok hallucinates, and its search mode has been notably error-prone: a March 2025 Tow Center study found Grok-3 Search gave incorrect answers on 94% of tested queries, the worst of eight engines. xAI reports Grok 4.1 cut its own web-quote hallucination rate from 12.09% to 4.22%. Verify regardless.

Does Grok hallucinate, and how often?

Yes, Grok hallucinates, and the picture varies sharply by version and by whether you use its search mode.

On Vectara's Hallucination Leaderboard (updated May 11, 2026), Grok-3 scored 5.8% hallucination on grounded summarization — competitive (Vectara, 2026). But on live web search, results were far worse: the Tow Center / Columbia Journalism Review study (March 2025, 1,600 queries across eight engines) found Grok-3 Search answered incorrectly on 94% of queries, the highest failure rate tested (CJR, 2025).

xAI has since claimed improvement: it reported that Grok 4.1 cut its web-quote hallucination rate from 12.09% to 4.22% on its internal evaluation (VentureBeat, Nov 2025). That is a vendor's own benchmark, so treat it as a directional claim, not an independent result.

Vendor numbers vs independent numbers

The 5.8% (Vectara) and 94% (Tow Center) figures come from independent third parties on defined tasks. The 12.09% → 4.22% figure is xAI's own. We flag the source so you can weight them accordingly — a self-reported improvement is not the same as an independent audit.

Grok's signature failure: confident real-time claims from noisy sources

Grok's most characteristic failure comes from its tight coupling to real-time posts on X. It answers current-events questions with speed and confidence, but it can amplify unverified or false claims circulating on the platform as if they were established fact.

The 2025 Tow Center finding — 94% incorrect on tested queries in Grok-3 Search — reflects this: fast, confident answers that frequently cite the wrong source or fabricate the attribution. Grok has also produced high-profile off-the-rails outputs, underscoring weak guardrails around contested or trending topics.

Worked example. During a breaking news event, ask Grok "what caused the outage?" and it may synthesise an answer from trending posts — naming a specific company, cause, and timeline — before any of that is confirmed. Hours later the official cause is different. The original answer was fluent, specific, and wrong, drawn from social speculation dressed up as reporting.

How to fix Grok hallucinations: manual and automated

The fix for Grok is to distrust speed on unsettled topics and to verify every real-time claim against a primary source.

Manual checks:

  • For breaking news, wait for a primary or official source; treat Grok's early answer as a lead, not a fact.
  • Check whether cited posts are from credible accounts or from speculation, and whether the post actually says what Grok claims.
  • Be skeptical of confident specifics — exact numbers, names, and timelines — on trending topics.

Automated with Mephistopheles: paste any Grok answer into /chat. Mephistopheles grounds each claim against independent, authoritative sources rather than the social feed Grok drew from, and returns a per-claim verdict with a hallucination-risk score. On our ~290-claim benchmark it catches ~88% of factual errors and passes ~98% of true statements clean. When a fast-moving claim has no authoritative source yet, it returns "unverifiable" instead of guessing — which is often the honest answer during a breaking event.

Frequently asked questions

Is Grok's search mode accurate?

It has been the weakest of the major engines tested. The March 2025 Tow Center / Columbia study found Grok-3 Search answered incorrectly on 94% of queries — the worst of eight AI search engines. xAI says newer versions improved, but on that independent test Grok's live search was highly unreliable. Verify every result.

Did Grok 4.1 fix hallucinations?

xAI reported that Grok 4.1 cut its web-quote hallucination rate from 12.09% to 4.22% on its own internal benchmark (Nov 2025). That is a meaningful directional improvement, but it is a vendor-reported figure, not an independent audit, and it does not cover all task types. Continue to verify.

Why does Grok repeat false information from social media?

Grok is closely coupled to real-time posts on X, so on trending or breaking topics it can restate unverified speculation as fact. Speed is its selling point and its weakness. Treat its live-event answers as leads to confirm against primary sources.

How do I verify a Grok answer?

Paste it into Mephistopheles, which grounds each claim against independent authoritative sources rather than the social feed, returns a verdict and risk score, and marks unconfirmed breaking-news claims as unverifiable.

Verify what your AI just told you.

Paste any AI answer and Mephistopheles checks each claim against independent sources — no sign-up to try.

Verify an answer →