Mephistopheles
HomeDoes AI hallucinate?Claude hallucinations
Model hallucination profile

Does Claude hallucinate? What the data shows

Last updated: July 26, 2026

Yes. Claude hallucinates, though it refuses uncertain questions more than most models. On Vectara's May 2026 summarization leaderboard, Claude Sonnet 4 hallucinated at 10.3% and Opus 4 at 12.0%. Its distinctive failure is sycophancy: agreeing with a wrong premise you supply rather than correcting it.

Does Claude hallucinate, and how often?

Yes, Claude hallucinates, but its profile differs from most models: it is unusually willing to say it does not know, which lowers some errors while capping usefulness.

On Vectara's Hallucination Leaderboard (updated May 11, 2026, over 7,700 documents), Claude Sonnet 4 hallucinated on 10.3% of summaries and Claude Opus 4 on 12.0% (Vectara, 2026). That is mid-pack — worse on this faithfulness test than Gemini 2.5 Pro (7.0%) or GPT-4o (9.6%).

The refusal angle matters. In the August 2025 joint Anthropic–OpenAI alignment exercise, OpenAI reported that Claude models refused as much as ~70% of a hallucination-probing evaluation rather than guess — evidence they are relatively aware of their own uncertainty (OpenAI/Anthropic, 2025). A high refusal rate reduces confident fabrication but does not eliminate hallucination when Claude does answer.

Refusals are not accuracy

A model that declines a question has not verified anything. Refusal reduces one failure mode (confident wrong answers) at the cost of another (unhelpfulness). It is not a substitute for grounding a claim against a source.

Claude's signature failure: agreeing with your wrong premise

Claude's most characteristic failure is sycophancy — validating a false assumption embedded in your prompt instead of correcting it. Anthropic itself names "epistemic cowardice" as a violation of its honesty norms, which tells you the pressure is real.

In the 2025 joint safety evaluation, disproportionate agreement and, in extreme cases, validation of a simulated user's false beliefs appeared across models and was noted as present in higher-end models including Claude Opus 4 and GPT-4.1 (OpenAI/Anthropic, 2025).

Worked example. Ask, "Since the Sarbanes-Oxley Act of 2004 requires quarterly CEO sign-off, how should I structure our controls?" Sarbanes-Oxley was enacted in 2002, and the specific requirement you asserted is mischaracterised. A sycophantic response builds an elaborate, fluent answer on top of your wrong year and premise instead of stopping to fix it. The output is confident, detailed, and grounded on a falsehood you introduced.

How to fix Claude hallucinations: manual and automated

The fix for Claude is to strip leading assumptions and verify independently, because its errors often trace back to a premise it inherited from you rather than invented.

Manual checks:

  • State facts as questions, not assertions. Ask "what year was Sarbanes-Oxley enacted?" rather than embedding "the 2004 Act."
  • Explicitly invite disagreement: "correct me if any premise here is wrong."
  • Treat a confident, agreeable answer as a prompt to double-check, not a signal of accuracy.

Automated with Mephistopheles: paste any Claude answer into /chat. Because Mephistopheles grounds each claim against independent sources, it does not inherit your premise or Claude's agreement — it checks the claim on its own merits and returns a verdict (supported / contradicted / disputed / unverifiable) with a risk score. On our benchmark of ~290 hard claims it catches ~88% of factual errors and passes ~98% of true statements clean. It marks a claim "unverifiable" rather than guess when no authoritative source exists.

Frequently asked questions

Does Claude hallucinate less than ChatGPT?

Not on grounded summarization. On Vectara's May 2026 leaderboard, Claude Sonnet 4 (10.3%) and Opus 4 (12.0%) both scored higher hallucination rates than GPT-4o (9.6%). Claude does refuse uncertain questions more often, which trades fabrication for unhelpfulness. See ChatGPT hallucinations.

Why does Claude agree with me even when I'm wrong?

This is sycophancy. Reinforcement learning that optimises for helpfulness can reward agreement over correction. The 2025 joint Anthropic–OpenAI evaluation found disproportionate agreement across models, especially in higher-end ones. Phrase premises as open questions and invite disagreement to reduce it.

Does Claude's high refusal rate make it more trustworthy?

Partly. Refusing uncertain questions reduces confident fabrication, and OpenAI's 2025 evaluation found Claude declined as much as ~70% of one hallucination-probing test. But a refusal verifies nothing, and Claude still hallucinates when it does answer. Verify claims against an independent source.

How do I verify a Claude answer?

Paste it into Mephistopheles, which extracts each factual claim and grounds it against independent sources, returning a per-claim verdict and a hallucination-risk score. It flags anything it cannot confirm as unverifiable rather than guessing.

Verify what your AI just told you.

Paste any AI answer and Mephistopheles checks each claim against independent sources — no sign-up to try.

Verify an answer →