ChatGPT vs Claude: which hallucinates less?
Last updated: July 26, 2026
On Vectara's HHEM summarization leaderboard (Nov 2025), both GPT-5 and Claude Sonnet 4.5 hallucinate above 10% on its harder dataset, and the gap between top models is small enough to treat as effectively tied. Neither reliably beats the other across tasks, so verify both. Which to trust depends on the task, not a single leaderboard number.
The honest short answer
There is no stable winner. On Vectara's HHEM leaderboard (announced November 19, 2025), which measures factual consistency on a summarization task, both GPT-5 and Claude Sonnet 4.5 exceed a 10% hallucination rate on Vectara's newer, harder dataset of over 7,700 long articles (Vectara, Nov 2025). Vectara notes this measures one summarization task, not all AI use, and small differences between closely ranked models should not be over-read.
So treat any headline "X hallucinates less than Y" claim with suspicion. Rankings shift with each model release and each benchmark, and the differences between frontier models are often smaller than the difference between an easy task and a hard one. The safe posture: verify the output regardless of which model produced it.
The numbers, dated and sourced
Below are the most recent public figures we could verify. They are task-specific and time-stamped on purpose; do not read them as a model's universal error rate.
| Benchmark (task) | Model | Hallucination rate | As of |
|---|---|---|---|
| Vectara HHEM (summarization, harder dataset) | GPT-5 | >10% | Nov 2025 |
| Vectara HHEM (summarization, harder dataset) | Claude Sonnet 4.5 | >10% | Nov 2025 |
| Stanford RegLab (legal research) | GPT-4 | ~43% | 2024 |
Two caveats. First, on Vectara's older, easier 1,000-document dataset leading models sat much lower (the best near 1%, most within roughly 1–5%), so the same models look very different across datasets; the task drives the number more than the model does. Second, the Stanford GPT-4 figure is legal-domain and over a year old, included here only to show how high rates climb on hard, high-stakes tasks (Stanford RegLab, 2024). We are not aware of a directly comparable dated Claude figure on the same legal test, so we do not assert one.
Characteristic failure modes
The interesting difference is not a single percentage but how each family tends to fail, which affects what you should check. These are tendencies observed across benchmarks and usage, not laws, and both models exhibit both patterns.
- Confident specifics. When a model lacks a fact, it may still produce a precise-looking name, date, or citation. This is the failure behind fabricated legal cases like Mata v. Avianca, and it is why you should scrutinize any exact figure regardless of the model.
- Misgrounding. A model can cite a real source that does not actually support the claim, the subtle failure the Stanford study highlighted. Reasoning and retrieval features reduce but do not remove it.
- Sycophancy under pressure. Both families can reverse a correct answer if you push back, because they optimize for a satisfying response, not for defending a fact.
The practical upshot: the failure modes overlap enough that a single verification step matters more than picking the "safer" model.
When to trust which, and why to verify either way
A reasonable heuristic, held loosely: for factual summarization of source documents, pick whichever model tops the current Vectara HHEM leaderboard for that task and always ground it in the source text. For open-ended factual questions with no supplied source, treat both as equally capable of confident error and verify every specific claim. Neither model should be relied on unverified for legal, medical, or financial stakes.
Because rankings move, revisit the leaderboard periodically rather than trusting a number you saw once; a model that led last quarter may not lead this one. Whatever you choose, run the output through a verifier. Paste any ChatGPT or Claude answer into /chat, and Mephistopheles grounds each claim against independent sources and returns a per-claim verdict plus a risk score. See the model-specific pages on ChatGPT hallucinations and Claude hallucinations.
Why "two models agreeing" is not proof
Running the same question through GPT-5 and Claude and getting the same answer feels reassuring, but agreement is not verification. Models trained on overlapping data can share the same wrong belief. Only an independent source outside the models confirms a fact.
Frequently asked questions
Does ChatGPT or Claude hallucinate less?
Neither reliably. On Vectara's HHEM summarization leaderboard (Nov 2025), both GPT-5 and Claude Sonnet 4.5 exceed 10% on the harder dataset, and they rank close enough to treat as effectively tied. Rankings also shift with each release, so verify the output rather than relying on one model being safer.
What is the hallucination rate of GPT-5 and Claude?
It depends entirely on the task. On Vectara's harder HHEM summarization dataset (Nov 2025), both exceeded 10%, while on Vectara's older, easier dataset leading models sat much lower (the best near 1%). There is no single universal rate; the task drives the number more than the model does.
Is Claude more accurate than ChatGPT?
Not consistently. Across public benchmarks the two trade places depending on task and dataset, and the differences between frontier models are often small. Treat 'more accurate' claims as task-specific and dated, and verify important claims regardless of model.
If ChatGPT and Claude give the same answer, is it correct?
Not necessarily. Models trained on overlapping data can share the same mistaken belief, so agreement is not evidence. Confirm the claim against an independent source outside both models before relying on it.
Verify what your AI just told you.
Paste any AI answer and Mephistopheles checks each claim against independent sources — no sign-up to try.