Does ChatGPT hallucinate? What the data shows
Last updated: July 26, 2026
Yes. ChatGPT hallucinates, most dangerously by inventing confident, plausible-looking citations. On Vectara's May 2026 summarization leaderboard, GPT-4o hallucinated at 9.6% and GPT-5.2 at 10.8%. In a 2023 study, GPT-4 fabricated about 18% of academic citations outright. It rarely signals when it is guessing.
Does ChatGPT hallucinate, and how often?
Yes, ChatGPT hallucinates, and the rate depends heavily on how you measure it. There is no single number, so treat any headline figure as domain- and task-specific.
On a grounded summarization task, ChatGPT is relatively strong. In Vectara's Hallucination Leaderboard (last updated May 11, 2026, using more than 7,700 documents), GPT-4o hallucinated on 9.6% of summaries and the reasoning model GPT-5.2 on 10.8% (Vectara, 2026). That benchmark measures whether a summary stays faithful to a supplied passage, not open-ended factual recall.
Open-ended tasks are far riskier. In Stanford RegLab's 2024 study Large Legal Fictions, GPT-4 produced a legal hallucination on 58% of verifiable legal queries (Dahl et al., Journal of Legal Analysis, 2024). These are different tasks measured by different teams; do not average them.
Why we cite a range, not one rate
Hallucination rate is not a fixed property of a model. It shifts with prompt, domain, and how "hallucination" is defined. We refresh these figures quarterly; the numbers above reflect studies published through mid-2026.
ChatGPT's signature failure: the confident fake citation
ChatGPT's most characteristic failure is the fabricated citation delivered with total confidence. It does not hedge, and the invented source is formatted perfectly.
In a 2023 study by Walters and Wilder (Scientific Reports) across 42 topics, GPT-4 fabricated about 18% of its citations entirely, and of the citations that referred to real work, roughly 24% still contained substantive errors such as a wrong volume, page, or author (GPT-3.5 was worse, at 55% fabricated). This is the pattern behind the widely reported Mata v. Avianca sanctions in 2023, where a lawyer filed a brief citing six cases ChatGPT had invented.
Worked example. Ask ChatGPT for "three peer-reviewed studies on intermittent fasting and cognition," and it may return a clean list: authors, a real-sounding journal, a year, a DOI. Two of the three papers do not exist. The author names are real researchers in the field; the titles and DOIs are synthesised. Nothing in the output signals which one is fake, which is exactly what makes it dangerous.
How to fix ChatGPT hallucinations: manual and automated
You reduce ChatGPT hallucinations by verifying independently, never by asking ChatGPT to check itself. A model that fabricated a citation will often defend it when asked "are you sure?"
Manual checks:
- Open every citation. A DOI that resolves to a different paper, or a case number that does not exist on the court's site, is a fabrication.
- Confirm the source actually contains the claim. A real citation is not a supported claim — the paper can exist yet not back the sentence.
- Ask for the source before the answer, so you can reject unsupported claims up front.
Automated with Mephistopheles: paste any ChatGPT answer into /chat. Mephistopheles extracts each factual claim, grounds it against independent sources, and returns a per-claim verdict — supported, contradicted, disputed, or unverifiable — plus a hallucination-risk score. On our own benchmark of roughly 290 deliberately hard claims, it catches about 88% of factual errors while passing about 98% of true statements clean (a ~2% false-alarm rate). It flags what it cannot verify instead of guessing.
Frequently asked questions
How often does ChatGPT make up citations?
In a 2023 study by Walters and Wilder, GPT-4 fabricated about 18% of its citations entirely, and about 24% of the real ones still had substantive errors. Later GPT versions have improved on some benchmarks, but no public data shows citation fabrication has dropped to zero. Always open and verify each source.
Is GPT-5 more accurate than GPT-4?
It depends on the task. On Vectara's May 2026 summarization leaderboard, GPT-4o scored 9.6% hallucination and the GPT-5.2 reasoning model scored 10.8% — slightly worse on that specific faithfulness test, because reasoning models tend to add unsupported detail. Newer is not automatically more grounded. Verify either way.
Can I ask ChatGPT to fact-check its own answer?
Not reliably. Because ChatGPT often does not know when it is wrong, asking "are you sure?" frequently produces a confident re-assertion of the same error, or a new fabrication. Verification needs an independent source, which is what Mephistopheles provides.
Verify what your AI just told you.
Paste any AI answer and Mephistopheles checks each claim against independent sources — no sign-up to try.