Mephistopheles
HomeHow it worksFact-check ChatGPT
How-to guide

How to fact-check ChatGPT answers

Last updated: July 26, 2026

To fact-check ChatGPT, treat every answer as unverified. Ask for its sources, then open those sources yourself and read what other sites say about them (lateral reading). Confirm each factual claim against an independent authority. Never ask ChatGPT to confirm itself, because it can restate a fabrication with full confidence.

Why ChatGPT answers need checking at all

ChatGPT predicts likely next words. It does not look up facts and it has no internal sense of true versus false, so a confident, fluent sentence can still be wrong. This failure mode is called a hallucination, and it happens most on names, dates, numbers, quotes, and citations.

The rates are not trivial. In a 2024 preregistered Stanford RegLab study, GPT-4 produced hallucinations on roughly 43% of legal research queries, and even purpose-built legal tools hallucinated on 17% (Lexis+ AI) to 33% (Westlaw AI-Assisted Research) of queries (Stanford RegLab, 2024). That study is legal-domain and now over a year old, so read it as a magnitude signal, not a universal rate. The wider lesson holds: you cannot assume any single answer is correct.

The most public example is Mata v. Avianca (2023), where a lawyer submitted a brief citing six court cases that ChatGPT had entirely invented. The attorneys were sanctioned $5,000 (Mata v. Avianca, Inc.). Tellingly, when the lawyer asked ChatGPT whether the cases were real, it insisted they were. That is the exact trap the manual method below is built to avoid.

The honest manual method: how to fact-check ChatGPT by hand

You can fact-check ChatGPT with no tools at all. The technique professional fact-checkers use is lateral reading, formalized as the SIFT method (Stop, Investigate the source, Find better coverage, Trace claims to the original) by digital-literacy researcher Mike Caulfield (UC Merced Library, on Caulfield's SIFT). Adapted for an AI answer, it looks like this.

  1. Stop and isolate the claims. Break the answer into individual factual statements. "Company X was founded in 1998 by Y" is two claims: the year and the founder. Vague opinions do not need checking; specific, checkable facts do.
  2. Ask ChatGPT for its sources, then leave. Prompt it: "List the specific sources for each claim, with titles and URLs." Use these only as leads. Do not trust that a cited source exists or that it says what ChatGPT claims. A real citation is not a supported claim.
  3. Investigate the source laterally. Open a new tab and search for what other reputable sites say about each source, rather than judging it from its own page. If ChatGPT cites a study, find the study and independent coverage of it.
  4. Find better coverage. For each claim, look for an independent, authoritative source that confirms it: the primary document, an official registry, a reference work, or two unrelated reputable outlets that agree.
  5. Trace to the original. Follow quotes and statistics back to where they first appeared. Numbers mutate as they are re-reported, so confirm the figure, the date, and the measurement definition at the source.

If you cannot find any independent source for a claim, the correct verdict is unverifiable, not "probably true." Absence of confirmation is itself a finding.

Why "just ask ChatGPT if it's true" does not work

Asking a model to grade its own answer feels efficient and fails for a structural reason: the model has no ground truth to check against. It generates a plausible response to "is this true?" the same way it generated the original claim, so it can restate a fabrication with total confidence, or reverse a correct answer under mild pushback.

This is exactly what happened in Mata v. Avianca: asked to confirm the fake cases, ChatGPT confirmed them. Self-consistency is not evidence. For the same reason, two different models agreeing is not proof either; they can share the same wrong belief. Real verification requires an independent source outside the model.

The one-click upgrade: verify every claim at once

The manual method is trustworthy but slow. Mephistopheles automates the same discipline. Paste any ChatGPT answer into /chat (or verify in place with the browser extension), and it extracts each factual claim, grounds each one against independent sources (semantic retrieval, web, Wikidata and Wikipedia, and tools), runs a tiered LLM judge, and returns a per-claim verdict plus a hallucination-risk score.

Every claim lands in one of four buckets in the verdict taxonomy:

VerdictWhat it meansWhat to do
SupportedAn independent source backs the claimReasonable to rely on; keep the source
ContradictedA source directly conflicts with itFix or remove it
DisputedReputable sources disagreePresent both sides; do not assert
UnverifiableNo authoritative source foundFlag for a human; do not publish as fact

On our own private benchmark of about 290 deliberately hard claims across roughly 20 domains, Mephistopheles catches about 88% of factual errors (each with a contradicting source) and passes about 98% of true statements clean, a roughly 2% false-alarm rate. It is a detector, not an oracle; see how accurate it is and where it says unverifiable rather than guess.

Frequently asked questions

How do I fact-check ChatGPT for free?

Break the answer into individual claims, ask ChatGPT to list its sources, then verify each claim yourself by reading independent, reputable sources (lateral reading / the SIFT method). Do not trust the model to confirm its own output. This costs nothing but time. You can also paste an answer into /chat to check many claims at once.

Can ChatGPT fact-check itself?

No, not reliably. A model has no ground truth to compare against, so it can confidently restate a fabrication or reverse a correct answer under pressure. In Mata v. Avianca, ChatGPT confirmed six court cases it had invented. Verification needs an independent source outside the model.

What is the SIFT method for checking AI answers?

SIFT stands for Stop, Investigate the source, Find better coverage, and Trace claims to the original. It was created by digital-literacy researcher Mike Caulfield and is used by professional fact-checkers. Applied to an AI answer, you isolate each claim and confirm it against independent, authoritative sources instead of trusting the model.

How often does ChatGPT get facts wrong?

It varies sharply by task and topic, so no single number applies. As a magnitude signal, a 2024 Stanford RegLab study found GPT-4 hallucinated on roughly 43% of legal-research queries. Hallucinations cluster on names, dates, numbers, quotes, and citations, so those deserve the most scrutiny.

Verify what your AI just told you.

Paste any AI answer and Mephistopheles checks each claim against independent sources — no sign-up to try.

Verify an answer →