Hallucination-risk score
Last updated: July 26, 2026
A hallucination-risk score is a single number summarising how likely an AI answer contains unsupported or false claims. Mephistopheles builds it from per-claim verdicts: an answer full of contradicted or unverifiable claims scores high risk; one where every claim is independently supported scores low.
What is a hallucination-risk score?
A hallucination-risk score is a single figure that summarises how much of an AI answer looks unsupported. It is a rollup, not a magic gauge: Mephistopheles first breaks the answer into claims, assigns each a verdict, then aggregates those verdicts into one score.
The point is triage. When you cannot read every source yourself, the score tells you whether an answer is safe to skim or needs line-by-line review. A high score means several claims were contradicted or could not be verified; a low score means the checkable claims held up against independent sources.
The term is coined by Mephistopheles. It is only as trustworthy as the verdicts underneath it, which is why the score always ships with the per-claim breakdown — you can see exactly which claim drove the risk up.
The four verdicts behind the score
Every claim gets one of four verdicts, and the score is built from them. See the full verdict taxonomy.
| Verdict | Meaning | One-line example |
|---|---|---|
| Supported | An independent source backs the claim. | "Water boils at 100°C at sea level" — confirmed by physics references. |
| Contradicted | An independent source refutes it. | "The Eiffel Tower is in London" — every source says Paris. |
| Disputed | Credible sources disagree. | "Coffee is good for your heart" — studies conflict. |
| Unverifiable | No authoritative source found. | "The CEO privately told staff X on Tuesday" — nothing public to check. |
Contradicted claims push risk up hardest; unverifiable claims raise it more mildly, because 'we couldn't check' is a caution, not proof of error. See unverifiable.
A worked example
Suppose an AI answer makes six claims. Four are supported, one is disputed, and one — a fabricated statistic — is contradicted. The score reflects a moderate-to-high risk driven mostly by that one contradicted claim, and the breakdown points you straight to it.
Because the score is only a summary, always read the flagged claims. A single contradicted number in an otherwise solid answer is exactly the thing you want surfaced, not averaged away. Paste an answer into the verifier to see a live score and its breakdown.
Frequently asked questions
Is a low hallucination-risk score a guarantee the answer is true?
No. A low score means the checkable claims were independently supported and none were contradicted. It is strong evidence, not a guarantee — some claims may be unverifiable, and no verifier is perfect. Treat a low score as 'safe to rely on with normal care', not 'certified fact'. See how accurate it is.
How is the score calculated?
The answer is split into atomic claims, each is grounded against independent sources and given a verdict, and the verdicts are aggregated — weighting contradicted claims most heavily. The exact weighting is tuned so the score tracks real risk, but it is always shown alongside the per-claim verdicts so you can audit it. See how it works.
Why not just give a pass or fail?
Because real answers are mixed: mostly right with one bad claim, or right-but-unverifiable. A single pass/fail hides that. A score plus per-claim verdicts lets you see exactly where the risk sits and decide what to double-check.
Verify what your AI just told you.
Paste any AI answer and Mephistopheles checks each claim against independent sources — no sign-up to try.