Mephistopheles
HomeGlossaryFalse-alarm rate
Glossary

False-alarm rate

Last updated: July 26, 2026

A false-alarm rate is how often a verifier flags a true statement as a problem. It's the false-positive rate. For a fact-checker it is the metric that matters most: if you can't trust it when it flags something, the tool is worthless. Mephistopheles passes about 98% of true statements clean — roughly a 2% false-alarm rate.

What is a false-alarm rate?

A false-alarm rate is how often a verifier raises a flag on a statement that is actually true. In detection terms it is the false-positive rate: a true claim wrongly marked contradicted or suspicious.

It is distinct from missing an error (a false negative). A checker has two ways to be wrong — it can cry wolf on a true claim, or it can wave through a false one. The false-alarm rate measures the first.

For a fact-checker, the false-alarm rate is arguably the single most important number. A tool that flags true statements as often as false ones trains you to ignore its flags — and a checker you learn to ignore protects nobody.

Why low false-alarm is THE metric

The whole value of a flag is that you act on it. The moment flags are frequently wrong, you stop trusting them, and the tool's real-world benefit collapses to zero — even if it technically catches many errors.

This is the classic trade-off between precision and recall. Precision asks: when it flags, is it right? Recall asks: of all real errors, how many did it catch? You can juice recall by flagging everything, but that destroys precision and floods you with false alarms.

Mephistopheles optimises for precision first. On its private benchmark of roughly 290 deliberately hard claims across about 20 domains, it passes about 98% of true statements clean — a false-alarm rate near 2% — while catching about 88% of factual errors, each with a contradicting source. Those are Mephistopheles' own measured numbers on that internal set, not an external benchmark; different data would give different results. See how accurate it is.

A worked example

Imagine two checkers on 100 true statements. Checker A flags 20 of them as wrong — a 20% false-alarm rate. Checker B flags 2. Even if both catch the same number of real errors, you will quickly ignore Checker A because most of its alarms are noise. Checker B's flags are rare, so when one appears you look — and that is the entire point.

This is why a fact-checker's honesty about its own false-alarm rate is part of the product. If it hid the number, you could not calibrate how much to trust a flag.

Frequently asked questions

What is a good false-alarm rate for a fact-checker?

Lower is better, and there is no universal threshold, but a false-alarm rate in the low single digits is the target for a checker whose flags you want to act on. Mephistopheles measures roughly 2% on its internal benchmark of hard claims. The right number depends on your data and how costly a false flag is in your workflow. See how accurate it is.

Is false-alarm rate the same as accuracy?

No. Accuracy blends several things. False-alarm rate isolates one specific failure: flagging a true statement. A tool can have high overall accuracy yet a false-alarm rate high enough that its flags are untrustworthy, so the two numbers must be reported separately.

How do you keep the false-alarm rate low?

By grounding each claim against independent sources before flagging, using a tiered judge, and returning unverifiable instead of contradicted when the evidence is merely absent. Declining to guess is the main lever that keeps false alarms down. See how it works.

Verify what your AI just told you.

Paste any AI answer and Mephistopheles checks each claim against independent sources — no sign-up to try.

Verify an answer →