Mephistopheles
HomeHow accurate
Measured accuracy

How accurate is Mephistopheles?

Last updated: July 26, 2026

On a private benchmark of about 290 deliberately hard factual claims across roughly 20 domains, Mephistopheles catches about 88% of factual errors (each with a contradicting source) and passes about 98% of true statements clean, a false-alarm rate near 2%. It flags what it cannot verify rather than guessing.

What are the numbers?

On a private benchmark of about 290 deliberately hard factual claims across roughly 20 domains, Mephistopheles catches about 88% of factual errors and passes about 98% of true statements clean. That second number is a false-alarm rate near 2%, and it is the one we care about most.

Here is why. A checker that screams at everything will catch every error and be useless, because you cannot tell a real flag from noise. The value of a fact-checker lives in its false-alarm rate: when Mephistopheles flags a claim, that flag has to be worth reading. A low false-alarm rate is what makes the flags trustworthy.

MetricWhat it measuresResult
Error recallShare of genuinely false claims caught (marked contradicted, with a source)~88%
True-statement pass rateShare of genuinely true claims passed clean~98%
False-alarm rateShare of true claims wrongly flagged~2%
Benchmark sizeDeliberately hard factual claims~290
Domain coverageIndependent subject areas~20

These are Mephistopheles' own measured numbers on its own benchmark. They are honest, not marketing-rounded, and we will not inflate them. If you want the mechanism behind them, see how it works.

How do we measure it?

We measure accuracy on a fixed benchmark of about 290 factual claims spread across roughly 20 domains, chosen to be hard on purpose: near-miss dates, plausible-but-wrong attributions, transposed figures, real papers cited for claims they do not actually support, and confident statements about niche topics.

The scoring rule is strict. A false claim only counts as caught if Mephistopheles marks it contradicted and attaches an independent source that actually contradicts it. A vague low-confidence shrug does not count as a catch. This is deliberately harsher than counting any non-supported verdict as a hit, because a flag without a source is not something you can act on.

Each benchmark claim carries a known ground-truth label. We run the full pipeline, compare each returned verdict against that label, and compute recall on the false claims and the pass rate on the true ones. The four-way verdict set is supported, contradicted, disputed, and unverifiable. See precision and recall for how these terms map to catching errors versus avoiding false alarms.

Why not 100%?

No honest verifier hits 100%, and a fact-checking product that claimed it should be the first thing you distrust. Two forces hold the number below perfect, and both are deliberate.

Unverifiable classes. Some claims have no authoritative independent source at the moment of checking: private data, very recent events past a model's knowledge cutoff, hyper-local facts, or genuinely contested questions. For these Mephistopheles returns unverifiable instead of guessing. That is the correct answer, but it means some true claims are not confirmed and some false ones are not caught, because the evidence simply is not there.

Deliberate under-flagging. We tune the judge to leave hedged, approximate, and rounded language alone. "Roughly 40 million" against an actual 39.7 million is not an error worth a red flag. Preserving the low false-alarm rate means occasionally letting a borderline claim pass rather than crying wolf. That is a conscious trade: we would rather miss a marginal case than teach you to ignore our flags.

Add the honest recall ceiling and the answer is simple: the ~88% and ~98% are the shape of an instrument built to be trusted when it speaks, not a machine pretending to omniscience.

What is it strong and weak at?

Mephistopheles is strongest on discrete, checkable facts that a source can confirm or contradict: dates, named entities, numbers, quotations, statutory citations, attributions, and "does this cited paper actually say this" checks. This is the class where a real citation is not the same as a supported claim, and independent grounding pays off most.

It is weaker, and openly says so, on a few classes:

  • Fresh events past retrieval coverage, where no indexed source exists yet.
  • Niche or private facts with no authoritative public source, which land as unverifiable.
  • Shared-belief errors: if the entire independent web and reference corpus repeat the same wrong thing, an independent checker inherits that ceiling. Grounding de-correlates random error, not shared error.
  • Matters of interpretation, prediction, or opinion, which are not factual claims to verify.

Not advice, and not a final arbiter

Mephistopheles is a detector, not an oracle. It surfaces claims and evidence for a human to review. It is not legal, medical, or financial advice, and a clean pass is not a guarantee of truth, only that no contradicting source was found. Treat every flag, and every unverifiable, as a prompt to look closer.

Will these numbers change?

Yes, and you should expect them to. These figures come from internal testing on a fixed benchmark, not from an external audit, and they move as the engine improves: better retrieval, stronger judges, wider source coverage, and a larger, harder benchmark all shift the numbers.

We report the current measured values honestly and update this page when they change materially. What will not change is the priority: the false-alarm rate stays low, because a checker you cannot trust when it flags is worthless. If a change traded a lower false-alarm rate for higher recall in a way that hurt trust, we would not ship it.

You can pressure-test the numbers yourself. Paste a known-wrong AI answer into the verifier, drop a document into the document checker, or run claims through the API and read the per-claim verdicts and risk score against facts you already know.

Frequently asked questions

How accurate is Mephistopheles at catching AI hallucinations?

On a private benchmark of about 290 deliberately hard factual claims across roughly 20 domains, Mephistopheles catches about 88% of factual errors, each backed by a contradicting source, and passes about 98% of true statements clean. It reports unverifiable when no authoritative source exists rather than guessing.

What is the false-alarm rate?

About 2%. Roughly 98% of genuinely true statements pass clean, so only about 2% are wrongly flagged. This is the headline metric: a low false-alarm rate is what makes a flag worth reading, since a checker that flags everything is useless.

Why doesn't Mephistopheles catch 100% of errors?

Two reasons, both deliberate. Some claims have no authoritative independent source, so they return unverifiable instead of a guess. And the judge deliberately leaves hedged or approximate language alone to protect the low false-alarm rate. A fact-checker claiming 100% accuracy should be distrusted.

Are these numbers independently audited?

No. They come from Mephistopheles' own internal testing on a fixed benchmark and are reported without inflation. They move as the engine improves. You can verify them yourself by running known-true and known-false claims through the verifier or the API and checking the verdicts.

What counts as a caught error in the benchmark?

A false claim only counts as caught if Mephistopheles marks it contradicted and attaches an independent source that actually contradicts it. A low-confidence shrug does not count. This strict scoring is harsher than counting any non-supported verdict as a hit, because a flag without a source cannot be acted on.

Verify what your AI just told you.

Paste any AI answer and Mephistopheles checks each claim against independent sources — no sign-up to try.

Verify an answer →