Fact check AI medical claims safely: de-identify, then verify
Last updated: July 26, 2026
To fact check AI medical claims safely, strip patient identifiers from the text before any external lookup, then ground each claim against PubMed-class sources for a per-claim verdict. Mephistopheles de-identifies first and flags results for clinician review; it does not diagnose or give medical advice.
Why AI medical content needs verifying, carefully
AI produces confident medical statements with confident-looking citations, and a large share do not hold up. In the 2023 Cureus study by Bhattacharyya and colleagues, of 115 references ChatGPT-3.5 generated for medical content, 47% were entirely fabricated and only 7% were both authentic and accurately presented (see the PubMed record). Rates differ by model and topic, and tend to be worse for less-common conditions, so treat any single figure as indicative rather than fixed.
In clinical work the stakes make silent errors unacceptable, which is exactly why a verifier that flags what it cannot confirm matters more here than almost anywhere else.
The de-identify-then-verify workflow
Grounding a claim means sending text to independent sources. For patient content that order has to be inverted so identifiers never leave first.
- De-identify. Names, dates, record numbers, and other identifiers are stripped from the text before any external grounding call. See de-identification.
- Extract claims. Each medical assertion is isolated as a discrete unit.
- Ground. The de-identified claim is checked against PubMed-class and other authoritative sources. See grounding.
- Verdict and flag. Each claim returns supported, contradicted, disputed, or unverifiable, flagged for a clinician to review.
Use /try for documents or /chat for pasted text. On Mephistopheles's own benchmark it catches about 88% of factual errors and clears about 98% of true statements, a roughly 2% false-alarm rate, but medical claims sit closer to the edges of the literature, so read every flag with clinical judgment.
Honest limits of de-identification
De-identification is risk reduction, not a guarantee
Stripping direct identifiers lowers exposure but does not eliminate re-identification risk. Quasi-identifiers such as a rare diagnosis, an unusual date combination, or a small-population location can still point back to an individual. The HIPAA Safe Harbor method removes 18 identifier categories and requires no actual knowledge that the remaining information could identify the person; that judgment is yours, not the tool's. Treat de-identify-then-verify as one control among your organization's safeguards, not a substitute for your own HIPAA analysis or a BAA where one is required.
Verify, do not diagnose
Mephistopheles checks whether a stated claim is supported by independent evidence. It does not diagnose, recommend treatment, or give medical advice, and its verdicts are inputs for a qualified clinician, not clinical decisions.
A supported verdict means the specific claim aligns with the sources checked at that time; it does not mean the statement is appropriate for a given patient. An unverifiable verdict is common for novel or narrow clinical questions and should trigger a look at primary sources and guidelines, not a workaround.
Frequently asked questions
Is it safe to fact check medical text that contains patient information?
Use the de-identify-then-verify path so identifiers are stripped before any external lookup. That materially reduces exposure, but it is risk reduction, not a guarantee: quasi-identifiers can remain, and you should still apply your own HIPAA analysis and any required agreements. Mephistopheles is a verification aid, not a compliance sign-off.
Does it diagnose or give medical advice?
No. It checks whether a stated claim is supported by independent sources and returns a verdict for clinician review. It does not diagnose, recommend treatment, or provide medical advice, and its output should be read as evidence about a claim, not a clinical decision.
Which sources does it check medical claims against?
It grounds claims against PubMed-class biomedical sources and other authoritative references, plus structured lookups. For narrow, novel, or rapidly evolving topics it will often return unverifiable rather than assert support, prompting you to consult primary literature and current guidelines.
What does the de-identification remove?
It targets direct identifiers such as names, dates tied to an individual, and record numbers, in line with the categories Safe Harbor addresses, before external grounding runs. It cannot guarantee removal of every re-identifying signal, so rare diagnoses or unusual fact patterns still warrant caution.
Verify what your AI just told you.
Paste any AI answer and Mephistopheles checks each claim against independent sources — no sign-up to try.