De-identification
Last updated: July 26, 2026
De-identification is removing the identifiers that could link data to a specific person, so it no longer counts as protected health information. Under HIPAA's Safe Harbor method, that means stripping 18 specified identifiers. De-identify-then-verify strips identifiers before any external grounding, so sensitive content can be checked without exposing who it's about.
What is de-identification?
De-identification is the process of removing information that could identify a specific individual, so the remaining data can be handled and shared with far less risk. In healthcare, properly de-identified data is no longer protected health information (PHI) under the HIPAA Privacy Rule.
HIPAA defines two paths. The Safe Harbor method requires removing 18 specified categories of identifiers and having no actual knowledge that the remainder could still identify someone. The Expert Determination method uses a qualified expert to certify the re-identification risk is very small. This page focuses on Safe Harbor because it is a concrete checklist.
De-identification reduces risk; it does not always reduce it to zero. Free text can carry identifying detail that a checklist misses, which is why residual-risk caution matters.
The HIPAA Safe Harbor 18 identifiers
Under 45 CFR 164.514(b)(2), the Safe Harbor method requires removing all 18 of these identifier categories, per HHS guidance (see the HHS de-identification guidance):
- Names
- Geographic subdivisions smaller than a state (with a limited exception for the first three ZIP digits when the population they cover is large enough)
- Dates (except year) related to an individual, and all ages over 89
- Telephone numbers
- Fax numbers
- Email addresses
- Social Security numbers
- Medical record numbers
- Health plan beneficiary numbers
- Account numbers
- Certificate or license numbers
- Vehicle identifiers and serial numbers, including license plates
- Device identifiers and serial numbers
- Web URLs
- IP addresses
- Biometric identifiers, including fingerprints and voiceprints
- Full-face photographs and comparable images
- Any other unique identifying number, characteristic, or code
Safe Harbor also requires no actual knowledge that the remaining data could re-identify the person.
De-identify-then-verify, and its limits
Verification needs to send claims to independent sources — web search, knowledge bases, judges. For sensitive content, doing that with identifiers intact would leak PHI. The de-identify-then-verify path strips identifiers before any external grounding: identifiers are removed, the de-identified claim is grounded, and the verdict is mapped back locally.
De-identification is risk reduction, not a guarantee
Automated de-identification can miss identifying detail embedded in free text — an unusual combination of facts can still point at a person even after the 18 identifiers are stripped. Treat de-id as a strong safeguard that lowers residual risk, not a promise of zero risk, and keep a human in the loop for genuinely sensitive material. This is not legal advice.
For how this is applied in practice, see verifying medical content and how it works.
Frequently asked questions
How many identifiers does HIPAA Safe Harbor require you to remove?
Eighteen. The HIPAA Safe Harbor method under 45 CFR 164.514(b)(2) lists 18 categories of identifiers — names, geographic detail below state level, most dates, contact details, SSNs, record and account numbers, device and biometric identifiers, and a catch-all for any other unique identifier — that must all be removed, per HHS guidance. You must also have no actual knowledge that the remaining data could re-identify the person.
Does de-identification make data completely safe?
No. It substantially reduces re-identification risk but doesn't eliminate it. Free text can carry identifying combinations of facts that survive stripping the 18 identifiers. De-identification is a strong safeguard, not a guarantee, so keep human oversight for sensitive material. This is not legal advice.
What is de-identify-then-verify?
It's the order of operations for checking sensitive content: strip identifiers first, then send the de-identified claim to external sources for grounding, then map the verdict back. This lets you verify claims in sensitive text without exposing who the text is about. See verifying medical content.
Verify what your AI just told you.
Paste any AI answer and Mephistopheles checks each claim against independent sources — no sign-up to try.