The difference between data anonymization and data de-identification in healthcare
Photo by Gustavo Leighton on Unsplash
insight

The difference between data anonymization and data de-identification in healthcare

By Jason Alan Snyder·May 11, 2026

HIPAA defines two methods for de-identification, but neither method equals anonymization. The distinction determines whether health data can be re-linked to a patient, whether it qualifies for regulatory safe harbor, and whether AI models trained on it carry legal exposure. Most healthcare organizations conflate the two terms and absorb risk they never priced.

Most healthcare organizations use "anonymized" and "de-identified" as if they mean the same thing. They do not. The legal distinction between these two terms determines whether your data qualifies for HIPAA safe harbor, whether it can be shared across state lines, and whether an AI model trained on it creates liability you never accounted for.

What is de-identification in healthcare?

De-identification is the process of removing or obscuring specific data elements so that the remaining information cannot reasonably identify an individual. Under HIPAA, de-identification has a precise legal definition codified in 45 CFR §164.514. It is not a general concept. It is a compliance standard with two approved methods.

The Expert Determination method (§164.514(b)(1)) requires a qualified statistical or scientific expert to determine that the risk of identifying any individual is "very small." The expert must document the methods and results of the analysis.

The Safe Harbor method (§164.514(b)(2)) requires the removal of 18 specific identifiers, including names, geographic data smaller than a state, dates more specific than a year, phone numbers, email addresses, Social Security numbers, medical record numbers, and biometric identifiers. If all 18 are removed and the covered entity has no actual knowledge that the remaining data could identify someone, the data qualifies as de-identified.

De-identified data under HIPAA is no longer considered protected health information (PHI). This means HIPAA's privacy and security rules no longer apply to it. But here is the critical point: de-identified data can sometimes be re-identified.

What is anonymization, and how does it differ?

Anonymization goes further. Anonymized data has been processed so that re-identification is not just unlikely but impossible. There is no key, no linkage table, no auxiliary dataset that could reconnect the data to the original individual. The process is irreversible by design.

HIPAA does not use the term "anonymization." It is a concept more commonly associated with GDPR and international privacy frameworks. Under GDPR, truly anonymized data falls outside the regulation entirely because it is no longer personal data. But GDPR sets a high bar: if any party, using any reasonably available means, could re-identify the data, it is not anonymous.

The practical difference: de-identification reduces the probability of re-identification. Anonymization eliminates it. De-identification is a spectrum. Anonymization is a binary state.

Another word for de-identified data is "pseudonymized" in some contexts, though pseudonymization technically means replacing identifiers with artificial ones while retaining a re-linkage key. Pseudonymized data is explicitly not anonymous under GDPR because the key exists.

Why the confusion matters for healthcare AI

The conflation of these terms creates three specific risks.

First, organizations share de-identified data assuming it carries no re-identification risk. Research has shown otherwise. A 2019 study in Nature Communications demonstrated that 99.98% of Americans could be re-identified in any dataset using just 15 demographic attributes, even when the dataset was "de-identified" under Safe Harbor rules.

Second, AI models trained on de-identified data can memorize patterns that effectively re-identify individuals. If a training dataset contains a rare disease diagnosis, a specific age range, and a ZIP code, the combination may be unique to a single patient even after Safe Harbor removal of the 18 identifiers.

Third, state privacy laws are diverging from HIPAA. Washington's My Health My Data Act, effective 2024, applies to consumer health data outside HIPAA's scope. California's CCPA treats de-identified data differently than HIPAA does. A dataset that qualifies as de-identified under federal law may not qualify under state law.

As MedPage Today has covered in the context of physicians writing about patients, even narrative clinical descriptions without names or dates can identify individuals in small populations. The ethics of data use do not end at HIPAA compliance.

De-identification vs pseudonymization vs anonymization

Privacy protection spectrum: what each method preserves and removes
Privacy protection spectrum: what each method preserves and removes

These three terms represent a spectrum of privacy protection:

Pseudonymization replaces direct identifiers with tokens or codes. A linkage key exists. The data can be re-identified by the key holder. GDPR considers pseudonymized data personal data. HIPAA does not specifically define pseudonymization.

De-identification removes or generalizes identifiers to reduce re-identification risk below a defined threshold. Under HIPAA, this is a legal standard with two approved methods. Re-identification remains theoretically possible.

Anonymization destroys all linkage paths permanently. No party can re-identify the data using any reasonably available means. The data loses its status as personal data or PHI entirely.

For healthcare AI, the method you choose determines what you can do with the data, who you can share it with, and what regulatory obligations persist.

Key statistics

DTI trust dimensions: how provenance and consent weight privacy method documentation
DTI trust dimensions: how provenance and consent weight privacy method documentation

  • HIPAA's Safe Harbor method requires removal of 18 specific identifiers before data qualifies as de-identified
  • 99.98% of Americans can be re-identified using just 15 demographic attributes in de-identified datasets (Nature Communications, 2019)
  • HIPAA de-identified data is exempt from the Privacy Rule, but at least 14 states have enacted or proposed laws with stricter standards than federal de-identification
  • SuperTruth's DTI Engine scores every health data record across 8 trust dimensions, including provenance and consent, which capture whether de-identification or anonymization methods were properly applied and documented
  • In the imaware case study, SuperTruth processed 105,000 diagnostic records and reduced standardization time from 3 weeks to 2 hours, a 95% reduction, by scoring data integrity before downstream use
  • Can anonymized data be identified?

    By definition, no. If it can be re-identified, it was never truly anonymized. The problem is that many organizations label data as "anonymized" when it was only de-identified. This mislabeling creates a false sense of security and, in GDPR jurisdictions, potential fines of up to 4% of global annual revenue.

    The real question is whether your organization can prove which method was applied, when, by whom, and whether the method held up against current re-identification techniques. That is a provenance problem. And provenance is the single most heavily weighted dimension in the Data Trust Index at 25%.

    Without documented provenance for your de-identification or anonymization process, you cannot demonstrate compliance to an auditor, a regulator, or a court. The data may be clean. But you cannot prove it.

    What this means for your data pipeline

    If you are building AI models on health data, the distinction between anonymization and de-identification is not academic. It determines your HIPAA exposure, your state-law compliance posture, your GDPR obligations if you operate internationally, and whether your training data can withstand an FDA audit of your model's data provenance.

    The DTI Engine scores every health data record 0 to 100 across 8 trust dimensions before your AI model sees it. Provenance and consent scoring capture whether de-identification or anonymization was applied correctly, documented fully, and remains valid given the data's current context. If your team is evaluating data for training, compliance, or clinical use, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.

    Further reading:

  • DTI™ Engine
  • Health systems solution
  • What HIPAA does not tell you about data trust
  • Why consent governance fails in healthcare data and what fixes it
  • The chain of custody problem in health data: why provenance is the hardest dimension
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0–100. Travels with every record permanently.

    See the DTI Engine
    Share
    The difference between data anonymization and data de-identification in healthcare | SuperTruth