The HIPAA Problem Health AI Companies Are Ignoring: Patient Consent Does Not Cover Model Training
When a patient signs a HIPAA notice at a hospital, they authorize the use of their data for treatment, payment, and healthcare operations. They do not authorize its use in training an AI model. Most health AI companies are building on a consent foundation that does not exist.
When a patient signs a HIPAA notice at a hospital, they are authorizing a specific set of uses: treatment by their care team, payment processing by their insurer, and certain healthcare operations. That authorization does not extend to AI model training.
Most health AI companies are training on patient data they were never authorized to use.
This is not a theoretical concern. It is a specific legal gap that regulators are beginning to close — and health AI developers who have not addressed it at the record level are building on a foundation that will not hold.
What HIPAA actually authorizes
HIPAA's Privacy Rule defines three categories of permissible use without patient authorization: treatment, payment, and healthcare operations. The healthcare operations category is broad enough that many health AI companies have argued their training activities fall within it.
That argument is becoming harder to sustain.
FDA's guidance on AI/ML-Based Software as a Medical Device distinguishes between using patient data to improve care for that specific patient and using patient data to train a general-purpose model intended for other patients. The second use is not obviously covered by the healthcare operations carveout.
FTC has taken action against health technology companies that used consumer health data for AI training without adequate notice. The argument that health AI training is a healthcare operation is less defensible every year.
The consent gap at the record level
Even setting aside regulatory risk, there is a practical problem: most health AI companies cannot demonstrate that any specific patient record in their training dataset was authorized for AI training use.
They can point to a general HIPAA notice. They cannot point to a record-level authorization that covers this specific model, trained for this specific purpose, using this specific patient's data.
FDA's SaMD guidance is moving toward requiring exactly that documentation. When the reviewer asks — and they will ask — the answer needs to be more than a HIPAA notice signed five years ago.
Five consent tiers, not one
The problem with current health AI consent practice is that it treats consent as binary: HIPAA-covered or not HIPAA-covered. In reality, consent has at least five meaningful levels for health AI:
Most health AI companies are training on Level 1 data while asserting Level 4 authorization. The documentation does not exist to support that assertion.
Tracking consent at the record level
The solution is not to re-consent every patient — that is not operationally feasible. The solution is to track consent state at the record level for every intended use, and to route training runs only to records whose documented consent covers AI training.
SuperTruth's ConsentOS implements exactly this. Five consent tiers tracked per record. Records authorized for AI training are distinguished from records authorized only for direct care. Revocation propagates in real time — a patient who withdraws consent at 9am is removed from training datasets by 9am, not at the next batch processing cycle.
Every training run produces a consent audit trail: which records were included, what their consent tier was at the time of inclusion, and whether any records have had their consent status changed since inclusion.
When the FDA reviewer asks, there is an answer.
Further reading: Health AI data trust infrastructure and Why AI models trained on unscored health data will fail in production

Jason Alan Snyder
Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.
About SuperTruth · LinkedIn · Substack · jasonalansnyder.com
See it in practice
DTI scores the record, not the patient.
8 dimensions. 0–100. Travels with every record permanently.