FDA AI guidance and data provenance: what is coming for healthcare AI developers
Photo by Compagnons on Unsplash
insight

FDA AI guidance and data provenance: what is coming for healthcare AI developers

By Jason Alan Snyder·April 26, 2026

The FDA's evolving guidance on AI in healthcare is converging on a single requirement most developers are not ready for: data provenance. Draft frameworks for AI-enabled medical devices and drug development now treat training data lineage as a core submission element, not an afterthought. Here is what is coming and what developers need to build now.

The FDA cleared or approved over 1,000 AI-enabled medical devices by late 2024. More than 170 arrived in a single year. But the agency's next regulatory phase is not about approving more devices faster. It is about asking a question most developers cannot answer: where did your training data come from, and can you prove it?

The regulatory trajectory is clear

Three concurrent FDA initiatives are converging on data provenance as a hard requirement for healthcare AI.

First, the FDA's draft guidance on AI-enabled medical devices, released under the Software as a Medical Device (SaMD) framework, explicitly calls for documentation of data management practices, including collection methodology, labeling procedures, and dataset representativeness. Second, the joint FDA/Health Canada/MHRA guiding principles for Good Machine Learning Practice list data quality and relevance as foundational requirements. Third, the FDA's 2024 guidance on artificial intelligence in drug manufacturing signals that provenance expectations extend beyond devices into pharmaceutical production pipelines.

These are not isolated documents. They form a regulatory pattern. The FDA is building toward a world where "we used de-identified EHR data" is not a sufficient answer.

What FDA data provenance requirements actually demand

The word "provenance" appears repeatedly across FDA AI healthcare guidance documents, but the practical requirements break down into specific obligations.

Developers must document the original source of every training dataset. They must record every transformation applied between collection and model ingestion. They must demonstrate that training populations are representative of intended-use populations. And they must maintain this documentation in a form that survives audit, not just internal review.

The 10 guiding principles published jointly by the FDA, Health Canada, and MHRA in 2021 set the expectation explicitly: "Data collection protocols should ensure that the relevant characteristics of the intended patient population are sufficiently represented." The 2024 and 2025 draft guidances are operationalizing this principle into concrete submission requirements.

This is where most healthcare AI developers hit a wall. Provenance is not a field in a database. It is a chain of custody that spans institutions, EHR systems, data brokers, and transformation pipelines. As we have written before, the chain of custody problem in health data is the hardest dimension to solve because it requires infrastructure, not just documentation.

What is coming for healthcare AI developers by 2026

The FDA's trajectory suggests three concrete changes developers should prepare for.

Pre-market submissions for AI-enabled devices will require structured provenance documentation. This means machine-readable records of data origin, transformation history, and consent status for every dataset used in training and validation. The era of narrative-only data descriptions in 510(k) submissions is ending.

Post-market surveillance will include data drift monitoring. The FDA's predetermined change control plan (PCCP) framework already anticipates that AI models change over time. Provenance tracking enables regulators to ask not just "did your model drift" but "did your training data's characteristics drift from your intended-use population." We have covered how temporal drift destroys AI model accuracy in detail.

Third, the FDA is signaling alignment with international frameworks. The EU AI Act's high-risk classification for medical AI includes explicit data governance requirements. Developers building for global markets face converging provenance mandates from multiple regulators simultaneously.

Key statistics

Data Trust Index: dimension weights reflecting regulatory priority
Data Trust Index: dimension weights reflecting regulatory priority

The numbers frame the scale of what is ahead.

The FDA has cleared or approved over 1,000 AI/ML-enabled medical devices cumulatively through 2024, with radiology accounting for roughly 75% of all clearances. The agency's list of FDA-approved AI medical devices grows by approximately 150 to 200 entries per year.

The SuperTruth Data Trust Index weights provenance at 25% of the total trust score, the single highest-weighted dimension, reflecting its regulatory importance. In our work with imaware, we standardized 105,000 diagnostic records and reduced data preparation time from 3 weeks to 2 hours, a 95% reduction. That process surfaced provenance gaps that would have been invisible without systematic scoring.

A 2023 study in Nature Medicine found that fewer than 30% of FDA-cleared AI devices had publicly available training data documentation sufficient to evaluate bias or representativeness. The gap between what the FDA will require and what developers currently provide is enormous.

Why HIPAA compliance is not enough

Many healthcare AI developers assume that HIPAA-compliant data handling satisfies regulatory requirements for data governance. It does not.

HIPAA governs privacy. FDA AI healthcare guidance governs fitness for purpose. A dataset can be perfectly de-identified under HIPAA Safe Harbor and still fail FDA provenance requirements because no one recorded which hospital system it originated from, what transformations were applied during de-identification, or whether the patient population was representative of the intended use. We wrote a full analysis of what HIPAA does not tell you about data trust and separately addressed why patient consent does not cover model training.

The FDA's emerging framework treats data provenance as a quality system requirement, not a privacy requirement. This distinction matters for every team building submission-ready AI.

How to build provenance infrastructure now

Data preparation time: before and after DTI scoring (imaware case study)
Data preparation time: before and after DTI scoring (imaware case study)

Developers who wait for final guidance to build provenance infrastructure will be two years behind.

Start by scoring existing training data across provenance, consent, recency, and quality dimensions. Identify gaps before a reviewer does. Build automated chain-of-custody logging into your data ingestion pipeline so every record carries its history forward. And establish DTI floor thresholds below which data does not enter training, not as a best practice, but as an auditable policy.

The FDA is not asking developers to solve provenance retroactively. It is telling them to build it into the development lifecycle from day one. The developers who treat data trust scoring as pre-training infrastructure will have a regulatory advantage that compounds with every submission.

SuperTruth's DTI Engine scores every health data record 0 to 100 across 8 trust dimensions before your AI model sees it. Provenance carries the highest weight at 25% because regulators weight it highest too. If your team is building AI for FDA submission and needs audit-ready provenance documentation, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.

Further reading:

  • DTI™ Engine
  • Pharma solution
  • How the FDA will audit your health AI's training data
  • Data provenance in healthcare AI: why chain of custody matters before training
  • What makes health data Platinum-grade for FDA regulatory submission
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0–100. Travels with every record permanently.

    See the DTI Engine
    Share