The chain of custody problem in health data: why provenance is the hardest dimension
Provenance carries the highest weight in the Data Trust Index for a reason: it is the dimension most likely to be broken, forged, or simply missing across health data systems. A single patient record can pass through 17 or more systems before reaching a decision point, and each handoff creates an opportunity for the chain of custody to fracture silently.
A single patient record touches an average of 17 different systems over the course of a chronic disease. Each handoff is an opportunity for the chain of custody to break. Most of the time, nobody notices until something goes wrong.
Provenance is the dimension that makes or breaks every other measure of health data trust. It is also the one most likely to be absent.
What is provenance in healthcare?
Provenance in healthcare means knowing where a data record originated, every system it passed through, every transformation applied to it, and every actor who touched it along the way. It answers a deceptively simple set of questions: Who created this record? When? In what system? Has it been altered, merged, or duplicated since?
Most EHR systems store a creation timestamp and a last-modified date. That is not provenance. That is two data points out of dozens required to reconstruct a record's full history. True provenance tracking requires a continuous, immutable log of every state change from origin to the moment a clinician, algorithm, or regulator reads the record.
The concept parallels provenance in forensics, where chain of custody refers to the documented, unbroken trail of evidence from collection to courtroom. If a gap exists, the evidence is inadmissible. Health data should be held to the same standard, but rarely is.
Why provenance is the hardest dimension to get right
SuperTruth's Data Trust Index scores every health data record 0 to 100 across eight dimensions. Provenance carries the highest weight at 25%. That weighting reflects two realities.
First, provenance failures cascade. If you cannot verify where a record came from, you cannot validate its quality, confirm its consent status, or trust its recency. Every downstream dimension depends on the chain of custody being intact.
Second, the infrastructure for tracking provenance barely exists. HL7 FHIR includes a Provenance resource, but fewer than 12% of health systems populate it consistently. Most interoperability layers pass data without metadata about origin. The record arrives, but its history does not.
The three domains of healthcare quality proposed by Donabedian (structure, process, and outcomes) assume that the data feeding each domain is trustworthy. Provenance is the precondition for that assumption. Without it, you are measuring outcomes with instruments you cannot calibrate.
Where the chain breaks
Health data chain of custody fails at predictable points.
EHR migration. When a hospital switches from one EHR vendor to another, records are bulk-exported and re-imported. Migration scripts flatten provenance metadata. A record that originated in a lab system in 2014 may arrive in the new EHR with a creation date of 2023.
Health information exchanges. HIEs aggregate records from multiple sources but often strip source-system identifiers during normalization. The record looks clean. Its origin is gone.
Patient matching errors. When two records are incorrectly merged, the provenance of both is corrupted. The National Institute of Standards and Technology estimates duplicate rates of 8 to 12% across large health systems.
AI training pipelines. Data scientists extract records, transform them into training sets, and rarely carry provenance metadata forward. The FDA is now asking how AI training data was sourced, and most organizations cannot answer.
What has contributed to slow adoption of electronic health records?
Several factors have slowed EHR adoption, but one of the least discussed is the provenance burden. Clinicians and administrators recognize that digitizing records does not automatically create trustworthy records. The cost of maintaining audit trails, the complexity of interoperability standards, and the absence of universal patient identifiers all create friction. Meaningful Use incentives accelerated adoption of EHR systems themselves, but did not require the kind of provenance infrastructure that would make records trustworthy across organizational boundaries. The result is widespread digitization with shallow data lineage. Records exist in electronic form, but their chain of custody is fragmented by design.
Key statistics
Provenance carries 25% of the total weight in the SuperTruth Data Trust Index, more than any other single dimension.
Fewer than 12% of health systems consistently populate the HL7 FHIR Provenance resource in their interoperability transactions.
NIST estimates patient record duplicate rates of 8 to 12% across large health systems, each duplicate corrupting the provenance of both source records.
SuperTruth standardized 105,000 diagnostic records for imaware, reducing processing time from 3 weeks to 2 hours, a 95% reduction driven primarily by automated provenance validation.
The imaware engagement saved over 200 hours per month and identified a patient segment driving 20% of revenue that had been invisible in unstandardized data.
How SuperTruth solves the provenance problem
The DTI Engine assigns a provenance sub-score to every record based on five factors: source system identification, transformation history, actor attribution, timestamp integrity, and lineage completeness. Each factor is scored independently, then weighted into the provenance dimension.
When we onboarded imaware's 105,000 diagnostic records, provenance validation was the first pass. Records without verifiable source-system metadata were flagged, not discarded. We traced each record back to its originating lab instrument, reconciled timestamps across systems, and rebuilt the chain of custody that had been lost during prior data handling. CEO Brodie Flanders described the result: "The lab industry has never had a trust standard. DTI created one."
This approach differs from what currently ranks in search results for health data provenance. Academic surveys catalog the problem. We score it. Every record gets a number. That number tells you whether the chain of custody is intact before you use the record for clinical decisions, AI training, or regulatory submission.
For health systems preparing to deploy AI, provenance scoring is not optional. It is the prerequisite. If your training data cannot prove where it came from, your model cannot prove what it learned. The hospital AI readiness checklist starts here.
For the VA system, where records span decades and dozens of legacy systems, the provenance challenge operates at a scale that demands automated scoring, not manual review.
What provenance makes possible
When every record carries a verifiable chain of custody, three things become possible that are currently blocked.
First, consent becomes auditable. You can verify not just that consent was given, but that it was given for the specific version of the record that exists today. Consent governance depends on provenance.
Second, AI models become explainable. Regulators can trace a model's output back through its training data to the original source records, verifying that nothing was corrupted in the pipeline.
Third, patients can trust the system. Tools like MyBio.Health give patients visibility into where their data lives and how it has been used. That visibility requires provenance infrastructure underneath.
The chain of custody problem is not a theoretical concern. It is the structural weakness underneath every health data system that has not been scored. Fix provenance first. Everything else follows.
To understand how the DTI Engine scores provenance across your data, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.
Further reading:

Jason Alan Snyder
Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.
About SuperTruth · LinkedIn · Substack · jasonalansnyder.com
See it in practice
DTI scores the record, not the patient.
8 dimensions. 0–100. Travels with every record permanently.