Why EHR data needs a trust score before any AI model trains on it
Over 90% of U.S. hospitals run on Epic or Cerner, yet no standard process exists to score the trustworthiness of EHR data before it enters an AI training pipeline. Without a trust score, models inherit every documentation gap, coding inconsistency, and consent violation baked into the source record. The Data Trust Index changes that by scoring every record 0-100 before any model touches it.
Every major health AI model trains on data that originated in an EHR. Epic and Cerner together hold records for more than 250 million patients in the United States. Yet not a single widely adopted standard exists to score the trustworthiness of those records before they enter a training pipeline.
That gap is not theoretical. It is the root cause of model failures, biased outputs, and regulatory exposure that health systems and AI developers are already facing.
What is actually wrong with EHR data
EHR data was designed for billing and clinical documentation, not for machine learning. The structure reflects that origin.
Duplicate records across merged health systems inflate certain patient populations. Missing fields in social determinants of health data create blind spots that models interpret as absence of need. Inconsistent coding practices across departments mean the same diagnosis can appear as three different SNOMED or ICD-10 entries within a single institution.
A 2023 JAMIA study found that 30% of clinical notes in large academic medical centers contained copy-pasted text from prior encounters, meaning the "new" data a model trains on may be months or years old. Epic and Cerner systems handle structured and unstructured data differently, so models trained on exports from one system often perform poorly when deployed against the other.
None of this shows up in a standard data quality check. You need a trust score.
Why trust is essential in the context of AI and data protection
Trust is not a soft concept when applied to AI training data. It is a measurable property that determines whether a model's outputs can be defended in a clinical, regulatory, or legal context.
If a model trains on records with broken consent chains, the resulting predictions carry legal liability under HIPAA and emerging state privacy laws. If a model trains on records with poor provenance, no one can trace a prediction back to its source data, which means no one can audit the model. The FDA's draft guidance on AI/ML-based Software as a Medical Device explicitly references the need for data quality documentation.
Trust in this context means: can you prove where the data came from, that the patient consented to this use, that the record is current, and that it agrees with other sources? Without those answers, the model is a black box built on sand.
How to prepare EHR data for AI model training
The standard approach treats data preparation as a technical pipeline: extract, transform, load, deduplicate, normalize. That is necessary but insufficient.
Before normalization, every record needs a trust assessment across multiple dimensions. The Data Trust Index scores records 0-100 across eight dimensions: Provenance (25%), Consent (20%), Recency (15%), Quality (10%), Concordance (10%), Validation (10%), Breadth (5%), and Stability (5%). Records below a defined trust floor never enter the training set.
This is how you prepare data for AI model training in healthcare. You do not just clean it. You score it, enforce a threshold, and document the scoring for every downstream consumer of the model.
The trust in automation scale, applied to health AI
The Trust in Automation scale, originally developed by Jian, Bisantz, and Drury, measures how much a human operator relies on an automated system's outputs. In healthcare AI, this scale maps directly to clinician adoption.
When a physician receives an AI-generated risk score derived from EHR data, their willingness to act on it depends on whether they trust the data underneath. If the hospital cannot explain the data's provenance, recency, or consent status, clinicians default to ignoring the output. That is rational behavior. The automation trust literature shows that a single unexplained failure drops operator trust by 30-50%, and recovery takes significantly longer than the initial trust-building period.
Scoring data before it reaches the model is how you build the foundation for clinician trust in the output.
Why you should not over-trust generative AI outputs built on EHR data
Generative AI tools trained on clinical data produce outputs that look authoritative. They generate discharge summaries, prior authorization letters, and clinical decision support recommendations in fluent clinical language.
That fluency masks the underlying data problem. A generative model trained on EHR records with a 30% copy-paste rate will reproduce stale clinical information with high confidence. A model trained on records with inconsistent coding will generate recommendations that reflect the loudest signal in the training data, not the most accurate one. Over-trusting these outputs without examining the input data trust score leads to clinical errors that are harder to catch precisely because the output reads as credible.
Key statistics
What this means for health systems deploying AI
If your organization is training or fine-tuning AI models on Epic or Cerner data, and you do not have a trust score on every input record, you are building on undocumented risk. Regulators, payers, and clinicians will all eventually ask the same question: how do you know the training data was trustworthy?
The answer cannot be "we ran a deduplication script." It needs to be a scored, auditable, dimension-level assessment of every record that entered the pipeline.
SuperTruth scores incoming EHR data at the point of ingestion, before it reaches a model. If your system is deploying clinical AI and needs to answer an auditor's questions, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.
Further reading:

Jason Alan Snyder
Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.
About SuperTruth · LinkedIn · Substack · jasonalansnyder.com
See it in practice
EHR data scored before any AI model sees it.
DTI integrates with Epic, Oracle Health, and all major EHR systems.