Data quality vs data trust: what is the difference and why it matters for healthcare AI
Data quality measures whether a record is accurate and complete. Data trust measures whether that record should be used at all. Healthcare AI needs both, but the industry has invested almost exclusively in quality while ignoring trust, and that gap is where clinical AI fails.
Most healthcare organizations already have data quality programs. They check for missing fields, validate formats, flag duplicates, and enforce schema compliance. These programs are necessary. They are also insufficient.
Data quality tells you whether a record is accurate. Data trust tells you whether that record should be used. The difference between data quality vs data trust is the difference between a lab result that is correctly formatted and a lab result you can trace back to a specific patient, a consented collection, a verified source, and a known chain of custody. Healthcare AI needs both. The industry has built infrastructure for one.
What data quality actually measures
Data quality is a property of the data itself. A quality check asks: Is this field populated? Is the value within an expected range? Does the format match the schema? Is this record a duplicate?
These checks matter. An EHR record with a missing diagnosis code or a transposed lab value can cascade errors downstream. When a clinical decision support tool ingests garbage, it outputs garbage. That reality drives the entire data quality industry.
But quality checks operate on a narrow surface. They evaluate the state of the data at a single point in time. They do not ask where the data came from, whether the patient consented to its use in model training, how many times the record was transformed between source and destination, or whether the record is still clinically current.
What data trust actually measures
Data trust is a property of the relationship between the data, its source, and its intended use. A trust assessment asks: Can I prove where this record originated? Was consent collected for this specific use case? How recently was this data validated against a primary source? Does this record agree with other records about the same entity?
The Data Trust Index (DTI) scores every health data record from 0 to 100 across eight dimensions: Provenance (25%), Consent (20%), Recency (15%), Quality (10%), Concordance (10%), Validation (10%), Breadth (5%), and Stability (5%). Quality is one dimension. It accounts for 10% of the total trust score.
That weighting is deliberate. A record can score perfectly on quality and still score below 50 on trust if it lacks provenance, consent, or recency. A perfectly formatted patient record collected without proper consent authorization is a compliance liability, not a training asset.
Why does data quality matter for AI?
AI models learn patterns from data. If the data contains systematic errors, missing values, or inconsistencies, the model learns those patterns too. A sepsis prediction model trained on records where vitals are frequently missing will learn to underweight vitals. A radiology AI trained on images with inconsistent labeling will produce inconsistent diagnoses.
A 2023 MedPageToday opinion piece noted that AI is "becoming embedded in the day-to-day delivery of healthcare," from documentation support to clinical decision tools. That embedding makes quality failures more consequential, not less. When a spreadsheet has a bad row, someone notices. When a model has a bad training set, nobody notices until a patient is harmed.
But quality alone cannot prevent the failures that regulators and clinicians are most concerned about. The FDA's emerging guidance on AI/ML-based software as a medical device focuses heavily on data provenance and lifecycle documentation. These are trust dimensions, not quality dimensions.
The C's frameworks and their limits
The data quality field has produced several frameworks organized around "C" words. The most common versions:
The 3 C's of data quality focus on Completeness, Consistency, and Correctness. These are the baseline checks most data pipelines already enforce.
The 4 C's of data quality add Currency (timeliness) to the original three, acknowledging that a record can be complete and correct but outdated.
The 5 C's of data quality extend further to include Conformity (schema and format compliance), creating a more comprehensive quality framework.
Each version improves on the last. None of them address provenance, consent, or concordance. None ask whether the data was collected under conditions that permit its intended use. None evaluate whether independent sources agree about the same entity. These frameworks were designed for data warehousing, not for training clinical AI systems where a wrong answer can mean a wrong treatment.
Key statistics
Where the gap causes real harm
Consider substance use disorder records. Federal 42 CFR Part 2 regulations impose strict consent requirements on this data. A record can pass every quality check and still be legally prohibited from use in a training set. A quality score will not flag this. A trust score will.
Consider pediatric data. Consent for minors involves different legal thresholds, and those thresholds change when the minor becomes an adult. Quality frameworks do not track consent lifecycle. Trust frameworks must.
Consider health equity. Data from underserved populations is often incomplete or collected under conditions with weaker consent infrastructure. Quality checks may flag these records as low-quality and exclude them, worsening the representation gap. Trust scoring can identify the specific dimension that is deficient and route the record for targeted remediation rather than blanket exclusion.
What this means for healthcare AI teams
If your organization is building or deploying clinical AI, a data quality program is table stakes. You need one. But a quality program alone leaves you exposed on the dimensions that regulators, payers, and patients increasingly care about: Where did this data come from? Did the patient consent to this use? Can you prove the chain of custody?
These are not theoretical concerns. The FDA is actively developing audit frameworks for AI training data. CMS compliance programs like CRUSH require provenance documentation for provider directories. NCQA credentialing standards demand primary source verification. Each of these maps to a trust dimension, not a quality dimension.
The DTI Engine scores every health data record 0 to 100 across 8 trust dimensions before your AI model sees it. Quality is one of those dimensions. The other seven are the ones most organizations have never measured. If your team is evaluating data for training, compliance, or clinical use, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.
Further reading:

Jason Alan Snyder
Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.
About SuperTruth · LinkedIn · Substack · jasonalansnyder.com
See it in practice
DTI scores the record, not the patient.
8 dimensions. 0–100. Travels with every record permanently.