Why Innovaccer, Datavant, and AWS Health Lake are not solving the data trust problem
Photo by David Pupăză on Unsplash
insight

Why Innovaccer, Datavant, and AWS Health Lake are not solving the data trust problem

By Jason Alan Snyder·April 30, 2026

Innovaccer, Datavant, and AWS HealthLake each solve a real infrastructure problem. None of them solve the data trust problem. They move, link, and store health data without ever answering the question that matters most: should this record be trusted?

Innovaccer unifies health data. Datavant links it across institutions. AWS HealthLake stores it in a FHIR-native format. All three platforms perform useful infrastructure functions. None of them answer the question that regulators, clinicians, and AI developers actually need answered: is this specific record trustworthy enough to act on?

That is not a minor gap. It is the gap.

What each platform actually does

Innovaccer aggregates clinical, claims, and operational data into a unified patient view. It normalizes formats, deduplicates records, and feeds analytics dashboards. Its value proposition is consolidation.

Datavant tokenizes patient identifiers so records can be linked across organizations without exposing PHI. It connects datasets that were previously siloed. Its value proposition is linkage.

AWS HealthLake ingests FHIR R4 data, applies NLP to unstructured clinical notes, and stores everything in a queryable format. Its value proposition is cloud-native storage and retrieval.

Each platform assumes the data it receives is fit for use. None of them score it.

The trust problem none of them address

Consolidating bad data produces a unified view of bad data. Linking records with unknown provenance creates a larger dataset of unknown provenance. Storing unverified clinical notes in FHIR format does not make them verified.

The biggest challenge of health information technology is not moving data between systems. It is knowing whether the data that arrives is accurate, current, properly consented, and traceable to its source. Infrastructure trust is not data trust, and conflating the two has cost the industry years of misallocated investment.

A 2023 study in JAMIA found that 25% of clinical records used for AI model training contained at least one clinically significant error. Innovaccer does not flag those errors. Datavant does not detect them during linkage. HealthLake does not score them at ingestion.

What are the challenges in delivering quality health care?

The challenges compound across four layers: fragmented records, stale data, inconsistent consent governance, and zero provenance tracking. A patient who visits three health systems in two states generates records in three EHR instances, often with conflicting medication lists, outdated diagnoses, and no shared consent framework.

Datavant can link those three records by token. It cannot tell you which medication list is current. Innovaccer can display all three in a single dashboard. It cannot tell you which diagnosis code was entered by a physician and which was auto-populated by a billing system. AWS HealthLake can store all of it in FHIR. It cannot tell you whether the patient consented to AI model training on any of it.

These are not edge cases. They are the default state of health data in the United States.

What are the major forces affecting the delivery of healthcare?

Three forces are converging: the FDA's increasing scrutiny of AI training data provenance, CMS requirements for accurate provider and quality data, and the commercial pressure to deploy AI models faster than data pipelines can support. Each force demands not just data availability but data trustworthiness.

The FDA's draft guidance on AI/ML-based software expects developers to document the provenance, quality, and representativeness of training data. That audit trail does not exist in Innovaccer, Datavant, or HealthLake. None of them produce a per-record trust score. None of them track chain of custody across transformations.

Which of the following is a major challenge in implementing systems thinking in healthcare?

The major challenge is that each system optimizes for its own layer without accounting for the layers above and below it. Datavant optimizes for linkage. Innovaccer optimizes for aggregation. HealthLake optimizes for storage. No one optimizes for the integrity of the record itself across its full lifecycle.

Systems thinking requires a common unit of measurement. In finance, that unit is a credit score. In health data, that unit should be a trust score: a per-record, per-dimension assessment that travels with the data regardless of which platform stores, links, or displays it.

Key statistics

imaware data processing: before and after DTI Engine
imaware data processing: before and after DTI Engine

SuperTruth's Data Trust Index scores every health record 0 to 100 across 8 dimensions. In a case study with imaware, SuperTruth standardized 105,000 diagnostic records, reduced processing time from 3 weeks to 2 hours (a 95% reduction), and saved over 200 hours per month. That same scoring process identified a patient segment driving 20% of imaware's revenue that had been invisible in their existing analytics.

Provenance accounts for 25% of the DTI score. Consent accounts for 20%. Together, those two dimensions represent 45% of what makes a health record trustworthy, and neither Innovaccer, Datavant, nor AWS HealthLake measures either one.

What the trust layer looks like

DTI scoring dimensions: what Innovaccer, Datavant, and HealthLake do not measure
DTI scoring dimensions: what Innovaccer, Datavant, and HealthLake do not measure

The DTI Engine scores records at the point of ingestion, before any model trains on them. Each record receives a score from 0 to 100 across Provenance (25%), Consent (20%), Recency (15%), Quality (10%), Concordance (10%), Validation (10%), Breadth (5%), and Stability (5%). Think of it as a FICO score for health data.

This score is not a replacement for what Innovaccer, Datavant, or HealthLake do. It is the layer that should exist underneath all of them. You can aggregate, link, and store health data. But if you cannot answer "should this record be trusted for this use case," you have infrastructure without accountability.

AI explainability means nothing if the training data was never verified. And EHR data needs a trust score before any AI model trains on it. These are not theoretical positions. They are operational requirements that regulators are beginning to enforce.

The DTI Engine scores every health data record 0 to 100 across 8 trust dimensions before your AI model sees it. If your team is evaluating data for training, compliance, or clinical use, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.

Further reading:

  • DTI™ Engine
  • Health systems solution
  • Infrastructure trust vs data trust: why most healthcare data platforms miss the point
  • Data quality vs data trust: what is the difference and why it matters for healthcare AI
  • The eight dimensions of health data trust: a practical guide
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0–100. Travels with every record permanently.

    See the DTI Engine
    Share
    Why Innovaccer, Datavant, and AWS Health Lake are not solving the data trust problem | SuperTruth