The $3.5 trillion cost of bad health data: what fragmentation actually costs the system
Photo by MARIOLA GROBELSKA on Unsplash

The $3.5 trillion cost of bad health data: what fragmentation actually costs the system

By Jason Alan Snyder·May 9, 2026

Bad health data costs the U.S. healthcare system an estimated $3.5 trillion annually through redundant testing, failed care coordination, billing errors, and AI models trained on unverified records. The problem is not a lack of data. It is a lack of trust in the data that already exists.

The U.S. spent $4.8 trillion on healthcare in 2023. Independent analyses from Gartner, AHIP, the National Academy of Medicine, and the Shilling Foundation converge on the same finding: between 25% and 30% of that spending produces no clinical value. That puts the cost of waste, inefficiency, and fragmentation somewhere around $3.5 trillion per year. Not because providers are careless. Because the data underneath every clinical, operational, and financial decision is fractured, stale, or unverifiable.

How bad data costs the US $3 trillion per year

Where the $3.5 trillion in healthcare data waste concentrates
Where the $3.5 trillion in healthcare data waste concentrates

The often-cited figure of $3 trillion in annual waste traces back to overlapping estimates. The National Academy of Medicine identified $765 billion in pure waste categories: unnecessary services, excess administrative costs, fraud, and missed prevention. IBM estimated that bad data alone costs the U.S. economy $3.1 trillion per year across all sectors, with healthcare absorbing a disproportionate share. Gartner pegs the average financial impact of poor data quality at $12.9 million per organization per year.

When you layer healthcare-specific costs on top of those cross-industry numbers, the total climbs. Duplicate testing because records did not transfer. Readmissions because discharge summaries sat in a fax queue. Denied claims because a procedure code was entered wrong three months ago. Each failure traces back to data that was missing, late, inconsistent, or never verified.

What is the real cost of bad data?

The real cost is not just dollars. It is time, trust, and clinical outcomes.

A 2023 survey by KLAS Research found that 69% of health system CIOs ranked data quality as a top-three barrier to AI deployment. Not compute. Not talent. Data. When records carry no provenance, no recency score, no consent verification, every downstream system inherits that uncertainty.

For patients, the cost is visceral. A cancer patient who moves from one health system to another may lose weeks while records are re-requested, re-faxed, and re-entered. We documented this pattern in detail: The fragmented health record: why the most valuable data in healthcare lives nowhere. The clinical delay is real. So is the financial one.

For AI, the cost compounds. Models trained on unscored data absorb every error, every duplicate, every outdated diagnosis code. The output looks confident. The foundation is rotten. That is why EHR data needs a trust score before any AI model trains on it.

What fragmentation in healthcare actually means

Fragmentation in healthcare describes the condition where a single patient's health information is scattered across multiple systems, institutions, and formats with no unifying layer of verification. A patient with three specialists, a primary care physician, a pharmacy, a wearable device, and a lab testing provider may have six or more partial records, none of which reference each other.

This is not a metadata problem. It is a trust problem. Each record fragment may be internally consistent but externally unverifiable. No system confirms whether the lab result from January matches the diagnosis from March or whether the consent given to one provider extends to the AI model another provider is training.

SuperTruth scores this explicitly. The Data Trust Index evaluates every record across eight dimensions, with Provenance weighted at 25%, Consent at 20%, and Recency at 15%. A record that cannot prove where it came from, when it was last updated, or whether the patient authorized its use scores low. Period.

What chronic condition is the most expensive to the US healthcare system?

Diabetes. The American Diabetes Association reported $412.9 billion in total costs attributable to diagnosed diabetes in 2022, including $306.6 billion in direct medical costs. That figure does not include undiagnosed cases or prediabetes, which affect an additional 96 million American adults.

Here is the data problem within the diabetes problem: fragmented records make it nearly impossible to track A1C trends, medication adherence, and comorbidity progression across systems. A patient who sees an endocrinologist, a cardiologist, and a primary care physician in three different health systems has three incomplete pictures. None of those providers can verify what the other two are seeing unless someone manually reconciles the data.

This is precisely where health data integrity for value-based care programs breaks down. Value-based contracts require longitudinal data. Fragmentation destroys longitudinality.

Key statistics

imaware data processing: before and after DTI Engine
imaware data processing: before and after DTI Engine

  • $3.5 trillion: estimated annual cost of healthcare waste driven by data fragmentation, redundancy, and bad data quality in the U.S.
  • $412.9 billion: total cost of diagnosed diabetes in 2022, the most expensive chronic condition in the U.S. healthcare system
  • 69%: share of health system CIOs who rank data quality as a top-three barrier to AI deployment (KLAS Research, 2023)
  • 95%: time reduction SuperTruth achieved standardizing 105,000 diagnostic records for imaware, from 3 weeks to 2 hours
  • 200+ hours/month: ongoing time savings for imaware after DTI Engine deployment, with identification of a patient segment driving 20% of revenue
  • The ROI of fixing data trust

    The imaware case study demonstrates what happens when you apply trust scoring at scale. SuperTruth processed 105,000 diagnostic records, standardized them across the eight DTI dimensions, and reduced processing time from three weeks to two hours. That is a 95% time reduction. The team recovered 200+ hours per month. And the scored data revealed a patient segment responsible for 20% of revenue that had been invisible in the raw records.

    As imaware CEO Brodie Flanders put it: "The lab industry has never had a trust standard. DTI created one."

    The math is straightforward. If bad data costs $12.9 million per organization per year (Gartner's cross-industry average), and trust scoring prevents even 30% of that waste, the ROI is measurable in the first quarter. For health systems deploying AI, the stakes are higher: an unscored training set does not just waste money. It produces clinical outputs no one can audit or defend.

    The DTI Engine scores every health data record 0 to 100 across 8 trust dimensions before your AI model sees it. If your team is evaluating data for training, compliance, or clinical use, schedule a conversation with the SuperTruth commercial team or (215) 918-4140.

    Further reading:

  • DTI™ Engine
  • Health systems solution
  • Siloed health data: the infrastructure problem nobody has solved yet
  • The eight dimensions of health data trust: a practical guide
  • Why AI models trained on unscored health data will fail in production
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0 to 100. Travels with every record permanently.

    See the DTI Engine
    Share