ICD-10 coding accuracy: how billing data becomes a health AI liability
Photo by Denny Müller on Unsplash

ICD-10 coding accuracy: how billing data becomes a health AI liability

By Jason Alan Snyder·July 13, 2026

ICD-10 codes were designed for billing, not clinical truth. When health AI systems train on claims data without verifying diagnostic coding accuracy, they inherit financial incentives as clinical signals, turning reimbursement artifacts into model predictions that affect real patients.

Every diagnosis in the U.S. healthcare system passes through a filter that was never designed for clinical accuracy. It was designed for payment. ICD-10 codes sit at the intersection of clinical documentation and revenue cycle management, and the distortions introduced at that intersection propagate silently into every AI model trained on claims data.

The problem is not that ICD-10 coding is imprecise. The problem is that its precision serves the wrong master. When a coder selects one of 72,000+ ICD-10-CM codes, the selection is shaped by documentation quality, reimbursement incentives, payer-specific rules, and time pressure. The resulting code may correlate with the patient's actual condition. Or it may correlate with what the institution needs to get paid.

Health AI systems that ingest this data without scoring it for diagnostic coding data trust are building on a foundation optimized for cash flow, not clinical reality.

Key statistics

DTI dimensions most affected by billing data distortion
DTI dimensions most affected by billing data distortion

  • The ICD-10-CM system contains over 72,000 diagnosis codes, yet studies consistently show that 20-30% of claims contain coding errors that would change clinical interpretation.
  • CMS estimates that improper payments in Medicare alone exceeded $46.8 billion in FY2023, with coding errors as a primary driver.
  • A 2023 AHIMA survey found that 63% of health information management professionals reported pressure to code for reimbursement optimization over clinical accuracy.
  • SuperTruth's work with imaware standardized 105,000 diagnostic records and reduced data processing time from 3 weeks to 2 hours, a 95% reduction, while uncovering data segments that drove 20% of revenue.
  • Research published by the Journal of General Internal Medicine found that up to 57% of diabetes diagnoses in claims data could not be confirmed by clinical chart review.
  • The billing incentive that distorts clinical truth

    ICD-10 coding exists because payers need a standardized way to process claims. The codes themselves are clinically derived, mapped to the WHO's International Classification of Diseases. But the act of assigning them is an economic decision as much as a clinical one.

    Hospitals and physician practices employ certified coders whose performance is measured partly by claim acceptance rates. When a claim is denied, the coder or clinical documentation improvement (CDI) specialist reviews the chart and often selects a different code. Not because the diagnosis changed, but because the payer's adjudication logic requires a different representation.

    This creates a systematic bias: the codes that survive in billing databases are the codes that got paid, not necessarily the codes that most accurately describe what happened to the patient.

    Upcoding (selecting a more severe diagnosis to increase reimbursement) is illegal but widespread. The Department of Justice recovered over $2.2 billion in healthcare fraud settlements in FY2023, with coding manipulation as a recurring theme. But the more insidious problem is not fraud. It is the everyday, legal optimization of code selection that tilts the entire dataset toward financial utility.

    How billing data becomes a health AI liability

    When AI developers need large-scale diagnostic data, claims databases are the easiest source. They are structured, coded, and available through data aggregators, health information exchanges, and direct payer partnerships. Compared to unstructured clinical notes, claims data is ready to ingest.

    But "ready to ingest" is not the same as "ready to trust."

    Consider what happens when a risk stratification model trains on claims data:

  • Patients with the same clinical condition but different insurance coverage may carry different ICD-10 codes, because their payers enforce different documentation thresholds.
  • Patients who were coded for a condition to justify a test may appear in the dataset as having that condition, even if the test ruled it out.
  • Chronic conditions may appear to resolve and recur based on whether the practice submitted claims for ongoing management, not based on actual disease trajectory.
  • Socioeconomic factors influence coding patterns: patients in safety-net hospitals are coded differently than patients in academic medical centers, not because their diseases differ but because their documentation infrastructure differs.
  • Every one of these distortions enters the training set as signal. The model treats it as clinical truth. And the resulting predictions carry the billing system's biases into clinical decision-making.

    This is not a theoretical concern. The UnitedHealth Group AI denial rate controversy showed what happens when AI systems act on billing-derived logic without adequate data trust checks. Patients were denied coverage based on algorithmic predictions that reflected payer economics, not patient need.

    Is medical billing and coding AI proof?

    No. Medical billing and coding is not immune to AI-driven change, but it is also not a problem that AI can solve by simply automating the existing process. AI coding assistants can increase throughput and catch certain errors. But if the underlying incentive structure rewards code selection based on reimbursement rather than clinical accuracy, automating the process just scales the distortion faster.

    The real question is whether AI can be used to separate the billing signal from the clinical signal. That requires a different kind of system: one that scores the trustworthiness of each data point before it enters a model, rather than one that simply processes codes more efficiently.

    How is AI affecting medical billing and coding?

    AI is reshaping medical billing and coding in three ways, and two of them make the data trust problem worse.

    First, AI-powered coding assistants like those from 3M, Optum, and smaller vendors now suggest ICD-10 codes based on clinical documentation. These tools reduce coder workload and can flag inconsistencies. This is genuinely useful for productivity.

    Second, AI-driven CDI (clinical documentation improvement) tools prompt clinicians to add specificity to their notes, which frequently results in higher-acuity codes. This is marketed as "accuracy improvement," but the accuracy is measured against reimbursement capture, not against independent clinical validation. The tools optimize for payment, not truth.

    Third, payers deploy AI to audit claims and identify potential upcoding or fraud. This creates an adversarial dynamic where provider-side AI optimizes codes upward and payer-side AI pushes them back down. The resulting code is a negotiated outcome, not a clinical determination.

    None of these applications address the fundamental problem: the code that ultimately lands in the billing database reflects a financial negotiation, and downstream AI systems cannot distinguish that code from a pure clinical assertion.

    Can medical coding and billing be replaced by AI?

    Not in its current form, and attempting full replacement without solving the trust problem would be dangerous. Human coders bring contextual judgment that current AI systems lack: understanding when a physician's documentation is ambiguous, knowing when a code technically fits but clinically misleads, recognizing when a payer's rules conflict with accurate representation.

    The AAPC (American Academy of Professional Coders) projects that AI will augment rather than replace coders, shifting their role from code selection to validation and exception handling. This is likely correct for the billing function itself.

    But the more important question is whether the output of billing and coding, whether performed by humans or AI, should be trusted as clinical input for health AI models. The answer is: not without independent verification.

    What are the consequences of inaccurate coding and incorrect billing?

    The consequences cascade across three domains.

    Financially, inaccurate coding triggers claim denials (currently averaging 10-15% of all claims submitted), delayed payments, compliance investigations, and in severe cases, False Claims Act liability with treble damages.

    Clinically, inaccurate codes create false patient histories. A patient coded for Type 2 diabetes to justify an A1C test may carry that diagnosis in perpetuity across every system that ingests claims data. Downstream providers, care coordinators, and AI risk models all see a diabetic patient where none exists.

    For AI systems, inaccurate coding corrupts training data at scale. A model trained on millions of claims records inherits millions of financially motivated code selections as ground truth. The model then reproduces those biases in its predictions, recommendations, and risk scores. As we have written about in the context of data poisoning attacks on health AI, bad training data does not just reduce accuracy. It creates systematic, directional errors that are nearly impossible to detect from model outputs alone.

    The concordance gap: when billing codes contradict clinical records

    Claims data accuracy: billing codes vs clinical chart confirmation
    Claims data accuracy: billing codes vs clinical chart confirmation

    One of the eight dimensions in SuperTruth's Data Trust Index is concordance: the degree to which a data point agrees with other records describing the same patient, encounter, or condition. Billing data fails the concordance test more often than any other data type.

    A patient's EHR may show a problem list with three active conditions. The claims data for the same patient during the same period may show seven or eight diagnosis codes, because every rule-out, every test justification, and every severity modifier generates its own code. The EHR and the claim describe the same encounter but tell different stories.

    When an AI model trains on claims data alone, it sees the seven-code version. When it trains on EHR data alone, it sees the three-condition version. Neither is complete, but the claims version is systematically inflated.

    This is why data quality alone is insufficient. A claims record can be perfectly formatted, fully coded, and delivered on time. It can score well on every traditional data quality metric. And it can still be clinically misleading. Trust requires measuring concordance, provenance, and validation, not just completeness and timeliness.

    What trust scoring does that coding automation cannot

    The top-ranking content for ICD-10 and AI focuses almost exclusively on using AI to improve coding speed and reduce denials. That framing accepts the billing system's incentives as given and asks how to optimize within them.

    SuperTruth asks a different question: before any AI model trains on a record, how trustworthy is that record as a representation of clinical reality?

    The DTI Engine scores every health data record from 0 to 100 across eight dimensions. For billing-derived data, three dimensions are particularly critical:

    Provenance (25% weight): Where did this code originate? Was it assigned by a certified coder from clinical documentation, auto-generated by a CDI tool, or carried forward from a prior encounter without re-verification? Provenance scoring distinguishes between a code that was clinically determined and one that was financially optimized.

    Concordance (10% weight): Does this billing code agree with other data sources for the same patient? If the claims data says diabetes but the lab data shows normal glucose levels across 12 months, the concordance score drops. That signal is invisible to any system that ingests claims data in isolation.

    Validation (10% weight): Has this code been independently verified against a clinical source? A code that has been cross-referenced with pathology results, imaging reports, or physician attestation scores higher than one that exists only in the billing system.

    These dimensions do not fix the billing system. They create a measurable layer of trust between the billing system and the AI systems that consume its output.

    The real-world cost of skipping this step

    The $3.5 trillion U.S. healthcare system runs on data that was generated for billing. That data now feeds AI models that make clinical predictions, determine insurance coverage, identify patients for clinical trials, and stratify populations for value-based care contracts.

    Every one of those use cases assumes that the input data reflects clinical reality. When it does not, the costs multiply. We have documented the fragmentation cost in detail: rework, denials, misallocated care, and failed AI deployments that traced back to data nobody scored before ingestion.

    The fix is not to stop using billing data. Claims databases are too large and too structured to ignore. The fix is to score every record before it enters a model, flag records where billing incentives likely distorted the clinical signal, and enforce trust floors that prevent low-concordance billing data from being treated as ground truth.

    Strategies to ensure accurate coding when using AI tools

    Organizations deploying AI in billing and coding workflows need a dual strategy: improve the coding process and independently verify the output before downstream use.

    For the coding process itself:

  • Require AI coding tools to surface confidence scores alongside suggested codes, so human reviewers can focus attention on uncertain selections.
  • Separate CDI metrics from revenue metrics. When documentation improvement is measured by reimbursement lift, the incentive to upcode is structural.
  • Implement dual-pathway validation where AI-suggested codes are cross-checked against lab results, imaging findings, and medication lists before submission.
  • For downstream data consumers:

  • Never treat claims data as clinical ground truth without concordance checking against at least one independent clinical data source.
  • Apply trust scoring (like the DTI) at the point of data ingestion, before any model training or analytics pipeline consumes the records.
  • Maintain provenance metadata that distinguishes between human-assigned codes, AI-suggested codes, and auto-carried codes from prior encounters.
  • Establish DTI floor thresholds for different use cases: a population health dashboard may tolerate a DTI score of 60, but an FDA regulatory submission should require 85+.
  • The path forward is trust scoring, not better billing

    The healthcare industry has spent decades trying to make billing more accurate. CMS updates ICD-10 codes annually. The OIG audits coding practices. AI tools now assist with code selection. And yet the fundamental tension remains: the codes are selected in a financial context and consumed in a clinical one.

    No amount of coding automation resolves that tension. What resolves it is an independent trust layer that sits between the data source and the data consumer, scoring every record for its reliability as a clinical signal regardless of how accurately it functioned as a billing instrument.

    That is what the Data Trust Index was built to do. Not to replace coders, not to optimize claims, but to answer a question that no billing system can answer about itself: should an AI model trust this record?

    The DTI Engine scores every health data record 0-100 across 8 trust dimensions before your AI model sees it. If your team is evaluating claims-derived data for model training, risk stratification, or regulatory submission, and you need to know which records carry billing distortion versus clinical truth, schedule a conversation with the SuperTruth commercial team or (215) 918-4140.

    Further reading:

  • DTI™ Engine
  • Health systems solution
  • Data quality vs data trust: what is the difference and why it matters for healthcare AI
  • Why EHR data needs a trust score before any AI model trains on it
  • The $3.5 trillion cost of bad health data: what fragmentation actually costs the system
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0 to 100. Travels with every record permanently.

    See the DTI Engine
    Share