Medicare Advantage risk adjustment data integrity: what upcoding means for AI
Photo by Piret Ilver on Unsplash
insight

Medicare Advantage risk adjustment data integrity: what upcoding means for AI

By Jason Alan Snyder·September 8, 2026

Medicare Advantage plans receive an estimated $12-25 billion annually in excess payments tied to risk adjustment coding practices that inflate patient severity. When AI systems train on this data, they inherit the financial incentives baked into every diagnosis code, producing models that confuse billing optimization with clinical reality.

Medicare Advantage plans enrolled over 33 million beneficiaries in 2024, capturing more than half of all Medicare-eligible Americans. CMS pays these plans through a risk adjustment model that ties monthly capitated payments to the documented severity of each enrollee's health conditions. The sicker the patient appears on paper, the higher the payment. That financial structure has created one of the largest data integrity problems in American healthcare.

What is Medicare upcoding?

Medicare upcoding is the practice of documenting diagnosis codes that reflect higher severity or additional conditions beyond what a patient's clinical presentation supports. In fee-for-service Medicare, upcoding typically involves billing a higher-level E&M code for an office visit. In Medicare Advantage, upcoding operates differently and at far greater scale. It centers on Hierarchical Condition Category (HCC) codes that feed the CMS risk adjustment model.

Every HCC code assigned to a Medicare Advantage enrollee increases the plan's monthly payment from CMS. A patient coded with diabetes without complications (HCC 19) generates less revenue than the same patient coded with diabetes with chronic complications (HCC 18). The clinical difference between those two codes can be ambiguous. The financial difference is not.

The Government Accountability Office estimated in 2024 that Medicare Advantage risk adjustment coding led to $12 billion in excess payments in a single year. Other analyses, including work from the Medicare Payment Advisory Commission (MedPAC), have placed the figure closer to $25 billion annually when accounting for coding intensity differences between MA and fee-for-service populations.

What are upcoding diagnoses?

Upcoding diagnoses are specific clinical conditions documented at a severity level higher than the medical evidence supports, or conditions recorded that lack sufficient clinical basis in the patient's chart. Common examples in Medicare Advantage include:

  • Documenting major depressive disorder (recurrent) when the chart supports a single episode
  • Adding chronic kidney disease stage III when lab values show stage II
  • Recording heart failure with reduced ejection fraction when the echocardiogram shows preserved function
  • Listing vascular disease based on peripheral artery disease risk factors rather than confirmed diagnosis
  • What is an example of upcoding? A Medicare Advantage plan conducts an in-home health risk assessment on a 72-year-old enrollee. The physician assistant documents diagnoses of diabetes with neuropathy, chronic obstructive pulmonary disease with exacerbation, and protein-calorie malnutrition. The patient's primary care records show well-controlled type 2 diabetes, mild intermittent asthma, and a BMI of 24 with no weight loss. Each upgraded diagnosis generates a higher HCC code and higher monthly payments. The clinical record does not support the severity documented. That is upcoding.

    The scale of the risk adjustment data integrity problem

    CMS applies a coding intensity adjustment to Medicare Advantage payments, currently reducing aggregate payments by 5.9%, to account for the known tendency of MA plans to code more aggressively than fee-for-service providers. MedPAC has argued this adjustment is insufficient, estimating that MA coding generates risk scores 7-10% higher than clinically equivalent fee-for-service populations.

    The DOJ has pursued multiple False Claims Act cases targeting MA risk adjustment practices. In 2024 and 2025, settlements and ongoing litigation involved UnitedHealth Group's Optum, Kaiser Permanente, and several smaller plans. The legal theory is straightforward: submitting diagnosis codes unsupported by medical records to inflate risk adjustment payments constitutes a false claim against the federal government.

    But the legal problem is only the surface. The deeper problem is what happens when this data becomes training material for AI.

    Key statistics

    Medicare Advantage coding intensity vs CMS adjustment
    Medicare Advantage coding intensity vs CMS adjustment

  • Medicare Advantage risk adjustment coding generates an estimated $12-25 billion in annual excess payments above fee-for-service equivalent costs (GAO, MedPAC)
  • Over 33 million Medicare beneficiaries are enrolled in MA plans as of 2024, representing more than 51% of all Medicare-eligible Americans
  • CMS applies a 5.9% coding intensity adjustment to MA payments, but MedPAC estimates actual coding intensity runs 7-10% higher than fee-for-service
  • In-home health risk assessments, used by MA plans to capture additional HCC codes, increased from 2.6 million in 2017 to over 5 million in 2023
  • SuperTruth's DTI Engine processing of 105,000 diagnostic records for imaware reduced standardization time from 3 weeks to 2 hours, a 95% reduction, while identifying trust-score failures that would have propagated into downstream models
  • How upcoded data poisons AI models

    AI models trained on Medicare Advantage claims data inherit the financial incentives embedded in every record. A risk prediction model trained on MA data will systematically overestimate disease severity for conditions where upcoding is most prevalent: diabetes complications, heart failure classifications, chronic kidney disease staging, and major depressive disorder.

    This is not a theoretical concern. Health plans are already deploying AI-driven chart review software that scans medical records for "missed" HCC codes. These tools are marketed as ensuring coding accuracy. In practice, they operate as one-directional systems: they find codes to add, rarely codes to remove. The AI identifies potential diagnoses that could be documented but were not, and surfaces them for clinician attestation.

    The result is a feedback loop. Upcoded data trains models. Models identify more opportunities to code aggressively. Aggressively coded records become the new training data. Each cycle amplifies the gap between documented severity and clinical reality.

    A readmission prediction model trained on this data will flag patients as higher risk than they clinically are. A population health stratification model will over-allocate resources to conditions that are over-documented rather than under-treated. A clinical decision support tool will recommend interventions calibrated to inflated baselines.

    Will medical billing and coding get replaced by AI?

    Not entirely, but AI is already transforming how coding happens, and not always in the direction of accuracy. Natural language processing tools scan clinical notes and suggest diagnosis codes. Computer-assisted coding software has been in use for over a decade. The newer generation of AI-powered tools goes further, reviewing entire charts and identifying diagnosis gaps that translate to revenue opportunities.

    The question is not whether AI will replace human coders. The question is whether AI will replace the judgment that human coders are supposed to exercise. A human coder reading a chart can weigh clinical context, recognize when documentation is ambiguous, and apply specificity requirements that prevent unsupported code assignments. An AI system optimized for revenue capture does not exercise that judgment. It finds patterns that correlate with higher payments.

    CMS has signaled awareness of this problem. The 2025 MA rate notice included language about algorithmic coding and the agency's intent to evaluate whether AI-assisted coding tools are contributing to coding intensity. But regulation lags deployment by years. Plans are using these tools now, generating coded data now, and that data is flowing into AI training pipelines now.

    The trust gap in MA coding data

    MA coding data trust fails across multiple dimensions that any serious data integrity framework should measure.

    Provenance is compromised when codes originate from in-home health risk assessments conducted by contracted vendors rather than treating clinicians. The provider who documents the diagnosis may have no prior relationship with the patient, no access to longitudinal records, and a financial incentive tied directly to the number of HCC codes captured.

    Recency collapses when plans persist diagnosis codes across calendar years without clinical re-validation. A patient coded with major depression in 2022 may continue carrying that HCC code in 2024 because no one re-evaluated whether the condition is still active. The code persists because removing it reduces revenue.

    Concordance breaks when the diagnosis documented in the risk adjustment submission does not match the clinical evidence in the EHR. A patient's primary care physician may document "prediabetes" while the MA plan's retrospective chart review reclassifies the same lab values as "diabetes with complications."

    Validation is absent when no independent party confirms that submitted codes reflect the clinical record. CMS conducts Risk Adjustment Data Validation (RADV) audits, but the program has been mired in legal challenges and methodological disputes for over a decade. The audit rate covers a fraction of submitted diagnoses.

    Why standard data quality checks miss the problem

    Traditional data quality frameworks evaluate whether a code is syntactically valid, whether required fields are populated, and whether the claim passes basic edit checks. Upcoded data passes all of these tests. The diagnosis code is a real ICD-10 code. The patient is a real enrollee. The claim is properly formatted. The provider has valid credentials.

    The problem is not that the data is malformed. The problem is that the data is financially motivated in ways that distort its clinical meaning. No standard ETL pipeline or data warehouse validation rule catches this. The code E11.65 (type 2 diabetes with hyperglycemia) is syntactically identical whether it reflects genuine uncontrolled diabetes or an aggressive interpretation of a single elevated fasting glucose.

    This is why data trust requires more than data quality. Quality asks: is the data complete and well-formed? Trust asks: does this record reflect what actually happened to this patient? Those are fundamentally different questions, and the distinction matters enormously when AI models consume the answers.

    What the risk adjustment problem means for downstream AI applications

    Every AI application that consumes diagnosis data from Medicare Advantage populations inherits the upcoding signal. The specific downstream failures include:

    Population health analytics. Risk stratification models that use HCC scores or diagnosis counts will systematically overestimate the burden of disease in MA populations compared to fee-for-service populations. Organizations operating across both will see artificial performance differences driven by coding practices, not clinical reality.

    Clinical AI development. Models trained on MA claims data for tasks like predicting deterioration, identifying care gaps, or recommending interventions will calibrate to an inflated baseline. When deployed in settings with different coding practices, these models underperform because the population they trained on was sicker on paper than in person.

    Real-world evidence generation. Pharmaceutical companies and regulatory bodies increasingly rely on claims data for real-world evidence. If the underlying diagnosis codes are inflated, the prevalence estimates, comorbidity profiles, and treatment outcome measurements derived from them are distorted. An RWE study on heart failure outcomes that uses MA data without accounting for coding intensity will produce systematically biased results.

    Value-based care measurement. ACO REACH and other CMS Innovation Center models use risk adjustment to set benchmarks. If the underlying coding data is inflated, benchmarks shift, and the measurement of whether a program saved money or improved outcomes becomes unreliable. This connects directly to the data quality requirements in ACO REACH and the broader CMS Innovation Center value-based model data requirements.

    Scoring MA data before AI touches it

    DTI dimension weights applied to Medicare Advantage risk adjustment data
    DTI dimension weights applied to Medicare Advantage risk adjustment data

    The only way to prevent upcoded MA data from corrupting AI models is to evaluate every record before it enters a training pipeline, an analytics layer, or a decision support system. That evaluation needs to go beyond format checks and syntactic validation.

    The Data Trust Index scores every health data record across eight dimensions: Provenance (25%), Consent (20%), Recency (15%), Quality (10%), Concordance (10%), Validation (10%), Breadth (5%), and Stability (5%). For Medicare Advantage risk adjustment data, three dimensions carry particular weight.

    Provenance scoring evaluates whether the diagnosis originated from a treating clinician with longitudinal knowledge of the patient or from a contracted vendor conducting a one-time assessment. A diagnosis documented by a patient's primary care physician who has seen them quarterly for three years scores differently than the same code from an in-home assessment vendor.

    Concordance scoring compares the diagnosis code against other data elements in the record. Does the diabetes complication code align with lab values, medication lists, and specialist referral patterns? Or does it stand alone, unsupported by corroborating clinical evidence?

    Recency scoring checks whether the diagnosis has been clinically re-validated within an appropriate timeframe. A condition documented once in 2022 and carried forward without clinical encounter in 2023 or 2024 scores lower than the same condition with ongoing documentation.

    When SuperTruth processed 105,000 diagnostic records for imaware, the standardization and trust scoring pipeline identified records that would have passed conventional quality checks but failed concordance and provenance thresholds. These are the records that, without trust scoring, flow directly into models and distort every output.

    The regulatory trajectory

    CMS is tightening. The 2025 MA rate notice expanded the RADV audit methodology. The DOJ's False Claims Act pipeline continues to grow, with several major cases in active litigation against plans and their coding vendors. Congressional attention to MA overpayments has increased, with MedPAC recommendations to Congress calling for more aggressive coding intensity adjustments.

    But regulation alone cannot solve the data trust problem for AI. Audits happen after the fact. AI training happens continuously. By the time a RADV audit identifies upcoded diagnoses in a plan's 2023 submissions, those codes have already been used to train models, generate analytics, and inform clinical decisions across multiple organizations.

    The only defense is a trust layer that operates at the point of data ingestion, before any model, any analyst, or any decision support system consumes the record. That is what the DTI Engine provides.

    As noted in our analysis of ICD-10 coding accuracy and its implications for health AI, the gap between billing data and clinical reality is not a bug in the system. It is a feature of a payment model that rewards documentation over health. Any AI system that ignores this gap is building on a foundation designed for a purpose other than truth.

    The DTI Engine scores every health data record 0-100 across 8 trust dimensions before your AI model sees it. If your team is evaluating Medicare Advantage data for training, compliance, risk adjustment validation, or clinical AI deployment, talk to the SuperTruth commercial team. Schedule a conversation or call (215) 918-4140.

    Further reading:

  • DTI™ Engine
  • Health plans solution
  • ICD-10 coding accuracy: how billing data becomes a health AI liability
  • ACO REACH data quality requirements and trust infrastructure needs
  • Optum data assets and the trust question for competing health systems
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0–100. Travels with every record permanently.

    See the DTI Engine
    Share
    Medicare Advantage risk adjustment data integrity: what upcoding means for AI | SuperTruth