Hospital Compare data quality: what public reporting gets wrong about outcomes
Photo by Indra Projects on Unsplash
insight

Hospital Compare data quality: what public reporting gets wrong about outcomes

By Jason Alan Snyder·September 11, 2026

Hospital Compare publishes quality data on over 4,000 hospitals, but the underlying data suffers from coding inconsistencies, risk adjustment gaps, and reporting lags that distort the outcomes consumers see. CMS public reporting data trust depends on layers of transformation that few users understand, and fewer question.

CMS Hospital Compare, now folded into the Care Compare portal, publishes quality metrics on more than 4,000 hospitals across the United States. Millions of patients and families use these star ratings and outcome measures to decide where to receive care. The problem is that the data underlying those ratings is unreliable in ways that are rarely visible to the consumer.

This is not a theoretical complaint. Coding inconsistencies, risk adjustment model limitations, reporting lags measured in years, and structural biases against safety-net hospitals all contribute to a public reporting system that presents certainty where uncertainty exists. The result is a system that looks like transparency but functions more like a distortion mirror.

Where does Hospital Compare obtain the data?

Hospital Compare draws from multiple data sources, each with its own quality limitations.

Medicare claims data forms the backbone of most outcome measures, including 30-day mortality rates, readmission rates, and complication rates. These claims are generated through billing processes, not clinical documentation. The data was created to secure payment, not to measure quality.

Hospital-reported quality measures come through the Inpatient Quality Reporting (IQR) program. Hospitals submit these measures to CMS to avoid payment reductions. The incentive structure means hospitals are motivated to report, but the motivation is financial compliance, not data accuracy.

HCAHPS (Hospital Consumer Assessment of Healthcare Providers and Systems) surveys provide patient experience data. These surveys reach a sample of discharged patients and carry well-documented response biases tied to age, health literacy, and primary language.

Structural measures, such as whether a hospital uses electronic health records or has specific clinical protocols, are self-reported. CMS does not independently verify most of these submissions.

What are some examples of poor data quality in healthcare?

The examples are not hypothetical. They are structural and well-documented.

Coding variation across hospitals. Two hospitals treating identical patients can produce different mortality statistics based solely on how their coders capture diagnoses. A hospital with more aggressive documentation improvement programs will code more comorbidities, which feeds into risk adjustment models and can paradoxically make their outcomes look better. A hospital with simpler coding practices will appear to have a sicker-than-expected patient population or, worse, will appear to have higher-than-expected mortality because the risk adjustment did not account for true severity.

Research published in JAMA Internal Medicine found that hospitals with higher rates of coding for certain conditions, like sepsis, showed apparent mortality improvements that were driven entirely by coding changes rather than actual clinical improvement.

Present-on-admission indicator errors. Patient safety indicators depend on distinguishing complications that occurred during a hospital stay from conditions present at admission. The accuracy of present-on-admission coding varies widely. A 2019 study found error rates in POA coding ranging from 5% to over 25% depending on the condition and the hospital.

Claims lag. Medicare claims data used for Hospital Compare outcomes reporting typically reflects performance periods that ended 12 to 36 months before publication. A hospital that implemented major quality improvements in 2023 may not see those improvements reflected in public data until 2025 or later. The consumer looking at Care Compare today is making decisions based on data from a hospital that may no longer operate the same way.

Survey response bias. HCAHPS response rates average around 25% to 30%. Patients who are older, English-speaking, and white are overrepresented. Hospitals serving diverse, lower-income populations systematically receive lower patient experience scores that reflect population characteristics more than care quality.

Key statistics

Hospital Compare data quality gaps by source type
Hospital Compare data quality gaps by source type

  • Hospital Compare star ratings use data with reporting periods that lag 12 to 36 months behind the current date.
  • HCAHPS survey response rates average 25% to 30%, with documented demographic skew in who responds.
  • Present-on-admission coding error rates range from 5% to 25% across hospitals and conditions.
  • CMS uses claims from approximately 12 million Medicare fee-for-service hospital stays annually to calculate outcome measures.
  • SuperTruth's work with imaware standardized 105,000 diagnostic records in 2 hours, a process that previously took 3 weeks, demonstrating what data trust infrastructure can do when applied to healthcare data at scale.
  • The risk adjustment problem

    Risk adjustment is supposed to level the playing field. The idea is that a hospital treating sicker patients should not be penalized for having higher mortality rates. CMS uses hierarchical logistic regression models that adjust for patient age, sex, and comorbid conditions derived from claims data.

    The flaw is circular. The comorbid conditions used for risk adjustment come from the same coding process that varies across hospitals. A hospital with a robust clinical documentation improvement (CDI) program captures more secondary diagnoses, receives stronger risk adjustment, and appears to have better-than-expected outcomes. A hospital without CDI resources, often a smaller or safety-net facility, gets weaker risk adjustment and looks worse.

    This is not a minor effect. Academic medical centers and large health systems with dedicated CDI teams consistently outperform community hospitals on risk-adjusted metrics. Some of that gap reflects genuine quality differences. Some of it reflects coding infrastructure differences. Hospital Compare does not distinguish between the two.

    The risk adjustment models also do not incorporate socioeconomic status, functional status, or social determinants of health in most outcome measures. A hospital serving a population with high rates of housing instability, food insecurity, and limited transportation access will have higher readmission rates that reflect social conditions, not clinical failure.

    What are the 5 most important quality indicators in a hospital?

    CMS organizes Hospital Compare measures into categories that represent the most commonly referenced quality indicators:

  • Mortality rates. 30-day risk-standardized mortality for conditions including heart attack, heart failure, pneumonia, COPD, stroke, and coronary artery bypass graft surgery.
  • Readmission rates. 30-day risk-standardized unplanned readmission rates for the same condition categories.
  • Patient safety indicators. Composite measures including hospital-acquired infections (CLABSI, CAUTI, SSI, MRSA, C. diff), patient safety for selected indicators (PSI-90), and falls.
  • Patient experience. HCAHPS survey domains covering communication with doctors, communication with nurses, responsiveness of hospital staff, pain management, cleanliness, quietness, discharge information, and overall rating.
  • Timely and effective care. Process measures including median time from arrival to departure for ED patients, appropriate use of prophylactic antibiotics, and sepsis bundle compliance.
  • Each of these categories carries known data quality issues. Mortality and readmission measures depend on claims coding accuracy. Patient safety indicators depend on POA coding and electronic surveillance definitions that vary across EHR systems. Patient experience depends on survey response rates and population characteristics. Process measures depend on hospital self-reporting.

    The star rating aggregation problem

    CMS condenses all of these measures into a single overall star rating from 1 to 5 stars. The methodology uses latent variable modeling to group measures, then calculates a weighted summary score that gets converted to stars using a clustering algorithm.

    This approach creates several data trust problems.

    First, the weighting is opaque to most consumers. Mortality measures receive the highest weight, but a consumer looking at a 4-star hospital has no way to know whether that hospital excels at keeping patients alive but has terrible infection rates, or vice versa.

    Second, the clustering algorithm means that the difference between a 3-star and 4-star hospital can be tiny in absolute terms. A hospital could move from 3 to 4 stars based on marginal improvements in a single measure domain while experiencing real deterioration in another.

    Third, hospitals with fewer eligible measures, often smaller and rural facilities, are excluded from star ratings entirely. Over 700 hospitals do not receive star ratings because they lack sufficient data. These tend to be the hospitals serving the most vulnerable populations, the exact populations that most need quality information.

    What is one important result of increased transparency and public reporting in healthcare?

    The most significant documented result is that public reporting has driven hospitals to invest in documentation and coding accuracy at least as much as it has driven actual clinical improvement.

    This is not entirely negative. Coding accuracy matters. Better documentation supports better care coordination, more accurate billing, and more reliable research. But when the primary response to public reporting is to hire more coders rather than more nurses, the transparency system is incentivizing the wrong behavior.

    Studies have shown measurable reductions in specific targeted conditions following public reporting. CLABSI (central line-associated bloodstream infections) rates declined significantly after CMS began publicly reporting them. But researchers have also documented definitional gaming, where hospitals reclassify infections using alternative definitions or attribute them to mucosal barrier injury rather than true CLABSI.

    The honest summary: public reporting has made hospitals more aware of measured outcomes and more sophisticated in how they present data. Whether it has made patients meaningfully safer is a question the current data cannot definitively answer, because the measurement system itself is subject to the same gaming pressures it was designed to counteract.

    The structural bias against safety-net hospitals

    Safety-net hospitals serve disproportionately higher shares of Medicaid patients, uninsured patients, and patients with complex social needs. Hospital Compare data systematically disadvantages these hospitals in at least three ways.

    First, as noted above, risk adjustment models do not account for social determinants. A hospital whose patients cannot afford medications, lack stable housing, or cannot arrange follow-up transportation will have higher readmission rates regardless of the quality of inpatient care provided.

    Second, safety-net hospitals have fewer resources for CDI programs, quality reporting infrastructure, and survey optimization. The administrative requirements of public reporting create a fixed cost that hits smaller and under-resourced hospitals hardest.

    Third, patient experience scores on HCAHPS are correlated with hospital resources and physical environment. Older facilities with shared rooms, limited parking, and fewer amenities receive lower ratings that reflect capital investment, not nursing quality.

    The net effect is a public reporting system that steers patients toward well-resourced hospitals and away from safety-net facilities, potentially accelerating the financial distress of the hospitals that serve the most vulnerable populations.

    Why CMS public reporting data trust requires scoring, not just publication

    Data Trust Index: 8 dimensions and their weights
    Data Trust Index: 8 dimensions and their weights

    The core failure of Hospital Compare is that it publishes data without scoring the trustworthiness of that data. A mortality rate calculated from complete, well-coded claims at a hospital with a mature CDI program is presented identically to a mortality rate calculated from incomplete, variably coded claims at a hospital without CDI resources.

    This is the difference between data quality and data trust. Quality asks whether the data meets a defined standard. Trust asks whether the data can be relied upon for a specific decision.

    A consumer choosing a hospital for heart surgery needs to know not just what the mortality rate is, but how much confidence to place in that number. Is the underlying coding consistent? How complete is the claims capture? How recent is the data? Does the risk adjustment account for the patient population this hospital actually serves?

    None of these questions are answered by the current Hospital Compare interface. The star rating presents a single number with false precision, implying a level of certainty the underlying data does not support.

    This is exactly the problem that data trust scoring addresses. The Data Trust Index (DTI) evaluates every health data record across 8 dimensions: Provenance, Consent, Recency, Quality, Concordance, Validation, Breadth, and Stability. Applied to hospital outcomes data, DTI would flag the records where coding variation introduces noise, where reporting lags make the data stale, and where risk adjustment gaps mean the published metric does not mean what consumers think it means.

    When SuperTruth worked with imaware to standardize 105,000 diagnostic records, the process revealed data integrity issues that were invisible at the surface level. Records that appeared complete had concordance failures. Data that looked current had provenance gaps. The same dynamic exists in Hospital Compare data at a much larger scale, affecting millions of patient decisions annually.

    What needs to change

    Hospital Compare will not become trustworthy through better star rating algorithms or more measures. It will become trustworthy when every published metric carries a confidence indicator that reflects the quality of the underlying data.

    This means scoring the data before it becomes a public metric. It means flagging hospitals where coding variation makes risk-adjusted outcomes unreliable. It means disclosing reporting lag in a way consumers can understand. It means adjusting for social determinants or, at minimum, disclosing when adjustment is absent.

    The alternative is the current state: a transparency system that millions of patients trust, built on data that the people who produce it know is unreliable. That gap between perceived trustworthiness and actual data integrity is exactly where harm occurs.

    The DTI Engine scores every health data record 0 to 100 across 8 trust dimensions before your AI model sees it. If your team is building on CMS public reporting data, evaluating hospital outcomes for network design, or training models on claims-derived quality metrics, the question is not whether the data is available. The question is whether the data is trustworthy. Schedule a conversation with the SuperTruth commercial team or call (215) 918-4140.

    Further reading:

  • DTI™ Engine
  • Health systems solution
  • ICD-10 coding accuracy: how billing data becomes a health AI liability
  • Readmission prediction model bias: how training data trust affects clinical AI
  • Health data completeness scoring: what missing fields cost AI model performance
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    EHR data scored before any AI model sees it.

    DTI integrates with Epic, Oracle Health, and all major EHR systems.

    See our health systems solution
    Share