Cost-effectiveness analysis data trust: what health economics models need
Cost-effectiveness analysis models in healthcare depend on data inputs that most organizations never verify. QALY calculations, transition probabilities, and utility weights all carry embedded assumptions about data quality, yet fewer than 15% of published CEA models report data provenance. Without a trust layer beneath the inputs, the outputs are economic fiction.
Every cost-effectiveness analysis (CEA) model produces a number. An incremental cost-effectiveness ratio. A cost per QALY gained. A budget impact estimate. Payers, HTA bodies, and formulary committees treat these numbers as decision-grade evidence. But the numbers are only as trustworthy as the data that generated them.
The problem is specific: health economics models consume data from dozens of sources, apply transformations that are rarely documented, and produce outputs that carry false precision. A model might report that a new therapy costs $47,200 per QALY gained, implying a level of certainty that the underlying data does not support. When the utility weights come from a 2009 study of 83 patients in a single center, when the transition probabilities are borrowed from a trial population that does not match the target population, and when the cost inputs reflect charges rather than actual payments, the ICER output is an artifact of compounded data quality failures.
This is a data trust problem. And it has direct consequences for which treatments patients receive.
What cost-effectiveness analysis actually requires from data
A standard CEA model, whether it is a decision tree, Markov model, or microsimulation, requires four categories of input data:
Clinical effectiveness data. Transition probabilities between health states. Relative risk reductions. Event rates. These typically come from clinical trials, meta-analyses, or observational studies.
Utility weights. Health-related quality of life values assigned to each health state, measured on a 0-to-1 scale where 1 represents perfect health and 0 represents death. These are used to calculate QALYs. They come from patient surveys using instruments like the EQ-5D, HUI, or SF-6D.
Cost data. Direct medical costs (hospitalizations, procedures, drug acquisition), direct non-medical costs (transportation, caregiver time), and sometimes indirect costs (productivity losses). Sources include claims databases, hospital charge masters, Medicare fee schedules, and micro-costing studies.
Time horizon and discounting parameters. The model's structural assumptions about how long to project outcomes and at what rate to discount future costs and health benefits.
Each of these categories has distinct data trust requirements. And each one fails in predictable ways.
QALY data requirements and the utility weight problem
The QALY is the standard outcome measure in cost-effectiveness analysis. One QALY equals one year of life lived in perfect health. The calculation multiplies survival time by a utility weight.
The utility weight is where data trust breaks down most often.
A 2020 systematic review in PharmacoEconomics found that the choice of utility elicitation instrument can change the cost-effectiveness result by 30% or more for the same intervention. EQ-5D-3L produces systematically different values than EQ-5D-5L. The SF-6D generates different utility values than the HUI3 for the same health state. These are not minor methodological differences. They change whether an intervention falls above or below a willingness-to-pay threshold.
Most CEA models do not generate original utility data. They borrow values from published literature. A 2021 analysis in Value in Health found that over 60% of utility weights used in NICE technology appraisals were sourced from studies conducted in different countries, different patient populations, or different time periods than the target analysis. The provenance chain from original patient survey to final model input is rarely documented.
This matters because a utility weight of 0.72 for moderate heart failure means something different when it comes from a UK population sample in 2005 versus a US clinical trial population in 2019. The number looks the same. The data behind it is not the same.
Health economics data quality starts with provenance
Provenance is the single most important trust dimension for CEA data, and it is the most commonly missing.
When ICER (the Institute for Clinical and Economic Review) evaluates a manufacturer's economic model, one of their first assessments is whether the data sources are transparent, appropriate, and well-documented. But even ICER's own reviews frequently note that manufacturers use utility values, cost inputs, or transition probabilities from sources that are poorly matched to the decision context.
The SuperTruth Data Trust Index weights provenance at 25% of the total trust score for exactly this reason. For health economics data specifically, provenance answers three questions:
A transition probability derived from a phase III clinical trial in treatment-naive patients has different provenance characteristics than one derived from a retrospective claims analysis in a mixed population. Both might be reported as "the probability of disease progression is 0.15 per cycle." Without provenance documentation, the model user cannot distinguish between them.
The cost data trust gap in economic evaluation
Cost data in CEA models is notoriously unreliable. A recent MedPageToday piece noted that healthcare rationing is universal but often hidden behind opaque economic justifications. The cost inputs that drive those justifications deserve scrutiny.
Consider the common practice of using Medicare reimbursement rates as a proxy for costs. Medicare pays hospitals based on DRG codes, not actual resource consumption. A DRG payment for a hip replacement is the same whether the patient required 2 days or 7 days of post-operative care. Using DRG payments as cost inputs systematically misrepresents the actual cost of treating patients who are sicker or healthier than average.
Claims-based cost data also suffers from lag. As we have covered previously, claims data can carry 30-to-90-day reporting delays. When a CEA model uses claims data from a specific time period, the analyst must verify that the data is complete for that period. Incomplete claims data biases cost estimates downward because late-arriving high-cost claims are missing.
Charge-to-cost ratios introduce another layer of distortion. Hospital charges bear little relationship to actual costs. The ratio used to convert charges to costs varies by department, by hospital, and by year. A 2019 study in Health Affairs found that charge-to-cost ratios varied by a factor of 3 across hospitals within the same state.
Why model structure assumptions need trust scoring
Beyond input data, the structural assumptions of a CEA model carry their own trust requirements.
A Markov model assumes that the probability of transitioning between health states depends only on the current state, not on how the patient arrived there. This "memoryless" property is a simplification. For chronic diseases where treatment history matters, this assumption introduces bias. The modeler's choice to use a Markov structure rather than a microsimulation is itself an implicit data quality decision, because it determines how much patient-level heterogeneity the model can capture.
Cycle length affects results. A model with annual cycles misses events that happen within the year. A model with monthly cycles requires monthly transition probabilities, which often do not exist in the source data and must be converted from annual rates using mathematical transformations that assume constant hazard rates.
Time horizon selection is another trust-relevant decision. A 5-year time horizon for a chronic disease model will produce different cost-effectiveness results than a lifetime horizon. The choice of time horizon should be driven by the disease's natural history and the expected duration of treatment effect, but it is often driven by data availability.
Key statistics
The following data points quantify the scope of the data trust problem in cost-effectiveness analysis:
How AI is changing the data trust requirement for CEA
AI-generated cost-effectiveness models are already arriving. Large language models can draft model structures, populate parameter tables from published literature, and run probabilistic sensitivity analyses. A recent MedPageToday article raised the concern that AI may drive health costs up rather than down, partly because AI systems can generate economic justifications for expensive interventions faster than humans can critically evaluate them.
This makes the data trust requirement more urgent, not less. When a human analyst builds a CEA model over six months, they develop intuition about which data sources are reliable and which are not. When an AI system populates a model in six minutes by scraping parameters from published literature, it has no mechanism for assessing whether a utility weight from a 2007 study of 34 patients is appropriate for a 2026 coverage decision.
The MedPageToday coverage on equitable AI development raises a parallel concern. If CEA models are built using data that systematically underrepresents certain populations, the resulting cost-effectiveness estimates will be biased. Utility weights derived from predominantly White, higher-income study populations may not reflect the quality-of-life experience of patients in underserved communities. The economic evaluation then encodes that bias into coverage and reimbursement decisions.
What a trust-scored CEA data pipeline looks like
A trust-scored approach to CEA data would evaluate every model input across the eight dimensions of the Data Trust Index before it enters the model.
Provenance (25% weight): Is the data source documented? Is the original study accessible? Has the data been through identifiable transformations?
Consent (20% weight): Were the patients whose data generated the utility weights or clinical outcomes appropriately consented for this use? This is especially relevant when patient-level data from registries or EHRs feeds directly into microsimulation models.
Recency (15% weight): Are the cost data current? Are the clinical effectiveness estimates from trials that reflect current standard of care? A transition probability from a 2010 trial may not apply to a 2026 treatment landscape.
Quality (10% weight): Are there missing values, implausible ranges, or coding errors in the source data?
Concordance (10% weight): Do the data sources agree with each other? If two studies report different utility weights for the same health state, that discordance should be flagged.
Validation (10% weight): Has the data been independently verified? Has the model been externally validated against observed outcomes?
Breadth (5% weight): Does the data cover the relevant subgroups? A model that uses data from a single site cannot claim generalizability.
Stability (5% weight): Has the data source produced consistent values over time, or does it fluctuate in ways that suggest measurement problems?
A DTI score of 70 or above for every major model input would represent a meaningful improvement over current practice, where most inputs have no quality assessment at all.
The ICER connection and payer decision-making
ICER's value assessments directly influence payer coverage decisions in the United States. Their models consume the same categories of data described above. When ICER reports that a therapy's cost per QALY falls between $100,000 and $150,000, that range reflects uncertainty in the input data. But the range is typically generated through probabilistic sensitivity analysis, which varies parameters within assumed distributions. It does not assess whether the parameters themselves are trustworthy.
A data trust layer beneath ICER-style models would distinguish between uncertainty that comes from natural variability (which sensitivity analysis handles) and uncertainty that comes from data quality failures (which it does not). A utility weight with high variance because patient preferences genuinely differ is a different problem than a utility weight with unknown provenance because it was copied from a conference abstract that cannot be verified.
As a recent MedPageToday opinion piece argued, delaying disease recurrence saves both lives and healthcare costs. But quantifying those savings requires cost-effectiveness models built on trustworthy data. If the cost of recurrence treatment is estimated from outdated charge data, or the utility decrement of recurrence is borrowed from an unrelated cancer type, the economic argument for the intervention is weaker than it appears.
Practical steps for health economics teams
Health economics teams building or reviewing CEA models can apply data trust principles immediately:
The cost of not trusting your CEA data
Bad data in a cost-effectiveness model does not just produce wrong numbers. It produces wrong decisions. A therapy that appears cost-effective at $50,000 per QALY based on optimistic utility weights and outdated cost data might actually cost $120,000 per QALY with current, verified inputs. The difference determines formulary placement, prior authorization requirements, and patient access.
Health plans making coverage decisions based on CEA models with unscored data are accepting unknown risk. Manufacturers submitting economic models to HTA bodies without provenance documentation are inviting rejection or revision. And patients are living with the consequences of economic evaluations built on data that nobody verified.
The fix is not more sophisticated modeling. It is more trustworthy data.
SuperTruth's trust layer turns RWE from a compliance liability into a competitive asset. If your team needs audit-ready provenance for FDA submission or payer negotiation, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140. The DTI Engine scores every health data record 0-100 across 8 trust dimensions before your economic model sees it. For teams building or evaluating cost-effectiveness models, this is the difference between an ICER that reflects reality and one that reflects accumulated data quality failures.
Further reading:

Jason Alan Snyder
Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.
About SuperTruth · LinkedIn · Substack · jasonalansnyder.com
See it in practice
DTI scores the record, not the patient.
8 dimensions. 0–100. Travels with every record permanently.