Medicaid managed care data trust: what state reporting requirements demand
Forty-two states now contract with managed care organizations to administer Medicaid benefits, and every one of those states imposes reporting requirements that assume the underlying data is trustworthy. Most MCO data fails that assumption before it reaches CMS. State Medicaid reporting requirements create a compliance surface that cannot be satisfied by volume alone; MCO data quality depends on provenance, recency, and validation that most plans never measure.
Forty-two states run Medicaid through managed care. The data requirements are staggering.
As of 2024, 42 states plus the District of Columbia contract with managed care organizations to deliver Medicaid benefits to more than 57 million enrollees. That figure represents roughly 72% of all Medicaid beneficiaries nationwide. Every one of those state contracts includes reporting mandates that flow upward to CMS, and every one of those mandates assumes the data submitted is accurate, timely, and complete.
The assumption is wrong more often than anyone wants to admit.
Medicaid managed care data trust is not a theoretical concern. It is the foundation on which rate-setting, quality measurement, program integrity, and federal oversight all depend. When MCO data quality fails, it does not just create a compliance gap. It distorts the entire system's understanding of how care is being delivered to the most vulnerable populations in the country.
What CMS actually requires from states and MCOs
The minimum requirement for reporting data to CMS is defined through a layered framework that begins with the Transformed Medicaid Statistical Information System (T-MSIS). T-MSIS collects claims, encounter, eligibility, and provider data from every state Medicaid program. States must submit T-MSIS files monthly, and these files must conform to over 500 data elements across eight file types.
Beyond T-MSIS, CMS now requires three additional managed care reporting tools through the Medicaid Data Collection Tool (MDCT) system:
CMS uses data submitted through MDCT-MCR to monitor compliance with the 2024 Medicaid Access Rule and to enforce the managed care provisions of 42 CFR Part 438. These are not optional filings. States that fail to submit accurate data risk corrective action plans, deferred federal matching funds, and public disclosure of noncompliance.
Does CMS data include Medicaid?
Yes. CMS collects and maintains Medicaid data alongside Medicare data, though through different systems. Medicare data flows primarily through the Chronic Conditions Data Warehouse (CCW) and Medicare claims processing systems. Medicaid data flows through T-MSIS and state-level Medicaid Management Information Systems (MMIS).
The distinction matters because Medicare and Medicaid data have fundamentally different provenance characteristics. Medicare data originates from a single federal payer. Medicaid data originates from 56 different state and territory programs, each with its own eligibility rules, benefit structures, MCO contracts, and reporting timelines. This fragmentation means that Medicaid data trust cannot be assumed from the federal level down. It must be built from the state level up.
For dual-eligible beneficiaries enrolled in both programs, data reconciliation becomes even more complex. Claims may appear in both Medicare and Medicaid systems with different coding, different dates of service, and different provider identifiers. Without a trust layer that scores each record independently, downstream analytics inherit every inconsistency.
The state reporting surface: what MCOs must actually produce
State Medicaid agencies impose reporting requirements on MCOs that go well beyond what CMS collects. A typical state contract requires MCOs to submit:
Each of these data streams has its own submission format, validation rules, and correction windows. Most states use proprietary file layouts that do not align with each other, creating a patchwork of reporting standards that MCOs operating across multiple states must navigate simultaneously.
How many states use MCOs for Medicaid?
Forty-two states and the District of Columbia use managed care organizations to deliver some or all Medicaid benefits. Only a handful of states, including Alaska, Connecticut, Wyoming, and Vermont (for most populations), operate primarily through fee-for-service Medicaid. The trend toward managed care has accelerated since the Affordable Care Act expanded Medicaid eligibility, with states viewing MCOs as a mechanism for predictable budgeting and care coordination.
The scale varies enormously. California contracts with over 20 MCOs. Texas operates through 18 managed care plans across multiple service areas. Florida, New York, and Pennsylvania each have more than a dozen contracted MCOs. In contrast, some states operate with as few as two or three plans covering the entire state.
This variation creates a fundamental MCO data quality challenge. A national MCO like Centene, Molina, or UnitedHealthcare's Community Plan must maintain different reporting configurations for every state contract. A single encounter record may need to be formatted, validated, and submitted differently in Ohio than in Arizona. The same quality metric may use different denominators, different exclusion criteria, and different measurement periods depending on the state.
What states have the worst Medicaid coverage?
Twelve states have not expanded Medicaid under the ACA as of early 2025, including Texas, Florida (partial), Georgia, Mississippi, Alabama, South Carolina, Tennessee, Wisconsin (covers adults up to 100% FPL but not through expansion), Kansas, and Wyoming. These states have the largest coverage gaps, where adults earning between 0% and 138% of the federal poverty level may have no affordable insurance option.
Mississippi, Texas, and Georgia consistently rank lowest on Medicaid coverage metrics, including the percentage of eligible individuals actually enrolled, the breadth of covered services, and provider reimbursement rates relative to Medicare. The Commonwealth Fund's 2024 Scorecard ranked Mississippi, Texas, and Oklahoma at the bottom for healthcare system performance, with Medicaid access as a primary driver.
But "worst coverage" and "worst data" are not the same thing. Some expansion states with robust managed care programs still submit encounter data with error rates above 20%. Some non-expansion states with smaller programs have cleaner data simply because the volume is manageable. The real question is not which states have the worst coverage but which states have data they can trust enough to make decisions on.
Key statistics
The encounter data trust gap
Encounter data is the backbone of Medicaid managed care oversight. States use it to set capitation rates, measure quality, detect fraud, and report to CMS. But encounter data submitted by MCOs consistently fails basic quality checks.
CMS's own Data Quality Atlas has flagged issues including missing diagnosis codes, implausible service dates, duplicate records, and provider identifiers that do not match state enrollment files. A 2023 OIG report found that multiple states could not verify the completeness of encounter data submitted by their MCOs because they lacked the infrastructure to compare MCO submissions against adjudicated claims.
The root cause is structural. MCOs adjudicate claims internally, then transform a subset of that claims data into encounter records for state submission. Every transformation introduces opportunities for data loss, format errors, and timing mismatches. A claim paid on March 1 may not appear as an encounter submission until May. By the time the state validates and submits to T-MSIS, six months may have passed since the service was delivered.
This latency is not just inconvenient. It makes the data unreliable for any time-sensitive analysis, including rate-setting, which directly determines how much federal and state money flows to each MCO. As we have written about in depth, claims data lag creates compounding problems for AI models that depend on recency as a trust dimension.
Why MCO data quality is a trust problem, not just a compliance problem
Compliance frameworks treat data quality as a binary: the file was submitted or it was not. The fields were populated or they were not. The values fell within acceptable ranges or they did not.
This is insufficient. A field can be populated with a value that passes validation but is still wrong. A diagnosis code can be technically valid but clinically implausible given the patient's age, gender, or service history. An encounter record can arrive on time but represent a service that was never delivered.
Medicaid managed care data trust requires measurement across multiple dimensions simultaneously. Provenance: where did this record originate, and through how many transformations did it pass? Recency: how old is this data relative to the event it describes? Concordance: does this record align with other records for the same beneficiary? Validation: has an independent process confirmed the accuracy of key fields?
The Data Trust Index scores every record across exactly these dimensions. Provenance accounts for 25% of the score because in Medicaid managed care, the chain of custody from point of service to MCO system to state MMIS to T-MSIS to CMS is long and lossy. Every handoff is a trust degradation point.
The rate-setting feedback loop
Capitation rates for Medicaid MCOs are set using historical encounter and cost data. If the encounter data understates utilization, rates are set too low, creating pressure on MCOs to restrict services. If encounter data overstates utilization or includes upcoded diagnoses, rates are set too high, wasting taxpayer funds.
This is not hypothetical. Multiple state audits have found that MCOs submitted encounter data with systematically higher acuity codes than the underlying claims supported. The mechanism is similar to what happens in Medicare Advantage risk adjustment, where diagnosis coding directly drives revenue.
In Medicaid managed care, the feedback loop is tighter. States use last year's data to set this year's rates. If the data was untrustworthy last year, this year's rates are wrong by definition. And next year's data will be shaped by the incentives created by this year's rates. Without an independent trust layer that scores data quality before it enters the rate-setting process, the entire cycle operates on an unverified foundation.
SDOH data and the expanding reporting surface
At least 30 states now include social determinants of health requirements in their Medicaid MCO contracts. These range from mandatory SDOH screening using standardized instruments to reporting on closed-loop referrals to community-based organizations.
This expansion of the reporting surface creates new data trust challenges. SDOH data collected through screening instruments has different provenance than claims data. It originates from patient self-report, often captured by front-desk staff or community health workers, and recorded in systems that may not be integrated with the MCO's claims adjudication platform.
The result is a data stream with high variability in completeness, consistency, and accuracy. A food insecurity screening completed in a pediatric clinic has different quality characteristics than one completed via a telephonic health risk assessment. Both may satisfy the contractual requirement to "screen for SDOH," but they carry fundamentally different trust profiles. For more on this specific challenge, see our analysis of social risk factor screening data trust and CBO data trust for Medicaid programs.
What trust scoring changes for Medicaid managed care
The current approach to Medicaid data quality is retrospective. States receive data, run validation checks months later, send error reports to MCOs, and wait for corrections that may arrive in the next quarter. CMS reviews T-MSIS submissions annually and publishes quality findings that are already outdated by the time they are released.
Trust scoring flips this to a prospective model. Every record receives a score at the point of ingestion, before it enters any downstream system. Records below a defined trust floor are flagged for remediation before they contaminate rate-setting models, quality reports, or federal submissions.
This is what the DTI Engine does. It evaluates provenance, consent status, recency, field completeness, concordance with other records, independent validation, breadth of data elements, and stability over time. The output is a 0-to-100 score that tells a state Medicaid agency or MCO exactly how much confidence to place in each record.
For Medicaid managed care, the most critical dimensions are provenance (tracking the chain from point of service through MCO systems to state submission), recency (measuring the gap between service delivery and data availability), and concordance (checking whether encounter records align with eligibility files, provider directories, and quality metrics).
The compliance case for trust infrastructure
CMS's 2024 Medicaid Access Rule significantly expanded reporting requirements for states and MCOs. The rule requires states to conduct secret shopper surveys, publish appointment wait time data, and demonstrate that MCO networks meet quantitative adequacy standards.
All of these requirements depend on data that can be trusted. A secret shopper survey is only meaningful if the provider directory used to select survey targets is accurate. Wait time data is only valid if appointment availability records are current. Network adequacy calculations are only reliable if credentialing data reflects actual provider status.
States that cannot demonstrate the trustworthiness of their underlying data will not be able to satisfy these requirements regardless of how many reports they generate. The volume of reporting is not the bottleneck. The quality of the data feeding those reports is.
What comes next
CMS has signaled that it will use MDCT-MCR data, specifically MCPAR, MLR, and NAAAR submissions, for cross-state comparisons and public transparency reporting. This means MCO data quality will become visible not just to state regulators but to beneficiaries, advocacy organizations, and competing plans.
The states and MCOs that build trust infrastructure now will be positioned to meet these requirements. The ones that continue to rely on retrospective validation and manual correction cycles will fall further behind with each reporting period.
The DTI Engine scores every health data record 0-100 across 8 trust dimensions before your AI model sees it or your compliance team submits it. If your team is managing Medicaid encounter data, preparing for the Access Rule, or building analytics on MCO submissions, talk to the SuperTruth commercial team. Schedule a conversation or call (215) 918-4140.
Further reading:

Jason Alan Snyder
Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.
About SuperTruth · LinkedIn · Substack · jasonalansnyder.com
See it in practice
DTI scores the record, not the patient.
8 dimensions. 0–100. Travels with every record permanently.