Utilization management data integrity: what concurrent review data needs
Photo by Solen Feyissa on Unsplash
insight

Utilization management data integrity: what concurrent review data needs

By Jason Alan Snyder·September 17, 2026

Concurrent review decisions depend on data that changes hour by hour, yet most utilization management systems treat clinical records as static inputs. When 30% of UM denials trace back to incomplete or stale data rather than clinical judgment, the integrity of the underlying record becomes the central problem. This post maps what concurrent review data actually requires across provenance, recency, and concordance.

A 2023 AMA survey found that 34% of prior authorization denials were later overturned on appeal. The number for concurrent review denials is harder to pin down because most payers do not publish it separately, but internal audits at large health plans consistently show that 25% to 30% of concurrent review disputes involve missing or outdated clinical information rather than genuine medical necessity disagreements. The denial was not wrong because the reviewer made a bad call. The denial was wrong because the data underneath the decision was incomplete.

Utilization management data integrity is not a secondary concern. It is the structural requirement that determines whether a concurrent review produces an accurate, defensible, and clinically appropriate result.

What utilization management actually is

Utilization management (UM) is the set of processes health plans and provider organizations use to evaluate whether a specific healthcare service is medically necessary, appropriate for the clinical setting, and delivered efficiently. It is not a single activity. It is a framework with three basic components: prospective review (before a service), concurrent review (during a service), and retrospective review (after a service).

The three components share a common dependency: every review decision is only as good as the clinical and administrative data available at the time the reviewer makes it.

What are the three types of utilization reviews?

The three types of utilization reviews are prospective, concurrent, and retrospective.

Prospective review happens before care begins. Prior authorization is the most familiar example. A clinician submits a request, the payer evaluates it against clinical criteria, and the decision determines whether the service will be covered.

Concurrent review happens while the patient is actively receiving care, typically during an inpatient stay. A reviewer evaluates whether continued hospitalization or a specific level of care remains medically necessary. This review repeats at defined intervals, sometimes daily for ICU patients.

Retrospective review happens after care is complete. The payer examines claims and medical records to determine whether the services delivered were appropriate and whether billing codes match the documented care.

Each type has distinct data requirements, but concurrent review has the most demanding data integrity needs because the patient's condition is actively changing. A prospective review can work with a static snapshot. A concurrent review cannot.

What are the CMS guidelines for utilization review?

CMS requires that Medicare Advantage plans, Medicaid managed care organizations, and QHP issuers on the exchanges maintain utilization review programs that meet specific standards. The key requirements include:

  • Clinical decisions must be made by qualified medical professionals using evidence-based criteria (42 CFR §438.210 for Medicaid, §422.566 for MA).
  • Adverse determinations must include the clinical rationale and the specific criteria used.
  • Plans must provide timely decisions: 72 hours for urgent concurrent review requests, 14 calendar days for standard pre-service requests, and 30 days for post-service reviews under MA rules.
  • CMS requires that UM decisions not be based solely on automated systems without physician review for adverse determinations.
  • The 2024 CMS Interoperability and Prior Authorization Final Rule (CMS-0057-F) mandates that payers implement Prior Authorization APIs using FHIR standards by January 2027, with specific data exchange requirements.
  • The regulatory framework assumes that the data feeding into these decisions is accurate and current. CMS does not have a specific "data integrity score" requirement for UM programs, but the requirement for evidence-based decisions implicitly demands trustworthy data. When a reviewer denies continued stay based on yesterday's lab values that have since changed, the decision fails the CMS standard even if the process was technically followed.

    What are the three key steps in the utilization review process?

    The three key steps are: (1) data collection and clinical documentation review, (2) application of evidence-based criteria such as InterQual or Milliman Care Guidelines, and (3) determination and communication of the decision.

    Step one is where data integrity matters most and where it fails most often. If the clinical documentation is incomplete, if the ADT feed is lagging, or if the reviewer is looking at records from a system that has not synced in 12 hours, the criteria application in step two is built on sand.

    Key statistics

  • 34% of prior authorization denials are overturned on appeal, per the 2023 AMA Prior Authorization Physician Survey, suggesting widespread data and criteria application failures across UM.
  • Claims data lags 30 to 90 days behind the point of care, making retrospective UM inherently dependent on stale information for post-discharge reviews.
  • The average acute care hospital stay generates 50 to 100 pages of clinical documentation, and concurrent reviewers typically see a fraction of it at the time of review.
  • CMS requires urgent concurrent review decisions within 72 hours, but ADT feeds at many hospitals update on batch cycles of 4 to 24 hours, creating decision windows built on incomplete data.
  • SuperTruth's DTI Engine scored 105,000 diagnostic records for imaware, reducing standardization time from 3 weeks to 2 hours. The same scoring logic applies to UM data feeds where recency and concordance determine whether a review decision is defensible.
  • Where concurrent review data breaks

    Concurrent review data failure types by frequency
    Concurrent review data failure types by frequency

    Concurrent review is unique among the three UM types because the underlying clinical state changes continuously. A patient admitted for pneumonia may develop sepsis on day two. A surgical patient may have an unexpected complication that changes their level-of-care need between morning rounds and the afternoon UM review.

    The data integrity failures in concurrent review cluster in five areas:

    Recency gaps. The most common failure. A reviewer evaluates continued stay using lab results from 18 hours ago, vital signs from the prior shift, or progress notes that the attending has not yet signed. The clinical picture at the time of decision does not match the clinical picture at the time of documentation.

    Provenance ambiguity. Clinical data arrives from multiple sources: the hospital's EHR, consulting physician notes, lab interfaces, imaging reports, pharmacy systems. When these sources disagree or when a reviewer cannot determine which source is authoritative, the decision rests on uncertain ground.

    Concordance failures. The diagnosis on the admission order does not match the working diagnosis in the progress note. The ICD-10 code assigned by the coder does not reflect the clinical narrative. The level of care documented by nursing does not align with what the physician ordered. These mismatches create ambiguity that concurrent reviewers must resolve in real time, often without the tools to do so.

    Completeness gaps. A concurrent reviewer needs the full clinical picture: vitals, labs, imaging, medications, nursing assessments, physician notes, and therapy evaluations. In practice, reviewers often work from partial records. A 2022 AHIP study noted that clinician burden from UM documentation requirements is a top complaint, but the flip side is that incomplete submissions force reviewers into guesswork.

    Consent and access constraints. For patients with behavioral health or substance use disorder records, 42 CFR Part 2 restrictions may limit what a concurrent reviewer can see. The reviewer makes a medical necessity determination without access to clinically relevant information that exists but is legally segmented.

    Why UM data trust is different from general clinical data quality

    Clinical data quality efforts typically focus on completeness, accuracy, and standardization of the record itself. UM data trust requires those same qualities plus three additional dimensions that general data quality programs often miss.

    Temporal precision. A lab value is not just accurate or inaccurate. It is accurate as of a specific timestamp, and that timestamp must be close enough to the review decision point to be clinically meaningful. A hemoglobin of 7.2 documented at 6 AM may be clinically irrelevant if two units of packed red blood cells were administered at 8 AM and the review happens at 2 PM. UM data trust requires recency scoring, not just accuracy scoring.

    Decision-context alignment. The data must be structured and tagged in a way that maps to the clinical criteria being applied. InterQual criteria for continued ICU stay reference specific vital sign thresholds, medication requirements, and monitoring needs. If the data exists in a free-text nursing note but is not extractable by the UM platform, it functionally does not exist for the review.

    Auditability. Every concurrent review decision is potentially subject to appeal, external review, and regulatory audit. The data that supported the decision must be preserved exactly as it existed at the time of review. If the underlying record changes after the decision (as clinical records often do with late entries and addenda), the audit trail must show what the reviewer actually saw.

    What concurrent review data specifically requires

    DTI dimension weights applied to concurrent review data
    DTI dimension weights applied to concurrent review data

    Mapping these requirements against the eight dimensions of data trust produces a clear picture of what concurrent review data needs:

    Provenance (25% of DTI weight). Every data element in a concurrent review must trace to a named source system, a specific clinician or device, and a documented chain of custody. When a reviewer sees a blood pressure reading, they need to know whether it came from an automated cuff at the bedside, a manual reading by a nurse, or a value transcribed from a paper record. The provenance determines the reliability.

    Consent (20%). The reviewer must have lawful access to all clinically relevant information. For behavioral health integration, substance use records, and HIV status in certain states, consent constraints create blind spots that must be documented and accounted for in the determination.

    Recency (15%). This is the dimension most likely to cause concurrent review failures. A recency score should reflect not just the timestamp of the data but the clinical velocity of the patient's condition. A stable chronic care patient's data from 24 hours ago may score acceptably. The same staleness for a septic ICU patient would score near zero.

    Quality (10%). Structured data that maps to InterQual or MCG criteria fields scores higher than unstructured text requiring NLP extraction. Validated lab values score higher than manually entered results.

    Concordance (10%). Diagnosis codes, clinical narratives, order sets, and billing codes must align. Discordance between these elements is the single most common trigger for concurrent review denials that are later overturned.

    Validation (10%). Has the data been independently verified? A lab result from a CLIA-certified lab carries more weight than a point-of-care test result entered manually. An imaging report signed by a board-certified radiologist carries more weight than a preliminary read.

    Breadth (5%). Does the reviewer have access to data from all relevant care settings? A patient transferred from another facility needs their transfer records integrated. A patient with outpatient records from a specialist needs those records available.

    Stability (5%). How frequently has this data element changed? A diagnosis that has been revised three times in 48 hours signals clinical uncertainty that the reviewer needs to factor into the determination.

    The AI layer compounds the problem

    Health plans increasingly use AI-assisted tools for concurrent review triage, flagging cases for human review based on algorithmic assessment of medical necessity. UnitedHealth Group's use of the nH Predict algorithm drew national scrutiny in 2023 when reports showed the system had a 90% denial override rate on appeal for post-acute care cases.

    The core issue was not that AI was making decisions. The issue was that the AI was making decisions on data that lacked the integrity required for those decisions. An algorithm trained on historical length-of-stay data cannot account for a specific patient's real-time clinical trajectory if the data feeding the algorithm is 12 to 24 hours stale.

    When AI enters the concurrent review process, every data integrity requirement intensifies. The model cannot exercise clinical judgment to fill gaps. It cannot call the nurse to ask about the patient's current oxygen requirements. It processes what it receives, and if what it receives is incomplete, stale, or discordant, the output is a confident-looking decision built on untrustworthy data.

    What needs to change

    Four structural changes would close the gap between what concurrent review data needs and what UM systems currently receive.

    Real-time data feeds, not batch updates. ADT feeds, lab interfaces, and clinical documentation should flow to UM platforms in near-real-time via FHIR-based APIs. The CMS-0057-F rule pushes in this direction for prior authorization. Concurrent review needs the same standard.

    Trust scoring at the point of ingestion. Every data element entering a UM review should carry a trust score that reflects its provenance, recency, and concordance. A reviewer should see not just the hemoglobin value but a confidence indicator that tells them how fresh, validated, and consistent that value is.

    Audit-grade provenance trails. When a concurrent review decision is made, the system should freeze the data state and preserve it with full chain-of-custody documentation. This protects both the payer and the patient in the event of an appeal.

    Concordance checks before criteria application. UM platforms should run automated concordance checks across diagnosis codes, clinical narratives, and order sets before a reviewer applies InterQual or MCG criteria. Flagging discordances before the review decision prevents the most common category of overturned denials.

    The cost of getting this wrong

    A denied concurrent review that is later overturned costs everyone. The patient experiences delayed or disrupted care. The provider organization spends 45 minutes to 2 hours on average per appeal. The payer processes the appeal, pays for the external review if it reaches that stage, and absorbs the administrative cost. Multiply this across the estimated 35 million prior authorization requests submitted annually (AMA data), plus the concurrent and retrospective reviews that do not appear in prior auth statistics, and the system-wide cost of data-driven UM failures reaches billions.

    The fix is not to eliminate utilization management. It is to ensure that the data feeding UM decisions meets the integrity threshold the decisions require.

    The DTI Engine scores every record 0 to 100 across eight dimensions before your AI model or your human reviewer sees it. If your team is building or evaluating UM platforms, deploying AI-assisted concurrent review, or facing appeal rates that suggest underlying data quality problems, talk to the SuperTruth commercial team. Schedule a conversation or call (215) 918-4140.

    Further reading:

  • DTI™ Engine
  • Health plans solution
  • Prior authorization automation data quality: what AI approvals require
  • Claims data lags 30 to 90 days. Here is what it costs AI models
  • UnitedHealth Group AI denial rate controversy: what data trust had to do with it
  • Why recency is the most underrated dimension in health AI data scoring
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0 to 100. Travels with every record permanently.

    See the DTI Engine
    Share