HL7 message quality scoring: what trust looks like for legacy EHR data
Photo by Quincy Follweiler on Unsplash

HL7 message quality scoring: what trust looks like for legacy EHR data

By Jason Alan Snyder·July 12, 2026

More than 90% of hospital interfaces still run on HL7 Version 2, a standard first published in 1987. Every one of those messages carries structural ambiguity that degrades downstream analytics, AI training, and clinical decision support. Scoring HL7 message quality is how trust gets attached to legacy EHR data before it causes harm.

More than 90% of live hospital interfaces still transmit data using HL7 Version 2. That is not a legacy footnote. It is the operational reality of health data exchange in 2025.

HL7 v2 was first published in 1987. Its pipe-delimited message format predates the modern internet, HIPAA, and every EHR system currently in production. Yet it remains the dominant protocol for ADT alerts, lab results, pharmacy orders, and clinical observations flowing between systems. The problem is not that HL7 v2 exists. The problem is that nobody scores its output before it feeds an AI model, a quality measure, or a clinical decision.

That is what HL7 message quality scoring fixes. Not the protocol itself, but the trust deficit it creates.

What does an HL7 message look like?

An HL7 v2 message is a structured block of ASCII text organized into segments, fields, and components separated by pipe characters (|), carets (^), and tildes (~). A typical ADT (Admit, Discharge, Transfer) message starts with an MSH segment (message header), followed by segments like PID (patient identification), PV1 (patient visit), and OBX (observation results).

Here is a simplified example:

` MSH|^~\&|EPIC|HOSP1|LAB|HOSP1|20250115120000||ADT^A01|MSG00001|P|2.5.1 PID|1||MRN12345^^^HOSP1||DOE^JOHN||19550312|M PV1|1|I|ICU^101^A|E|||1234^SMITH^JANE|||MED||||||||V123456 `

Each segment contains positional fields. PID-5 holds the patient name. PID-7 holds date of birth. PV1-3 holds the patient location. The structure is technically standardized, but the standard permits enormous variation. Fields can be left empty. Coded values differ between institutions. Date formats vary. And the standard itself has evolved through versions 2.1 through 2.9, with hospitals often running different versions on different interfaces simultaneously.

This structural flexibility is precisely what makes HL7 v2 so durable and so dangerous when consumed without validation.

Why legacy EHR data trust requires message-level scoring

HL7 v2 messages are not self-describing. They carry no metadata about their own completeness, accuracy, or provenance. A message with an empty PID-8 (sex) field looks identical, structurally, to one where that field was intentionally omitted versus one where the source system failed to map it.

This creates a compounding trust problem. When thousands of these messages feed a data warehouse, the warehouse inherits every ambiguity. When that warehouse feeds an AI model, the model trains on gaps it cannot distinguish from valid nulls.

Legacy EHR data trust is not about whether the EHR system works. It is about whether the data leaving that system, through HL7 v2 interfaces, arrives at its destination with enough integrity to support the decisions being made from it.

Scoring HL7 message quality means evaluating each message against measurable dimensions: Is the patient identifier present and structurally valid? Is the timestamp formatted correctly and within a plausible range? Are coded fields using the expected vocabulary? Is the sending facility identified? Are required segments present?

Without this scoring, there is no way to answer a basic question: can we trust this data?

Key statistics

  • HL7 Version 2 handles an estimated 90%+ of all clinical data exchange in U.S. hospitals, according to HL7 International.
  • The HL7 v2 standard permits over 80 message types and 200+ trigger events, creating wide variation in what constitutes a "complete" message.
  • SuperTruth's DTI Engine scored 105,000 diagnostic records from imaware, reducing standardization time from 3 weeks to 2 hours, a 95% reduction.
  • The DTI framework evaluates data across 8 dimensions with Provenance weighted at 25%, the single highest-weighted factor for legacy data trust.
  • Research published in BMC Public Health comparing HL7 and legacy syndromic surveillance formats found no difference in completeness for age, chief complaint, and gender, but significant variation in structured coding consistency.
  • The anatomy of HL7 message quality failure

    HL7 message quality failure categories by frequency in production interfaces
    HL7 message quality failure categories by frequency in production interfaces

    HL7 v2 quality problems cluster into five categories that any scoring system must address.

    Missing required fields. A PID segment without a medical record number (PID-3) makes patient matching impossible. A PV1 segment without a visit number (PV1-19) breaks encounter linking. These are not edge cases. They are routine in production interfaces, especially those built years ago and never re-validated.

    Inconsistent coding. HL7 v2 allows locally defined code tables. One hospital maps "M" for male in PID-8. Another uses "1". Another uses "Male". Without a scoring layer that checks coded values against expected vocabularies, downstream analytics silently fragment.

    Timestamp drift. MSH-7 carries the message timestamp. OBR-7 carries the observation date. When these diverge by hours or days, it may reflect legitimate asynchronous processing or it may reflect a clock synchronization failure at the sending system. Scoring flags the discrepancy. Ignoring it propagates it.

    Segment ordering violations. The HL7 v2 standard specifies segment ordering rules. OBX segments should follow their parent OBR. NK1 (next of kin) segments follow PID. Production interfaces sometimes violate these rules, and many receiving parsers silently accept the violations, creating records that look valid but are structurally incorrect.

    Truncation and encoding errors. HL7 v2 fields have maximum lengths that vary by version. A patient name exceeding 48 characters in v2.3 gets truncated by some interfaces and passed intact by others. Special characters in free-text fields (OBX-5 for observation values) cause parsing failures when escape sequences are not properly applied.

    Each of these categories maps directly to a dimension in the Data Trust Index.

    How the DTI scores HL7 message quality

    DTI dimension weights applied to HL7 message scoring
    DTI dimension weights applied to HL7 message scoring

    The DTI Engine scores every health data record from 0 to 100 across eight dimensions. For HL7 v2 messages, the mapping is direct.

    Provenance (25% weight): Does the MSH segment identify the sending application, sending facility, and message control ID? Can we trace this message to a specific source system at a specific time? For legacy EHR data, provenance is the hardest dimension to satisfy because many older interfaces strip or overwrite origin metadata during translation.

    Consent (20% weight): Does the data carry consent indicators, or does the receiving system have a consent governance layer (like ConsentOS) that maps consent status to the message's patient identifier? HL7 v2 has no native consent segment in most message types.

    Recency (15% weight): Is the message timestamp current relative to the use case? A lab result from 2019 fed into a 2025 predictive model without recency scoring creates temporal drift, a well-documented source of AI model degradation.

    Quality (10% weight): Are required fields populated? Are coded values valid? Are data types correct? This is the dimension most people think of when they hear "HL7 message quality," but it represents only 10% of overall trust.

    Concordance (10% weight): Does this message agree with other records for the same patient? If PID-7 in an ADT message says the patient was born in 1955 but a lab message for the same MRN says 1965, concordance scoring catches it.

    Validation (10% weight): Has this data been cross-referenced against an external source? For HL7 messages carrying lab results, validation might mean confirming the performing lab's CLIA number. For provider data, it means primary source verification.

    Breadth (5% weight): How many data elements does this message carry relative to what the message type could carry? An ADT^A01 with only MSH, PID, and PV1 segments is structurally valid but informationally thin.

    Stability (5% weight): Does this interface produce consistent output over time? An interface that intermittently drops NK1 segments has a stability problem that individual message scoring alone will not catch.

    Is EHR data considered secondary data?

    Yes. EHR data is classified as secondary data when used for purposes beyond the original clinical encounter. The primary purpose of an EHR record is to support direct patient care. When that same record is extracted, aggregated, and used for population health analytics, AI model training, quality reporting, or research, it becomes secondary use data.

    This distinction matters for HL7 message quality because secondary use imposes requirements the original data was never designed to meet. A clinician entering a progress note does not think about whether their documentation will train a sepsis prediction model three years later. An HL7 interface built to notify the lab of an admission does not anticipate that the ADT message will feed a social determinants algorithm.

    Secondary data use amplifies every quality gap. A missing zip code in an ADT message is clinically irrelevant for a bed assignment but destroys a geographic disparity analysis. HL7 data integrity scoring bridges this gap by making quality visible before secondary use begins.

    This is why scoring EHR data before any AI model trains on it is not optional. It is the prerequisite.

    What EMR system does Legacy Health use?

    Legacy Health, a six-hospital system based in Portland, Oregon, uses Epic as its primary electronic medical record system. This is relevant to the HL7 message quality discussion because Epic, like all major EHR vendors, generates HL7 v2 messages for outbound interfaces. The quality of those messages depends not just on Epic's configuration but on the interface engine (often Rhapsody, Mirth Connect, or Cloverleaf) and the local build decisions made by the health system's integration team.

    No EHR system produces uniformly high-quality HL7 output. The quality varies by interface, by message type, by department, and by how recently the interface was validated. This is why message-level scoring matters more than vendor-level assumptions.

    What is the latest standard developed under HL7?

    HL7 FHIR (Fast Healthcare Interoperability Resources) is the latest standard from HL7 International. FHIR uses RESTful APIs and JSON or XML resources instead of pipe-delimited messages. It is designed for web-based data exchange and is the foundation of the ONC's interoperability rules, CMS's Patient Access API requirements, and TEFCA.

    FHIR R4, published in 2019, is the current normative release. FHIR R5, published in 2023, adds resources for genomic data, subscription-based notifications, and refined consent models. We covered the trust architecture implications of this version change in detail in our analysis of FHIR R4 vs FHIR R5.

    But here is the critical point: FHIR does not replace HL7 v2 in most production environments. It supplements it. Hospitals run FHIR APIs for patient-facing apps and regulatory compliance while simultaneously running dozens of HL7 v2 interfaces for internal clinical workflows. The legacy data problem does not go away because a new standard exists. It persists until someone scores the old data.

    The cost of unscored HL7 data

    Bad health data costs the U.S. healthcare system an estimated $3.5 trillion annually when you account for duplicated tests, misrouted claims, failed care coordination, and AI model errors. A meaningful portion of that cost traces back to HL7 v2 messages that were accepted, stored, and acted upon without quality validation.

    Consider a concrete example. A hospital sends an ADT^A08 (update patient information) message with an incorrect insurance group number in the IN1 segment. The receiving system updates the patient record. The claim is submitted with the wrong payer information. The claim is denied. A billing specialist spends 45 minutes resolving it. Multiply by thousands of messages per day.

    Now consider the AI use case. A predictive model for readmission risk trains on two years of ADT messages. 12% of those messages have missing or incorrect discharge disposition codes in PV1-36. The model learns that missing disposition is a feature, not a gap. It begins predicting readmission risk based on data absence rather than clinical signal. Nobody catches this because nobody scored the training data.

    This is the problem that unscored health data creates in production.

    Building an HL7 data integrity scoring pipeline

    A practical HL7 message quality scoring system operates at three levels.

    Message-level scoring evaluates individual messages as they arrive. Each message receives a composite score based on field completeness, coding accuracy, structural validity, and timestamp plausibility. Messages below a configurable threshold are flagged for review or quarantined before they enter the data warehouse.

    Interface-level scoring aggregates message scores over time for each interface. An interface that produces consistently low-scoring messages indicates a configuration problem at the source. This is the stability dimension of the DTI applied to HL7 feeds.

    Corpus-level scoring evaluates the entire body of HL7-sourced data for a patient, an encounter, or a population. This is where concordance scoring operates. A single high-quality ADT message means little if it contradicts the demographics in a lab message from the same day.

    The DTI Engine performs all three levels. It scores incoming data at the point of ingestion, before it reaches a model, a report, or a clinical workflow. For organizations dealing with legacy HL7 data, this means retroactively scoring historical data to establish a trust baseline and prospectively scoring new messages to maintain it.

    Why HL7 v2 is not going away

    The industry has been predicting the death of HL7 v2 for over a decade. It has not happened. Here is why.

    HL7 v2 interfaces represent millions of dollars of integration investment at every major health system. Replacing them requires re-engineering workflows, re-testing downstream systems, and re-training staff. The ROI of that replacement is hard to justify when the existing interfaces "work" in the sense that messages flow.

    FHIR adoption is growing, but primarily for net-new use cases: patient apps, payer interoperability, research APIs. The internal clinical data backbone of most hospitals will run HL7 v2 for years. The fragmented health record problem persists precisely because this infrastructure is too entrenched to replace quickly.

    This means HL7 message quality scoring is not a transitional need. It is a permanent requirement for any organization that wants to trust data flowing through legacy interfaces.

    What trust actually looks like for legacy EHR data

    Trust is not a feeling. It is a score.

    For legacy EHR data flowing through HL7 v2 interfaces, trust means every message has a provenance trail. Every field has been evaluated against expected values. Every timestamp has been checked for plausibility. Every coded element has been validated against the appropriate vocabulary. And every score is recorded, auditable, and available to downstream consumers.

    That is what the DTI provides. Not a replacement for HL7 v2. Not a migration path to FHIR. A trust layer that makes the data you already have reliable enough to act on.

    SuperTruth scores incoming EHR data at the point of ingestion, before it reaches a model. If your system is deploying clinical AI and needs to answer an auditor's questions about training data quality, the DTI Engine provides the answer in a quantified, reproducible format. Contact Louis Simeonidis at schedule a conversation or (215) 918-4140.

    Further reading:

  • DTI™ Engine
  • Health systems solution
  • Why EHR data needs a trust score before any AI model trains on it
  • Data quality vs data trust: what is the difference and why it matters for healthcare AI
  • The eight dimensions of health data trust: a practical guide
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    EHR data scored before any AI model sees it.

    DTI integrates with Epic, Oracle Health, and all major EHR systems.

    See our health systems solution
    Share