Health equity data: measuring what we do not see in traditional health systems
Photo by Yang🙋‍♂️🙏❤️ Song on Unsplash
insight

Health equity data: measuring what we do not see in traditional health systems

By Jason Alan Snyder·April 23, 2026

Roughly 80% of health outcomes are driven by factors outside the clinical setting, yet most health systems collect structured data on fewer than half of those factors. Health equity data measurement requires capturing what traditional EHRs were never designed to see: housing instability, food access, transportation barriers, and the trust dynamics that determine whether patients share information at all.

Roughly 80% of health outcomes are driven by factors outside the clinical setting. Yet most health systems collect structured data on fewer than half of those factors. The result is a measurement crisis that makes health equity invisible to the very institutions tasked with improving it.

Traditional EHRs were designed to document encounters, not to capture the conditions that determine whether a patient shows up for that encounter in the first place. Housing instability, food access, transportation barriers, neighborhood safety, and the trust dynamics that shape whether a person shares accurate information with a provider: none of these live neatly in an ICD-10 code.

This is the core problem with health equity data measurement. We are trying to close gaps we cannot see because our data infrastructure was never built to see them.

What are the 4 pillars of health equity?

Most frameworks identify four pillars: access, quality, outcomes, and experience. Access means whether people can physically and financially reach care. Quality means whether the care delivered meets evidence-based standards regardless of who receives it. Outcomes means whether survival rates, complication rates, and disease burden are distributed equitably across populations. Experience means whether patients feel respected, heard, and safe.

Each pillar requires different data. Access data comes from geographic and insurance coverage records. Quality data comes from clinical process measures. Outcome data comes from registries and claims. Experience data comes from surveys, behavioral signals, and community-level trust indicators.

The problem is that most health systems only measure the first two with any consistency. Outcome and experience data, particularly for underserved populations, remain fragmented or missing entirely.

What are the biggest barriers to health equity?

The barriers are structural, not incidental. They include income inequality, residential segregation, differential insurance coverage, implicit bias in clinical decision-making, and the digital divide that limits telehealth access for rural and low-income populations.

But there is a barrier underneath all of these: data trust. Communities that have experienced medical exploitation, from the Tuskegee syphilis study to contemporary algorithmic bias in risk prediction tools, are less likely to share health information. When patients withhold data or avoid care systems entirely, the resulting datasets underrepresent them. AI models trained on those datasets then perpetuate the disparity. As a recent MedPageToday opinion piece argued, financial return on investment should not be the deciding factor in healthcare; mission ROI, including equity outcomes, must carry equal weight.

This is why health disparities data trust is not a soft concept. It is a technical prerequisite. Without it, the data pipeline is contaminated at the source.

What are the 3 C's of healthcare?

The 3 C's are cost, convenience, and continuity. Cost determines whether a patient can afford care. Convenience determines whether logistical barriers like distance, wait times, and scheduling friction prevent them from receiving it. Continuity determines whether their care is coordinated across providers and over time.

For underserved populations, all three C's tend to fail simultaneously. A Medicaid beneficiary in a rural county may face high out-of-pocket costs for transportation, zero convenient specialists within 60 miles, and no continuity because their records are scattered across disconnected systems. MedPageToday's coverage of bringing house calls into the 21st century highlighted how home health services can address convenience and continuity gaps, but only if the data infrastructure supports coordination.

The fragmented health record problem is directly relevant here. When the most valuable patient data lives nowhere, continuity becomes impossible.

How do we measure health equity?

Measuring health equity requires three capabilities most systems lack.

First, stratified data collection. Every clinical and operational metric must be breakable by race, ethnicity, language, disability status, sexual orientation, gender identity, and socioeconomic indicators. The National Academy of Medicine has called for this since 2009. Most hospitals still cannot do it reliably.

Second, SDOH data integration. Social determinants of health data from community-based organizations, public records, and patient-reported sources must be linked to clinical records. This is where tools like DataSpine become essential, providing geographic SDOH scoring at the community level rather than relying on patients to self-report in a clinical encounter.

Third, trust scoring. Raw health equity data is only useful if it is accurate, consented, recent, and validated. A dataset with 40% missing race/ethnicity fields is not just incomplete; it actively distorts any equity analysis built on top of it. This is why health equity AI data must be scored before any model touches it. SuperTruth's Data Trust Index scores every record 0 to 100 across eight dimensions, including provenance, consent, recency, and concordance, so that equity analyses are built on data that has been verified, not assumed.

Key statistics

Health equity data gaps: what systems collect vs. what matters
%22%7D%7D%5D%7D%2C%22legend%22%3A%7B%22display%22%3Atrue%2C%22position%22%3A%22top%22%7D%2C%22plugins%22%3A%7B%22filler%22%3A%7B%22propagate%22%3Afalse%7D%7D%7D%7D) Health equity data gaps: what systems collect vs. what matters

The numbers tell a stark story about the current state of health equity measurement.

  • Only 55% of U.S. hospitals collect race and ethnicity data in a standardized format, according to the American Hospital Association.
  • The CDC estimates that social determinants of health account for approximately 80% of health outcomes, yet fewer than 25% of health systems systematically screen for SDOH.
  • SuperTruth's work with imaware standardized 105,000 diagnostic records, reducing processing time from 3 weeks to 2 hours, a 95% time reduction that freed over 200 hours per month for analysis rather than data wrangling.
  • A 2023 JAMA Network Open study found that algorithmic bias in risk prediction tools underestimated illness severity for Black patients by up to 50%, directly tied to training data that lacked equity stratification.
  • Health systems that implemented standardized equity data collection saw a 30% improvement in identifying at-risk populations within the first year, per Commonwealth Fund research.
  • The missing layer: community-sourced equity data

    Hospital systems will never capture the full picture alone. Community-based organizations hold data on food insecurity, housing conditions, domestic violence, substance use, and immigration-related care avoidance that clinical systems cannot access.

    But CBO data has its own integrity challenges. It is often collected in spreadsheets, stored inconsistently, and shared without formal consent governance. The trust gap between community health organizations and SDOH data quality is real and well documented.

    Bridging this gap requires infrastructure, not just good intentions. CBO data trusts for Medicaid value-based programs offer one model: shared data governance frameworks where community organizations retain control while contributing to population health analytics.

    Why equity data without trust scoring fails

    Data Trust Index: 8 dimensions of health data scoring
    Data Trust Index: 8 dimensions of health data scoring

    Collecting more equity data without scoring it for integrity creates a dangerous illusion of progress. A health system might report that it has "complete" SDOH data on 90% of its patient panel, but if 60% of those records are self-reported, unvalidated, and over two years old, the resulting equity metrics are fiction.

    Health equity AI data must meet the same rigor as clinical trial data. Provenance must be documented. Consent must be explicit and current. Recency must be enforced. Concordance across sources must be checked. This is exactly what the Data Trust Index was built to do.

    Without this layer, health equity measurement becomes a compliance exercise rather than a tool for actual change.

    To understand how trust-scored health equity data can strengthen your population health analytics and AI readiness, contact Louis Simeonidis, SVP Commercial Operations, at louis@supertruth.ai or (215) 918-4140.

    Further reading:

  • DTI Engine
  • Health systems solution
  • Community health organizations and SDOH data quality: the trust gap
  • CBO data trust for Medicaid value-based programs: what community organizations need
  • Rural health data gaps and how synthetic data fills them without compromising trust
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    EHR data scored before any AI model sees it.

    DTI integrates with Epic, Cerner, and all major EHR systems.

    See our health systems solution
    Share