Cross-sector data trust for housing-health integration
Photo by zero take on Unsplash

Cross-sector data trust for housing-health integration

By Jason Alan Snyder·September 22, 2026

Housing and health data sit in separate systems governed by separate laws, creating a trust gap that no data-sharing agreement alone can close. Cross-sector integration requires scoring every record for provenance, consent, and recency before it enters a model. Without that foundation, social determinants housing intelligence remains unreliable.

Housing instability drives 18% of total healthcare spending in the United States, according to a 2019 Health Affairs analysis of social determinants spending attribution. Yet the data systems that track housing and the data systems that track health were never designed to talk to each other. They run on different standards, different consent frameworks, different timelines, and different definitions of what a "record" even is.

The result: cross-sector data trust fails before the integration even starts.

The two-system problem

Health data lives in EHRs, claims databases, and HIE networks governed by HIPAA. Housing data lives in HUD systems, Homeless Management Information Systems (HMIS), public housing authority databases, and community-based organization intake forms governed by a patchwork of federal, state, and local rules.

HMIS alone operates under HUD Data Standards that define 107 data elements, most of which have no FHIR equivalent. The 2023 HMIS Data Standards Manual requires fields like "prior living situation" and "length of stay in prior living situation" that have no structured representation in any major EHR system.

HIPAA regulates health data. The McKinney-Vento Act governs some housing data for families with children. HMIS data falls under HUD's privacy and security standards, which are separate from and in some cases incompatible with HIPAA's consent framework. A patient who consents to share health data through an HIE has not consented to share their HMIS record. A tenant who gives intake information to a housing authority has not consented to have that record linked to their Medicaid claims.

This is not a technical problem. It is a consent architecture problem.

Why data-sharing agreements are not enough

The top-ranking content on this topic focuses on data-sharing agreements and MOUs as the solution to cross-sector integration. That framing misses the core issue. A data-sharing agreement governs who can access what. It does not verify that the data being shared is accurate, current, or properly consented for the downstream use.

Consider a common scenario: a Medicaid managed care organization wants to identify members experiencing housing instability so it can route them to supportive services. The MCO signs a data-sharing agreement with the local Continuum of Care to receive HMIS data. The HMIS data arrives.

Now what?

The HMIS record uses a client identifier that does not map to the MCO's member ID. The address field in the HMIS record may be a shelter address, a last known address, or blank. The record's timestamp reflects when data was entered, not when the housing event occurred. The 2022 HUD Annual Homeless Assessment Report found that point-in-time counts undercount the homeless population by an estimated 40% because of methodology limitations.

A data-sharing agreement does not fix any of this. What fixes it is scoring every record before it enters a model.

The five trust dimensions that break in cross-sector integration

DTI dimensions most affected by housing-health integration
DTI dimensions most affected by housing-health integration

| | Value (%) | |---|---| | Provenance | 25 | | Consent | 20 | | Recency | 15 | | Concordance | 10 | | Validation | 10 | | Quality | 10 | | Breadth | 5 | | Stability | 5 |

When housing data meets health data, five of the eight DTI dimensions face specific failure modes.

Provenance

Provenance accounts for 25% of the DTI score because knowing where a record came from is the single most important trust signal. In cross-sector integration, provenance breaks in both directions.

Health records may arrive through an HIE with a chain of custody that spans three or four systems. By the time a diagnosis code reaches the housing-focused analytics model, the original clinician's attestation is buried under layers of transformation. Housing records are worse. A caseworker at a community-based organization may enter data into a spreadsheet that gets uploaded to HMIS monthly. The provenance of that record is a person's memory, typed into Excel, batch-loaded into a federal system.

The Office of the National Coordinator's 2024 interoperability roadmap acknowledges that provenance metadata is "inconsistently captured" even within health-only data exchange. Across sectors, it is nearly absent.

Consent

Consent accounts for 20% of the DTI score. Cross-sector integration creates what we call the consent layering problem: the original consent given for data collection does not cover the downstream analytical use.

A person experiencing homelessness who provides information at a shelter intake is consenting to receive services. They are not consenting to have their shelter stay linked to their emergency department visits in a predictive model that flags them for care management outreach. The National Health Care for the Homeless Council's 2023 policy brief identified consent as the single largest barrier to housing-health data integration, with 67% of surveyed programs reporting that consent processes were "inadequate for cross-sector use."

Recency

Recency accounts for 15% of the DTI score. Housing status changes faster than almost any other social determinant. A person can move from housed to unhoused to sheltered to housed again within 90 days. Claims data lags 30 to 90 days. HMIS data is often updated quarterly. By the time a housing record and a health record are linked, the housing situation may have already changed.

The 2023 Annual Homeless Assessment Report showed that the median length of a shelter stay was 22 days. A record that reflects a shelter stay from 60 days ago may no longer represent reality.

Concordance

Concordance measures whether the same fact is recorded the same way across systems. Housing-health integration fails concordance routinely. A health system may record a Z59.0 ICD-10 code for homelessness. The HMIS system records the same person with a "prior living situation" code of 1 (emergency shelter) or 16 (place not meant for habitation). These are related but not identical assertions. A Medicaid MCO may flag a member as "housing insecure" based on an address change pattern in claims data. That flag may or may not align with what the HMIS record shows.

The Gravity Project's SDOH Clinical Care HL7 Implementation Guide attempted to create FHIR-based mappings between housing screening instruments and clinical records. As of 2024, adoption of these mappings remains limited, with the Gravity Project's own tracking showing fewer than 30 EHR implementations using the full housing domain value set.

Validation

Validation asks whether a record has been independently confirmed. In health data, validation might come from a second clinician's note, a lab result, or a claims adjudication. In housing data, validation is rare. A self-reported housing status at a screening is typically accepted at face value. A caseworker's observation is documented but not cross-referenced. The AHC-HRSN screening tool, which CMS developed for the Accountable Health Communities model, includes two housing questions, but neither is validated against actual housing records.

Z-codes capture the problem in miniature

Z-code capture vs estimated prevalence of housing instability
Z-code capture vs estimated prevalence of housing instability

| | Value (%) | |---|---| | Z-code capture rate | 1.59 | | Estimated actual prevalence | 24 |

ICD-10 Z-codes are the primary mechanism through which housing-related social determinants enter the clinical data stream. Z59 codes cover homelessness, inadequate housing, and housing instability. The capture rate is abysmal.

A 2023 JAMA Network Open study found that Z-code documentation for social determinants occurred in only 1.59% of Medicare fee-for-service claims, despite screening rates that suggest prevalence is 10 to 20 times higher. For housing-specific Z-codes, the capture rate was even lower.

This means that any model trained on claims data to identify housing instability is working with a denominator problem so severe that the model's outputs cannot be trusted. The data is not wrong. It is absent. And absence, in a predictive model, looks like a negative finding. The model concludes that the patient does not have housing instability when in reality the clinician never asked or never coded the answer.

We wrote about this gap in depth in our analysis of SDOH screening program data quality and Z-code capture rates.

What geocoding gets wrong about housing intelligence

Social determinants housing intelligence increasingly relies on address-level geocoding to infer housing quality, neighborhood risk, and environmental exposures. The assumption is that a patient's address can be linked to census tract data, which can then be linked to area deprivation indices, environmental hazard databases, and housing quality indicators.

The assumption breaks for the populations that need it most.

People experiencing homelessness often have no stable address. People in transitional housing may use a shelter address or a PO box. People doubling up with family members may use an address that does not reflect their actual living conditions. The AHRQ Social Determinants of Health Database uses census tract-level measures that assume residential stability. For the 1.5 million people who used a shelter in 2023 according to HUD's AHAR data, census tract linkage is meaningless.

We covered the broader geocoding accuracy problem in our piece on address-level SDOH data trust and the census tract mismatch problem.

What cross-sector data trust actually requires

The path forward is not more data-sharing agreements. It is not more interoperability standards, though those help. It is a trust layer that sits between the data sources and the models that consume them.

That trust layer must do four things.

First, score every record for provenance before it enters the integration pipeline. A housing record from a caseworker's direct observation is not the same as a housing record inferred from an address change in claims data. Both may be useful. They are not equally trustworthy.

Second, enforce consent at the record level, not the dataset level. Cross-sector integration requires knowing that this specific person consented to this specific use of this specific record. Dataset-level consent agreements do not provide that granularity.

Third, timestamp the assertion, not just the entry. When a caseworker enters a housing status observation into HMIS three weeks after the observation, the system needs to distinguish between the observation date and the entry date. Models that treat entry timestamps as event timestamps will draw false conclusions about housing trajectories.

Fourth, score concordance across systems before merging records. If the health system says the patient is housed (no Z-code) and the HMIS says the patient entered a shelter last month, the integration layer needs to surface that discordance rather than silently picking one record over the other.

The CBO data problem

Community-based organizations are the frontline of housing-health integration. They conduct screenings, provide referrals, and track outcomes. Their data infrastructure is, in most cases, a spreadsheet.

The 2023 NACHC data capacity survey found that 41% of community health centers lacked the IT infrastructure to share SDOH data electronically. For smaller CBOs that are not federally qualified health centers, the number is almost certainly higher, though comprehensive survey data does not exist.

This means that cross-sector data trust requires meeting CBOs where they are. Asking a three-person housing nonprofit to implement FHIR-based data exchange is not realistic. What is realistic is scoring the data they produce, in whatever format they produce it, and attaching that score to the record as it moves into the health system's analytics pipeline.

We explored this problem specifically in our analysis of CBO data trust for Medicaid value-based programs.

The Medicaid integration mandate

CMS is pushing this integration forward whether the infrastructure is ready or not. The 2024 Medicaid Access Rule requires state Medicaid programs to develop quality strategies that address social determinants, including housing. Several states, including California, Massachusetts, and North Carolina, have already launched Medicaid waiver programs that require housing-health data integration.

California's CalAIM Community Supports program requires managed care plans to track housing-related community supports and report outcomes to the state. North Carolina's Healthy Opportunities Pilots pay community organizations to deliver housing services and require data reporting that links housing interventions to Medicaid utilization outcomes.

These programs generate data. They do not, by default, generate trustworthy data. The reporting requirements specify what must be reported. They do not specify the provenance, consent, or recency standards the data must meet. That gap is where models fail.

Scoring housing-health data before the model sees it

Every record that enters a housing-health integration pipeline carries risk. A housing record with unknown provenance, expired consent, and a 90-day-old timestamp is not a record that should drive a care management decision. But without a scoring mechanism, that record looks the same as a verified, recently observed, properly consented record.

The DTI framework scores every record across eight dimensions before any model acts on it. For housing-health integration specifically, the three dimensions that matter most are provenance (where did this record come from and who attested to it), consent (did this person agree to this specific cross-sector use), and recency (does this record reflect the person's current situation).

A record that scores below the trust floor does not get suppressed. It gets flagged. The model still sees it, but it sees the score too. A housing status observation from a caseworker last week with documented consent for cross-sector sharing scores differently than an inferred housing status from a 60-day-old address change in claims data. Both contribute to the picture. Only one should drive a decision.

The DTI Engine scores every record 0 to 100 across eight dimensions before your AI model sees it. If your team is building housing-health integration for a Medicaid managed care program, a health system's population health platform, or a community-based organization network, the question is not whether you can get the data. It is whether the data is trustworthy enough to act on. Schedule a conversation or call (215) 918-4140.

Further reading:

  • DTI™ Engine
  • Health systems solution
  • Housing instability data quality: eviction records and health data integration trust
  • Address-level SDOH data trust: geocoding accuracy and the census tract mismatch problem
  • Community health organizations and SDOH data quality: the trust gap
  • Sources

  • Health Affairs, "Spending On Social Services Per Capita And Health Care Spending," 2019, https://www.healthaffairs.org/doi/10.1377/hlthaff.2018.05286
  • HUD Exchange, "HMIS Data Dictionary," 2024, https://www.hudexchange.info/resource/3824/hmis-data-dictionary/
  • HUD Exchange, "HMIS Data Standards Manual," 2024, https://files.hudexchange.info/resources/documents/HMIS-Data-Standards-Manual-2024.pdf
  • HUD Exchange, "HMIS Privacy and Security Standards," https://www.hudexchange.info/resource/1391/hmis-privacy-security-standards/
  • National Health Care for the Homeless Council, "Data Sharing Issue Brief," 2023, https://nhchc.org/wp-content/uploads/2023/06/NHCHC-Data-Sharing-Issue-Brief-2023.pdf
  • HUD, "2022 Annual Homeless Assessment Report Part 1," 2022, https://www.huduser.gov/portal/datasets/ahar/2022-ahar-part-1-pit-estimates-of-homelessness-in-the-us.html
  • HUD, "2023 Annual Homeless Assessment Report Part 1," 2023, https://www.huduser.gov/portal/datasets/ahar/2023-ahar-part-1-pit-estimates-of-homelessness-in-the-us.html
  • ONC, "Interoperability Roadmap," 2024, https://www.healthit.gov/topic/interoperability/interoperability-roadmap
  • HL7 Gravity Project, "SDOH Clinical Care Implementation Guide," 2024, https://build.fhir.org/ig/HL7/fhir-sdoh-clinicalcare/
  • The Gravity Project, https://thegravityproject.net/
  • JAMA Network Open, "Social Determinants of Health Z-Code Documentation," 2023, https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2800742
  • CMS, "Accountable Health Communities Model," https://innovation.cms.gov/innovation-models/ahcm
  • AHRQ, "Social Determinants of Health Database," https://www.ahrq.gov/sdoh/data-analytics/sdoh-data.html
  • NACHC, "Data Capacity Survey," 2023, https://www.nachc.org/resource/2023-data-capacity-survey/
  • CMS, "Medicaid and CHIP Managed Care Access, Finance, and Quality Final Rule," 2024, https://www.cms.gov/newsroom/fact-sheets/medicaid-and-childrens-health-insurance-program-chip-managed-care-access-finance-and-quality-final
  • California DHCS, "CalAIM Community Supports," https://www.dhcs.ca.gov/CalAIM/Pages/Community-Supports.aspx
  • North Carolina DHHS, "Healthy Opportunities Pilots," https://www.ncdhhs.gov/about/department-initiatives/healthy-opportunities/healthy-opportunities-pilots
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0 to 100. Travels with every record permanently.

    See the DTI Engine
    Share
    Cross-sector data trust for housing-health integration | SuperTruth