Sexual orientation and gender identity data collection trust in healthcare
Photo by Etienne Boulanger on Unsplash
insight

Sexual orientation and gender identity data collection trust in healthcare

By Jason Alan Snyder·August 18, 2026

Only 19% of federally qualified health centers consistently collect sexual orientation and gender identity data from patients. The gap is not a technology problem. It is a trust problem, and it cascades into every AI model trained on incomplete demographic fields.

Only 19% of federally qualified health centers consistently collect sexual orientation and gender identity (SOGI) data from patients, according to a 2023 analysis published in the American Journal of Preventive Medicine. The other 81% either collect it inconsistently, leave the fields blank, or do not ask at all.

This is not a technology limitation. EHR systems have supported SOGI fields since Meaningful Use Stage 3 required their availability in 2016. The fields exist. The problem is that patients do not trust what happens after they answer, and clinicians do not trust the data after it is recorded. That dual trust failure makes SOGI data one of the lowest-quality demographic categories in health records, and it poisons every downstream AI model that depends on complete population data.

Why SOGI data collection rates remain low despite federal mandates

The Office of the National Coordinator for Health IT (ONC) finalized requirements in 2015 for certified EHR technology to include structured SOGI data fields. CMS followed with quality reporting measures that reference SOGI completeness. By regulation, the infrastructure has been in place for nearly a decade.

Yet collection rates tell a different story. A 2022 study in JAMA Network Open found that among 1.6 million patient encounters at a large academic health system, sexual orientation data was recorded for only 25.1% of patients. Gender identity data fared slightly better at 35.6%, but over a third of those entries were marked "unknown" or "choose not to disclose."

The gap exists because mandating a field does not create the conditions for a patient to answer it honestly. And an unanswered or inaccurately answered SOGI field is worse than a missing one, because a missing field signals incompleteness while a wrong answer signals false completeness.

The trust deficit that drives non-disclosure

SOGI data collection barriers ranked by patient-reported impact
SOGI data collection barriers ranked by patient-reported impact

Patients withhold SOGI data for specific, rational reasons. A 2021 survey by the Williams Institute at UCLA School of Law found that 56% of LGBTQ+ respondents feared their data would be used to discriminate against them in insurance coverage decisions. Another 43% worried about information being shared with third parties without consent.

These fears are not abstract. Recent MedPage Today coverage on perceived discrimination in chronic disease management confirms the clinical stakes. Among patients with type 2 diabetes and hypertension, higher perceived discrimination in healthcare settings is associated with worse outcomes. LGBTQ+ patients who distrust data collection are not being paranoid. They are responding to a pattern where information volunteered in clinical settings has historically been used against them, from pre-Obergefell insurance exclusions to conversion therapy referrals documented in medical charts.

The trust problem compounds when patients see SOGI data requested on intake forms without any explanation of why it is being collected or how it will be used. A 2022 National Academies of Sciences report found that collection rates increased by 20 to 30 percentage points when clinical staff explained the purpose of SOGI questions and assured patients that responses were voluntary and confidential.

What happens to AI models when SOGI data is missing or wrong

Missing SOGI data does not create a neutral gap. It creates a systematic bias. When AI models train on population health data where 75% of sexual orientation fields are empty, those models learn to treat the visible 25% as representative. They are not.

The 25% who do disclose SOGI data tend to skew toward patients at large academic medical centers, patients under 40, and patients in states with anti-discrimination protections. Rural LGBTQ+ patients, older adults, and people in states without employment protections are disproportionately absent from the data.

This means a readmission prediction model trained on this data will systematically undercount LGBTQ+ risk factors. A care gap identification algorithm will miss screening recommendations specific to LGBTQ+ populations. A social determinant of health model will lack the demographic layer necessary to identify minority stress as a contributing factor to chronic disease.

As we documented in our work on race and ethnicity data quality, the difference between self-reported and inferred demographic data is not a minor quality issue. It is a structural one. For SOGI data, this structural problem is even more severe because inference is essentially impossible. You cannot impute someone's sexual orientation from claims data, and any attempt to do so introduces ethical violations alongside statistical ones.

Key statistics

The numbers paint a clear picture of the SOGI data trust crisis in healthcare:

  • 19% of federally qualified health centers consistently collect SOGI data (American Journal of Preventive Medicine, 2023)
  • 25.1% of patient encounters at a large academic health system included recorded sexual orientation data (JAMA Network Open, 2022)
  • 56% of LGBTQ+ respondents feared SOGI data would be used for insurance discrimination (Williams Institute, 2021)
  • 20-30 percentage point increase in collection rates when staff explained the purpose of SOGI questions (National Academies, 2022)
  • 95% time reduction achieved when SuperTruth standardized 105,000 diagnostic records for imaware, demonstrating the DTI framework's ability to score and normalize sensitive demographic data at scale
  • How SOGI data quality degrades across EHR systems

    Even when SOGI data is collected, quality degrades through multiple mechanisms.

    First, terminology inconsistency. One EHR system may offer "gay," "lesbian," "bisexual," "queer," and "other" as sexual orientation options. Another may offer only "heterosexual," "homosexual," and "bisexual." A third may use free-text entry. When these records are aggregated for population health analytics or AI training, the same patient identity gets coded differently depending on which system collected it.

    Second, temporal decay. SOGI data is typically collected once during patient registration and never updated. A patient who identified as heterosexual at age 19 and bisexual at age 28 will carry stale data in every system that does not prompt re-collection. As we have written about in data retention policy and trust scores, aging data loses integrity. SOGI data is particularly vulnerable because identity can evolve over a lifetime.

    Third, consent ambiguity. Patients may disclose SOGI information to a primary care provider with an implicit understanding that it stays in that clinical context. When the same data flows to a health information exchange, a payer, or a research consortium, the patient's original consent intent has been exceeded. This is the consent layering problem applied to one of the most sensitive data categories in healthcare.

    The political and regulatory climate makes trust harder

    SOGI data collection does not exist in a vacuum. Policy positions on LGBTQ+ healthcare, gender-affirming care restrictions, and anti-discrimination enforcement shift with election cycles and judicial decisions. MedPage Today's coverage of where political candidates stand on healthcare issues underscores how volatile this policy environment remains.

    For patients, this volatility increases the perceived risk of disclosure. A SOGI data field that felt safe to complete in 2021 may feel dangerous in 2025 if a patient's state has since restricted gender-affirming care and expanded the circumstances under which medical records can be subpoenaed.

    For health systems, the volatility creates a compliance paradox. Federal quality measures increasingly require SOGI data completeness, while state laws in some jurisdictions restrict or complicate the collection and use of gender identity data. Systems operating across state lines face contradictory obligations.

    The result is a data trust environment where both patients and institutions have rational reasons to avoid engagement with SOGI fields. And the data quality suffers accordingly.

    What trust scoring reveals about SOGI data readiness

    DTI dimension weights applied to SOGI data scoring
    DTI dimension weights applied to SOGI data scoring

    When you apply a structured trust framework to SOGI data, the deficiencies become measurable rather than anecdotal.

    The Data Trust Index (DTI) scores health data records 0 to 100 across eight dimensions: Provenance (25%), Consent (20%), Recency (15%), Quality (10%), Concordance (10%), Validation (10%), Breadth (5%), and Stability (5%). Applied to SOGI data, here is what each dimension reveals:

    Provenance (25%): Where did this SOGI data originate? Was it self-reported by the patient on a validated instrument, entered by a registration clerk based on a verbal exchange, or inferred from billing codes? The provenance chain for SOGI data is frequently broken or undocumented.

    Consent (20%): Did the patient consent to SOGI data being used for population health analytics? For AI model training? For research? The consent layer for SOGI data is almost never granular enough to answer these questions.

    Recency (15%): When was this SOGI data last confirmed or updated? If the answer is "at initial registration seven years ago," the recency score drops substantially.

    Quality (10%): Is the data structured using standardized vocabularies (SNOMED CT, HL7 gender identity codes)? Or is it free text that requires NLP to interpret?

    Concordance (10%): Does the SOGI data in the EHR match what appears in connected systems, insurance records, or patient portal entries? Discordance signals either data entry error or identity evolution that has not been captured.

    Validation (10%): Has the data been cross-referenced or confirmed through any secondary mechanism?

    Breadth (5%): Does the record capture both sexual orientation and gender identity, or only one?

    Stability (5%): Has the data remained consistent across updates, or does it fluctuate in ways that suggest data quality issues rather than genuine identity changes?

    Most SOGI data records in current EHR systems would score below 30 on the DTI scale. That places them in the Bronze tier at best, far below the threshold for any responsible AI training or population health analytics.

    What actually improves SOGI data trust

    Improving SOGI data quality is not primarily a technical challenge. It is a trust architecture challenge that requires coordinated action across five domains.

    1. Purpose transparency at the point of collection. Patients need to know why SOGI data is being collected, who will see it, and what it will not be used for. The 20 to 30 percentage point improvement documented by the National Academies came from this single intervention.

    2. Granular consent governance. SOGI data should carry consent metadata that specifies permitted uses. A patient may consent to their provider seeing sexual orientation data for clinical decision-making while declining its inclusion in research datasets or AI training pools. This is exactly the kind of problem that dynamic consent architectures are designed to solve.

    3. Standardized vocabulary adoption. Health systems need to adopt HL7 and SNOMED CT SOGI value sets consistently, not create custom picklists that fragment data across systems. Without terminology standardization, concordance scores will remain low.

    4. Recency enforcement. SOGI data should be re-confirmed at clinically appropriate intervals, not treated as a static registration field. Annual confirmation during wellness visits is a practical starting point.

    5. Trust scoring before downstream use. No SOGI data should flow into an AI model, a population health dashboard, or a research dataset without a trust score that quantifies its provenance, consent status, and recency. This is not optional data hygiene. It is the minimum standard for responsible use of sensitive demographic information.

    The cost of getting this wrong

    The downstream costs of poor SOGI data quality are specific and measurable.

    Health plans that cannot accurately identify LGBTQ+ members will fail to offer targeted preventive services, including HIV PrEP outreach, cervical cancer screening for transgender men, and mental health resources for populations with elevated suicide risk. These failures increase downstream acute care costs.

    Research institutions that train models on SOGI-incomplete data will produce findings that do not generalize to LGBTQ+ populations, perpetuating the evidence gaps that have historically disadvantaged these communities.

    Health systems that collect SOGI data without adequate trust infrastructure face regulatory exposure. If a patient's SOGI data is breached or used without consent, the legal and reputational consequences are amplified by the sensitivity of the information.

    The pattern mirrors what we documented in disability status data quality: when a demographic category is simultaneously sensitive and undercollected, AI systems trained on that data will produce outputs that are clinically inadequate for the populations most in need.

    Building trust before building models

    The current SERP landscape around SOGI data collection focuses almost entirely on the clinical case for collection and the implementation mechanics of asking the questions. Those are necessary but insufficient. What is missing from the conversation is a rigorous framework for evaluating whether collected SOGI data meets the quality threshold for any given downstream use.

    A SOGI data field that was self-reported, collected with explicit consent for research use, standardized to HL7 vocabulary, and confirmed within the last 12 months is a fundamentally different data asset than one that was entered by a clerk during a rushed registration, carries no consent metadata, uses a custom picklist, and has not been updated in six years. Both appear as populated fields in an EHR. Only one is trustworthy.

    Trust scoring makes this distinction visible and actionable. Without it, health systems, researchers, and AI developers are building on a foundation they cannot verify.

    The DTI Engine scores every health data record 0 to 100 across 8 trust dimensions before your AI model sees it. For organizations working with SOGI data, this means quantifiable provenance, consent verification, and recency enforcement on every record before it enters a training set, a dashboard, or a regulatory submission. If your team is evaluating demographic data quality for compliance, clinical AI, or population health analytics, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.

    Further reading:

  • DTI™ Engine
  • Health systems solution
  • Race and ethnicity data quality: what self-reported vs inferred data means for AI
  • Disability status data quality: the hidden demographic missing from most health AI
  • Mental health data: the most sensitive consent domain in healthcare AI
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0–100. Travels with every record permanently.

    See the DTI Engine
    Share