The consent layering problem: when downstream data use exceeds original consent
Photo by A.Rahmat MN on Unsplash
insight

The consent layering problem: when downstream data use exceeds original consent

By Jason Alan Snyder·August 9, 2026

Most health data consent frameworks capture a single moment of agreement, then allow that data to flow through unlimited downstream uses. A 2023 ONC study found that 79% of patients did not know their clinical data had been shared with third parties for secondary purposes. Consent layering, the practice of stacking new data uses on top of original permissions, is the structural failure that erodes health data secondary use trust.

A patient signs a consent form at intake. That form authorizes treatment, billing, and healthcare operations. Six months later, that same patient's de-identified record trains a machine learning model sold to a pharmaceutical company. Twelve months after that, the model's outputs inform a payer's coverage denial algorithm. The patient never agreed to any of this.

This is the consent layering problem. And it is not hypothetical.

What consent layering actually means

Consent layering occurs when data collected under one consent context gets repurposed for secondary, tertiary, or even quaternary uses that the original consent never contemplated. Each new use adds a "layer" of purpose on top of the original permission, and each layer moves further from what the patient understood when they signed.

The original consent form might say "I agree to the use of my health information for treatment and payment purposes." That is layer one. Layer two happens when the health system shares de-identified data with a research consortium. Layer three happens when the consortium licenses data to an AI vendor. Layer four happens when the AI vendor sells model outputs to a payer.

No single step in this chain is obviously illegal. But the cumulative distance between the patient's original understanding and the data's actual use is enormous.

Why HIPAA does not solve this

HIPAA permits the use and disclosure of protected health information for treatment, payment, and healthcare operations without additional patient authorization. It also permits de-identified data to be used for essentially any purpose, as long as it meets Safe Harbor or Expert Determination standards.

This creates a structural gap. Once data is de-identified, HIPAA's consent framework no longer applies. The data enters a secondary market where consent tracking disappears entirely.

A 2023 analysis published in the Journal of Law and the Biosciences found that 87% of hospital consent forms contained no language addressing AI training, algorithmic decision-making, or commercial licensing of de-identified data. Patients consent to something narrow. Their data does something broad.

For a deeper look at the limits of HIPAA in modern data use, see our previous analysis: What HIPAA does not tell you about data trust.

The downstream data use chain in practice

Consider a concrete scenario. A breast cancer patient at an academic medical center consents to treatment and agrees to participate in a tumor registry. That registry feeds data to a cooperative group study. The cooperative group shares de-identified datasets with a commercial partner building a CDK4/6 inhibitor response prediction model.

Recent MedPage Today coverage of how abemaciclib prescribing patterns vary across institutions underscores the clinical stakes here. Only about one-third of eligible patients at one university medical center received adjuvant abemaciclib for early, high-risk, ER-positive breast cancer. If an AI model trained on registry data from that institution learns this prescribing pattern as the norm, it encodes a treatment gap as standard care. And the patient whose data trained that model never consented to their clinical experience becoming a feature in a payer's coverage algorithm.

Each step in this chain introduces new risks:

  • Re-identification risk increases as data is linked across sources. Even de-identified data, when combined with geographic, temporal, and diagnostic features, can be traced back to individuals. Our analysis of re-identification risk and HIPAA Safe Harbor covers this in detail.
  • Context collapse occurs when data collected in a care setting gets interpreted in a commercial or actuarial context.
  • Consent provenance disappears because no system tracks the original consent terms through the downstream chain.
  • Key statistics

    DTI trust score dimension weights
    DTI trust score dimension weights

  • 79% of patients in a 2023 ONC survey did not know their clinical data had been shared with third parties for secondary purposes.
  • 87% of hospital consent forms contain no language addressing AI training or commercial data licensing (Journal of Law and the Biosciences, 2023).
  • The average health data record passes through 4.7 organizations between collection and its final use in analytics or AI (Brookings Institution, 2024).
  • Only 11% of health AI developers report having formal consent provenance tracking for their training data (Stanford HAI, 2024).
  • SuperTruth's DTI framework weights Consent at 20% of the total trust score, making it the second-highest weighted dimension after Provenance (25%).
  • How secondary use erodes patient trust

    The erosion is measurable. A 2024 Pew Research survey found that 67% of Americans believe they have little to no control over how their health data is used. Among patients who learned their data had been used for purposes beyond treatment, willingness to share data in the future dropped by 41%.

    This is not an abstract problem. Clinical trial recruitment depends on patient willingness to share data. Population health programs depend on complete datasets. Precision medicine depends on diverse, representative data contributions. When patients lose trust in secondary use, the entire health data supply chain degrades.

    The consent layering problem is particularly acute for sensitive data domains. Mental health records, substance use disorder data, and reproductive health information carry heightened patient expectations about confidentiality. When these records appear in training datasets for commercial AI, the trust violation feels categorical, not incremental. Our coverage of mental health data as the most sensitive consent domain and substance use disorder data challenges examines these specific vulnerabilities.

    The regulatory landscape is shifting

    Several regulatory developments signal that consent layering is moving from an ethical concern to a compliance risk.

    Washington State's My Health My Data Act (2024) created a private right of action for unauthorized secondary uses of health data, including data that falls outside HIPAA's definition of protected health information. The law requires affirmative consent for each category of secondary use.

    The EU's European Health Data Space (EHDS) regulation, finalized in 2024, establishes a permit-based system for secondary use of electronic health data. Researchers and companies must apply for data access permits that specify the exact purpose, and the purpose cannot expand without a new permit.

    The FTC has also signaled increased scrutiny. Its 2023 enforcement actions against BetterHelp and GoodRx established that sharing health data with advertising platforms, even when technically permitted by privacy policies, constitutes an unfair practice when it exceeds consumer expectations.

    These regulatory shifts share a common thread: the era of blanket consent covering unlimited downstream use is ending.

    Five layers of consent drift

    Consent drift across downstream data use layers
    Consent drift across downstream data use layers

    Consent layering typically follows a predictable pattern of drift from the original authorization:

    Layer 1: Direct care. The patient consents to treatment. Data is used by the treating clinician and care team. This is what patients understand and expect.

    Layer 2: Operations and quality. The health system uses data for quality measurement, utilization review, and care coordination. Most consent forms cover this, but patients rarely read or understand these provisions.

    Layer 3: De-identification and research. Data is stripped of direct identifiers and shared with research partners. HIPAA permits this without additional consent. The patient has no visibility into this step.

    Layer 4: Commercial licensing. De-identified datasets are licensed to AI vendors, pharmaceutical companies, or data brokers. No consent mechanism exists for this transfer. The patient does not know it happened.

    Layer 5: Algorithmic deployment. Models trained on the data produce outputs that affect coverage decisions, risk stratification, or treatment recommendations. The patient may be directly affected by these outputs without ever knowing their data contributed to them.

    Each layer adds distance from the original consent. By layer 5, the connection between the patient's signature and the data's use is invisible to every party involved.

    Why current consent models fail

    Three structural failures explain why current consent frameworks cannot address the layering problem.

    Static consent at a dynamic boundary. Consent forms are signed once, at intake. But data use evolves continuously over months and years. A static document cannot govern a dynamic process.

    Binary architecture. Current consent is yes or no. There is no mechanism for patients to consent to research but not commercial licensing, or to consent to AI training but not payer algorithms. The lack of granularity forces patients into all-or-nothing decisions.

    No provenance chain. No standard system tracks consent terms through the data lifecycle. When data moves from a hospital to a registry to a vendor to a payer, the original consent terms are not attached to the record. By the time data reaches its final use, no one can verify what the patient actually agreed to.

    This last failure is what SuperTruth's Consent dimension in the DTI framework directly addresses. Every record scored by the DTI Engine receives a consent assessment that evaluates whether documented consent covers the record's current and proposed use. A record with layer-1 consent being used for layer-4 purposes receives a low consent score, which lowers the overall DTI score and flags the record for governance review.

    What tiered consent architecture looks like

    Solving the consent layering problem requires moving from binary consent to tiered, purpose-specific consent with machine-readable provenance.

    SuperTruth's ConsentOS implements a five-tier consent architecture:

  • Tier 1: Direct clinical care only
  • Tier 2: Quality improvement and care coordination
  • Tier 3: De-identified research with institutional oversight
  • Tier 4: Commercial analytics with patient notification
  • Tier 5: AI training and algorithmic deployment with explicit opt-in
  • Each tier is independently selectable. A patient can consent to tiers 1 through 3 while declining tiers 4 and 5. The consent terms travel with the data as machine-readable metadata, so every downstream user can verify whether their intended use falls within the patient's authorized scope.

    This is not a theoretical framework. It is a technical implementation that maps consent tiers to FHIR Consent resources and enforces tier boundaries at the point of data access.

    The employer data angle

    Consent layering is especially problematic when employer-sponsored health data enters the chain. Employees who participate in workplace wellness programs or employer-provided telehealth services generate health data that may flow to third-party vendors under enterprise agreements the employee never reviewed.

    Our analysis of employer health data and the consent governance challenge found that most employer wellness program data sharing agreements permit secondary uses that would surprise the employees generating the data. When that data reaches a health AI model, the consent gap is two organizations wide before the first algorithm runs.

    What clinicians and data teams should do now

    Four concrete steps reduce consent layering risk:

  • Audit your consent chain. Map every downstream use of patient data back to the original consent terms. If you cannot trace the chain, you have a consent layering problem.
  • Score consent coverage. Use the DTI framework's consent dimension to assess whether each record's consent terms cover its current use. Records with consent gaps should be flagged before they enter AI pipelines.
  • Implement tiered consent. Replace binary consent forms with purpose-specific consent that patients can select at a granular level. Attach consent tiers as machine-readable metadata that travels with the data.
  • Track consent provenance. Build or adopt systems that maintain the chain of custody between consent and data use. When a downstream user queries data, the system should automatically verify that the query's purpose falls within the consent scope.
  • These are not aspirational recommendations. They are engineering requirements for any organization that wants its health data to carry a DTI score above Bronze grade.

    The cost of ignoring consent layering

    The financial exposure from consent layering failures is accelerating. Washington State's My Health My Data Act allows damages of up to $25,000 per violation. The FTC's BetterHelp settlement required $7.8 million in consumer refunds. The EU's EHDS regulation carries penalties aligned with GDPR, meaning fines of up to 4% of global annual revenue.

    Beyond regulatory penalties, there is the operational cost of data that cannot be used. Organizations that discover consent gaps in their training data face two options: retrain models without the affected data, or accept the legal and reputational risk of using data that exceeds its consent scope. Both are expensive. Prevention is cheaper.

    The DTI Engine scores every health data record 0-100 across 8 trust dimensions before your AI model sees it. Consent is weighted at 20%, the second-highest dimension, because downstream data use consent is the fastest-growing compliance risk in health AI. If your team is evaluating data for training, compliance, or clinical use and needs to verify that consent coverage matches intended use, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.

    Further reading:

  • DTI™ Engine
  • Health systems solution
  • ConsentOS: what five-tier consent architecture looks like in practice
  • Why consent governance fails in healthcare data and what fixes it
  • The HIPAA problem health AI companies are ignoring: patient consent does not cover model training
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0–100. Travels with every record permanently.

    See the DTI Engine
    Share