Health data broker accountability: what the FTC approach means for AI vendors
The FTC has settled with multiple health data brokers for unlawful sale of sensitive location and health data, and its enforcement model is shifting from reactive penalties to structural prohibitions. For AI vendors building on brokered health data, these actions redefine what accountability looks like, and what data trust infrastructure is now required to avoid regulatory exposure.
The FTC banned two data brokers from selling sensitive health data in 2024. Not fined. Banned. That distinction matters for every AI vendor whose training pipeline touches brokered health data.
The settlements with Outlogic (formerly X-Mode Social) and InMarket Solutions did not just impose financial penalties. They imposed structural prohibitions: no selling, no disclosing, no using sensitive location data tied to health facilities. The FTC also required deletion of previously collected data and the algorithms derived from it. That last part is the one most AI vendors have not internalized.
If the data you trained on gets retroactively prohibited, the model trained on it may need to be destroyed. That is not a hypothetical. That is the FTC's stated enforcement position.
Why health data broker accountability is different now
For years, health data brokers operated in a regulatory gap. HIPAA covers covered entities: hospitals, insurers, clearinghouses. It does not cover the data broker that aggregates location signals from mobile apps to infer who visits an oncology clinic, a substance abuse treatment center, or a reproductive health facility.
The FTC stepped into that gap using its Section 5 authority over unfair and deceptive practices. The agency's theory is straightforward: if a consumer did not know their location data would be sold to identify their health conditions, the sale is unfair. If a company promised privacy and then sold the data, the sale is deceptive.
This is not abstract. The FTC's complaint against Outlogic specifically cited the sale of location data that could reveal visits to hospitals, mental health clinics, and addiction treatment centers. The complaint against InMarket cited similar practices with geofencing data around sensitive health locations.
The enforcement signal is clear: health-adjacent data is now health data in the FTC's view, regardless of whether it meets HIPAA's technical definition.
How does the FTC regulate AI?
The FTC does not have a single AI-specific statute. Instead, it applies existing authority under Section 5 of the FTC Act to AI systems that cause consumer harm. The agency has published guidance making clear that AI-powered tools are not exempt from truth-in-advertising rules, fair lending obligations, or data protection requirements simply because a model made the decision instead of a human.
Three enforcement mechanisms matter most for health AI vendors:
Algorithmic disgorgement. The FTC can require companies to delete not just the improperly collected data but also the models trained on it. This happened in the Everalbum case (2021), where the company had to destroy facial recognition models built on deceptively obtained photos. The same logic applies to health AI models trained on brokered data obtained without proper consent.
Unfairness doctrine. A practice is unfair if it causes substantial injury to consumers, is not reasonably avoidable by consumers, and is not outweighed by countervailing benefits. Selling inferred health conditions derived from location data meets all three prongs.
Deception doctrine. If a company represents that data is anonymized or that users consented to its use, and those representations are false, the FTC treats it as deceptive regardless of whether the company intended to deceive.
For AI vendors, this means the provenance of every training data record matters. Not just at the point of collection, but at every step in the chain of custody.
What does the FTC consider a deceptive ad?
The FTC defines a deceptive advertisement as any representation, omission, or practice that is likely to mislead a consumer acting reasonably under the circumstances. The representation must be material, meaning it would affect the consumer's decision.
In the health data context, this standard applies directly to how AI vendors and data brokers describe their data practices. If a data broker tells app developers that user data will be used only for analytics but then sells it for health profiling, that is deceptive. If an AI vendor tells health systems that its training data was ethically sourced but cannot prove consent chains, that claim is vulnerable.
The FTC has been explicit that claims about AI accuracy and data sourcing are subject to the same substantiation requirements as any other advertising claim. Saying "our model was trained on consented data" without being able to prove it is a regulatory risk, not a marketing choice.
How does the FTC protect consumers?
The FTC protects consumers through enforcement actions, rulemaking, and public guidance. In the health data space, the agency has been unusually aggressive since 2022, bringing more data broker cases in two years than in the prior decade.
Key consumer protection mechanisms include:
The FTC also coordinates with state attorneys general, who have brought parallel actions under state consumer protection laws. Washington, Connecticut, and California have enacted specific health data privacy statutes that create additional liability layers for data brokers and the AI vendors who buy from them.
What types of data are prohibited in public AI systems?
The FTC's enforcement actions and the emerging state law landscape point to several categories of health data that face outright prohibition or severe restrictions in AI systems:
For AI vendors, the category that matters most is inferred health data. You do not need to collect a diagnosis code to create a health data liability. If your model infers a health condition from behavioral, location, or purchase data, the FTC treats that inference as health data subject to the same protections.
The consent chain problem for AI vendors
Most AI vendors do not collect health data directly. They buy it from brokers, who aggregate it from app developers, who collect it from consumers who clicked "I agree" on a terms-of-service document. The FTC has made clear that this chain does not insulate downstream users from liability.
The consent given to an app developer for "analytics purposes" does not transfer to a data broker for "health profiling purposes," and it certainly does not transfer to an AI vendor for "model training purposes." Each link in the chain requires its own valid consent, and the FTC evaluates consent based on what a reasonable consumer would understand, not what the fine print technically permits.
This creates a concrete problem for health AI vendors. If you cannot trace every training data record back to a specific, informed consent event that covers your intended use, you have a regulatory exposure. The FTC's algorithmic disgorgement remedy means that exposure could cost you not just a fine but your entire model.
Key statistics
The scope of the health data broker accountability problem is measurable:
What fairness and bias mitigation require in this context
The FTC's data broker enforcement creates a second-order problem for AI fairness. When health data is collected through location tracking and behavioral inference, the data disproportionately reflects populations who use mobile apps, visit brick-and-mortar health facilities, and live in areas with dense geofencing coverage.
Rural populations, elderly patients, and communities with limited smartphone adoption are systematically underrepresented. Populations that avoid formal health settings due to stigma, immigration status, or distrust are also invisible in brokered data sets.
AI models trained on this data do not just have a privacy problem. They have a bias problem. The FTC has signaled that disparate impact from AI systems can constitute an unfair practice under Section 5, which means that biased training data is not just an ethical concern but a legal one.
Addressing this requires more than removing protected attributes from training data. It requires scoring every record for representativeness, completeness, and provenance before it enters a training pipeline. This is exactly what data trust scoring was designed to do.
Transparency and explainability as regulatory requirements
A key aspect of transparency in healthcare AI systems is the ability to trace every output back to the data that produced it. The FTC's enforcement model makes this more than a best practice. If the agency investigates your model, you need to show where the training data came from, what consent covered its use, and whether any of it originated from a prohibited source.
This is not the same as model explainability. Explaining why a model produced a specific prediction is useful, but it does not answer the FTC's core question: was the data lawfully obtained and properly consented for this use?
Provenance tracking, consent chain documentation, and data retention policies are the transparency mechanisms that matter for regulatory compliance. A model that can explain its reasoning but cannot prove its data was clean is still exposed.
The data trust infrastructure AI vendors need
The FTC's approach to health data broker accountability creates a clear infrastructure requirement for AI vendors. You need three capabilities that most vendors currently lack:
1. Provenance scoring at the record level. Every data record entering your pipeline needs a documented chain of custody showing where it originated, how it was collected, and what transformations it underwent. Batch-level provenance is not sufficient. The FTC evaluates data practices at the individual consumer level.
2. Consent verification that covers your specific use. Generic consent for "data analytics" does not cover health AI model training. You need consent documentation that specifically addresses how you will use the data, and you need a system that flags records where consent does not match intended use.
3. Temporal validity tracking. The FTC's data deletion requirements mean that data you were legally allowed to use yesterday may be prohibited today. Your infrastructure needs to track consent revocations, regulatory changes, and data expiration in near-real time.
These three capabilities map directly to the Provenance, Consent, and Recency dimensions of the Data Trust Index. They are not optional features. They are the minimum infrastructure required to operate in the FTC's current enforcement environment.
What this means for health systems buying AI
Health systems evaluating AI vendors now have a new due diligence requirement. Before deploying any AI tool, they need to ask: where did the training data come from, and can the vendor prove every record was lawfully obtained?
If the vendor used brokered data, the health system needs to verify that the broker's data practices comply with FTC requirements. If the vendor cannot produce consent documentation for its training data, the health system inherits the regulatory risk.
This is not theoretical. The FTC's algorithmic disgorgement remedy means that a health system could deploy an AI tool, integrate it into clinical workflows, and then be forced to remove it because the underlying model was trained on prohibited data. The operational disruption from that scenario dwarfs the cost of proper data trust scoring upfront.
The regulatory trajectory is acceleration, not stabilization
The FTC's 2024 data broker settlements are not the endpoint. They are the starting point. The agency has signaled through speeches, blog posts, and proposed rules that it intends to expand enforcement against health data brokers and the AI vendors who rely on them.
State legislatures are moving even faster. Washington's My Health My Data Act, effective March 2024, creates a private right of action for consumers whose health data is sold without consent. Similar bills are advancing in multiple states. The patchwork of state laws creates compliance complexity that only scalable, automated data trust infrastructure can address.
AI vendors who treat FTC compliance as a one-time legal review rather than an ongoing infrastructure requirement will find themselves in the same position as the data brokers the FTC just banned: explaining to regulators why they thought the rules did not apply to them.
The DTI Engine scores every health data record 0 to 100 across 8 trust dimensions before your AI model sees it. If your team is evaluating data for training, compliance, or clinical use and needs to prove that every record in your pipeline meets FTC, state, and HIPAA requirements, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.
Further reading:

Jason Alan Snyder
Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.
About SuperTruth · LinkedIn · Substack · jasonalansnyder.com
See it in practice
DTI scores the record, not the patient.
8 dimensions. 0–100. Travels with every record permanently.