Amazon Clinic and the retail health data trust gap
Photo by Declan Sun on Unsplash

Amazon Clinic and the retail health data trust gap

By Jason Alan Snyder·May 13, 2026

Amazon Clinic collected health data from millions of consumers before quietly shutting down in 2024. The retail health data trust gap it exposed remains unsolved: consumer health records generated outside traditional clinical settings lack provenance, consent governance, and quality scoring. Until health data has a trust layer, retail health platforms will keep creating records that cannot be safely used for AI, research, or clinical decision-making.

Amazon Clinic launched in November 2022, treated patients across all 50 states by early 2023, and shut down in late 2024. In between, it generated health records for millions of consumers who filled out intake questionnaires, received asynchronous diagnoses, and got prescriptions filled through Amazon Pharmacy. Those records still exist. The question nobody answered before, during, or after Amazon Clinic's operation is simple: can any downstream system actually trust that data?

The answer matters far beyond Amazon. CVS Health, Walmart Health (also shut down in 2024), Walgreens, and dozens of digital-first platforms have created a growing pool of consumer health data that sits outside traditional clinical infrastructure. This data has no standardized provenance chain. Its consent governance ranges from buried Terms of Service clauses to nonexistent. And its quality has never been independently scored.

This is the retail health data trust gap. It is the single largest unaddressed risk in consumer health data today.

What happened to Amazon Clinic?

Amazon launched Amazon Clinic as a virtual health service offering message-based consultations for 30+ common conditions, including allergies, acne, urinary tract infections, and hair loss. Patients completed structured intake forms, a third-party clinician reviewed the information asynchronously, and prescriptions were routed to Amazon Pharmacy.

By mid-2023, Amazon had also completed its $3.9 billion acquisition of One Medical, a primary care membership service with physical clinics. The two services overlapped in ways that created strategic confusion. Amazon Clinic handled low-acuity, high-volume virtual visits. One Medical handled longitudinal, in-person primary care.

Amazon quietly discontinued Amazon Clinic in December 2024, folding virtual care capabilities into the One Medical platform. The shutdown was not widely publicized. Amazon framed it as consolidation, not failure.

But the health data generated during Amazon Clinic's two-year run did not disappear. It remains in Amazon's systems, governed by Amazon's privacy policies rather than HIPAA in many cases, since Amazon Clinic operated as a technology platform connecting patients to independent clinicians rather than as a covered entity itself.

Why did Amazon introduce Amazon Clinic?

Amazon's entry into healthcare followed a clear commercial logic. The company already operated Amazon Pharmacy after acquiring PillPack in 2018. Adding a virtual care layer created a closed loop: consult, diagnose, prescribe, dispense. Every step happened within Amazon's platform.

The broader strategy targeted the estimated $4.3 trillion U.S. healthcare market. Amazon saw what every retail company sees: healthcare spending is the largest single category of consumer expenditure, and most of it flows through fragmented, analog systems. Amazon Clinic was designed to capture the low-acuity, high-frequency segment of that market.

There was also a data play. Amazon already possessed the world's most detailed consumer purchase history. Adding health data to that profile would create a behavioral dataset of unprecedented depth. A consumer who buys antacids weekly, searches for acid reflux symptoms, and then consults Amazon Clinic about GERD represents a data trifecta: purchase behavior, search behavior, and clinical information, all linked to a single identity.

This is precisely what made privacy advocates nervous from the start.

Is Amazon selling my data to any third parties?

Amazon's privacy policy for health services states that it does not sell personal health information to third parties. However, "selling" is a narrow legal term. Amazon's policies permit the use of health data to "improve services," "develop new features," and "provide personalized experiences." These are broad permissions that allow internal use of health data across Amazon's business units without constituting a "sale."

The distinction matters. Under HIPAA, covered entities face strict limits on how they use Protected Health Information (PHI). But Amazon Clinic's structure placed much of the data handling under Amazon's general privacy policy rather than HIPAA's framework. The third-party clinicians who treated patients were covered entities. Amazon, as the technology platform, occupied a gray zone.

Washington state's My Health My Data Act, which took effect in 2024, specifically targets non-HIPAA health data collected by technology companies. Several other states have passed or proposed similar legislation. These laws exist because the federal framework has a gap: consumer health data generated outside traditional clinical relationships has weaker protections than most consumers assume.

As we have written before, HIPAA Safe Harbor is not sufficient for modern AI. When retail platforms collect health data under general consumer privacy policies, the gap becomes even wider.

What are the drawbacks of Amazon Clinic?

The drawbacks of Amazon Clinic fell into three categories: clinical, structural, and data-related.

Clinically, asynchronous message-based care limits diagnostic accuracy. Patients completed structured questionnaires, but clinicians could not perform physical examinations, observe non-verbal cues, or order immediate follow-up tests. For the conditions Amazon Clinic treated, this was often adequate. But the model incentivized throughput over thoroughness.

Structurally, Amazon Clinic created health records that were disconnected from patients' existing medical histories. A patient's primary care physician had no automatic access to Amazon Clinic visit records. Prescriptions written through the platform did not flow into the patient's EHR. This fragmentation, which we have examined extensively in The fragmented health record, created clinical blind spots.

From a data trust perspective, the drawbacks were the most concerning. Amazon Clinic records lacked provenance documentation (where did the data originate, who handled it, what transformations occurred). Consent was bundled into Terms of Service agreements that most consumers accepted without reading. Data quality was unscored and unverified. And recency tracking, the ability to know how current a record is and when it was last validated, was nonexistent as a formal dimension.

These are not abstract concerns. They determine whether the data can be safely used for AI model training, clinical decision support, population health analytics, or research.

The retail health data gap is structural, not reputational

Most coverage of Amazon Clinic's data risks focused on consumer trust: will people trust Amazon with their health information? An AMA survey found that only 37% of patients trusted big tech companies to handle their health data responsibly. STAT News documented the demographic gap between One Medical's clinic locations and average U.S. neighborhoods.

These are real concerns. But they address the wrong layer of the problem.

The fundamental issue is not whether consumers trust Amazon. It is whether downstream systems, AI models, research platforms, clinical decision tools, payer algorithms, can trust the data Amazon Clinic generated. Consumer sentiment is a perception problem. Data trust is an engineering problem.

Every health data record has measurable properties: where it came from (provenance), whether the patient authorized its specific uses (consent), how recently it was validated (recency), whether it meets structural standards (quality), whether it agrees with other records about the same patient (concordance), whether it has been independently verified (validation), how many data dimensions it covers (breadth), and whether it remains consistent over time (stability).

Amazon Clinic data scores poorly on most of these dimensions. Not because Amazon is uniquely bad, but because no retail health platform has built infrastructure to address them.

Key statistics

Patient trust in health data custodians
Patient trust in health data custodians

Amazon completed the One Medical acquisition for $3.9 billion in February 2023, creating the financial foundation for its combined virtual and in-person care strategy.

An AMA survey found only 37% of patients trust big tech companies with their health data, compared to 75% who trust their personal physician.

The U.S. healthcare market represents $4.3 trillion in annual spending, with low-acuity virtual visits being the fastest-growing segment before Amazon Clinic's shutdown.

SuperTruth's work with imaware standardized 105,000 diagnostic records and reduced processing time from 3 weeks to 2 hours, a 95% reduction, demonstrating that trust scoring at scale is operationally feasible.

An estimated 40% of consumer health interactions now occur outside traditional clinical settings, through telehealth platforms, retail clinics, wellness apps, and direct-to-consumer testing, creating a growing volume of health data with no trust layer.

What happens to health data after the platform shuts down?

Walmart Health closed all 51 of its clinics in 2024. Amazon Clinic shut down the same year. Babylon Health went bankrupt in 2023. Each of these platforms generated health records during their operational periods. Each left behind data that now sits in systems with uncertain governance.

When a traditional healthcare provider closes, there are established protocols for medical record retention and transfer. State laws typically require records to be maintained for 7 to 10 years. Patients must be notified. Successor custodians must be designated.

Retail health platforms operate under different, often weaker, obligations. Amazon Clinic's data likely falls under Amazon's general data retention policies. For patients who want their records, the path to access is unclear. For researchers or AI developers who might want to use aggregated Amazon Clinic data, the provenance chain is broken before it begins.

This is the lifecycle problem that retail health data introduces. Traditional health data has a custodial chain, imperfect but documented. Retail health data has a Terms of Service agreement and a corporate parent that may or may not continue operating the relevant business unit.

Why consent governance breaks in retail health

Amazon Clinic obtained consent the way every consumer technology platform does: through a click-through agreement during account creation. This consent was broad, covering data use for service delivery, improvement, analytics, and related purposes.

Clinical consent works differently. HIPAA requires specific authorization for uses beyond treatment, payment, and healthcare operations. Research use requires IRB approval and often separate informed consent. Marketing use requires explicit patient authorization.

Retail health platforms collapse these distinctions. A single click-through covers everything from treatment to analytics to product development. The patient has no mechanism to consent to clinical use of their data while restricting its use for algorithm training or internal analytics.

We built ConsentOS specifically to address this problem. Five-tier consent architecture allows patients to grant different levels of authorization for different uses of their data. Treatment: yes. Research: yes, with restrictions. AI training: no. Marketing: no. Each tier is independently tracked and enforced.

No retail health platform has implemented anything comparable. The gap between the consent consumers think they gave and the consent the platform actually obtained is one of the largest unaddressed risks in consumer health data.

The data quality problem nobody measured

Amazon Clinic intake questionnaires were designed for clinical efficiency, not data quality. Patients self-reported symptoms, medical history, and medication lists. There was no validation against existing medical records. There was no concordance check against pharmacy data, lab results, or prior diagnoses.

Self-reported health data is notoriously unreliable. Studies show that patients misreport medication adherence 30% to 50% of the time. Symptom descriptions vary dramatically based on health literacy, cultural context, and the specific phrasing of questions.

None of this means Amazon Clinic data is worthless. It means it is unscored. Without a trust score, no downstream system can determine which records are reliable and which are not. An AI model trained on a mix of high-quality and low-quality records will produce unreliable outputs. A research study that includes unvalidated self-reported data alongside clinician-verified records will draw flawed conclusions.

This is exactly the problem we have documented with unscored health data in AI training pipelines. The problem is not unique to Amazon. But Amazon's scale makes it uniquely consequential.

The trust layer retail health never built

Data Trust Index: 8 dimensions of health data trust scoring
Data Trust Index: 8 dimensions of health data trust scoring

Amazon built exceptional logistics infrastructure for healthcare. Pharmacy fulfillment. Scheduling systems. Payment processing. Clinical workflow tools. These are infrastructure problems, and Amazon solves infrastructure problems better than almost any company on earth.

But infrastructure trust is not data trust. We have written about this distinction before. AWS can guarantee 99.99% uptime for its health data lake. That tells you nothing about whether the data inside it is accurate, current, properly consented, or safe to use for a specific purpose.

The trust layer that retail health needs sits between the platform and every downstream use of the data. It scores every record before it enters a training pipeline. It enforces consent at the use level, not the platform level. It tracks provenance from the moment of data creation through every transformation and transfer. It validates quality against independent sources.

This layer does not exist in any retail health platform today. Not Amazon's. Not CVS's. Not Walmart's before it shut down. Not any of the dozens of telehealth startups generating consumer health data at scale.

What this means for the industry

The retail health data trust gap will only grow. Consumer health interactions are migrating out of traditional clinical settings at an accelerating rate. Direct-to-consumer lab testing, remote patient monitoring, wellness apps, retail clinics, and virtual care platforms are generating health data that feeds into AI models, population health tools, and research datasets.

Every one of these records needs a trust score before any downstream system touches it. Not a binary pass/fail. A granular score across multiple dimensions that tells the consuming system exactly what it is working with.

Amazon Clinic's shutdown does not solve the problem it created. The data remains. The gap remains. And the next platform to enter retail health will face the same structural challenge unless the industry builds the trust layer that Amazon never did.

The DTI Engine scores every health data record 0 to 100 across 8 trust dimensions before your AI model sees it. If your team is evaluating consumer health data for training, compliance, or clinical use, and you need to know which records you can actually trust, schedule a conversation with the SuperTruth commercial team or (215) 918-4140.

Further reading:

  • DTI™ Engine
  • Health systems solution
  • Digital health app data: the gap between consumer trust and clinical trust
  • Re-identification risk: why HIPAA Safe Harbor is not sufficient for modern AI
  • Infrastructure trust vs data trust: why most healthcare data platforms miss the point
  • Why AI models trained on unscored health data will fail in production
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0 to 100. Travels with every record permanently.

    See the DTI Engine
    Share
    Amazon Clinic and the retail health data trust gap | SuperTruth