Synthetic data in healthcare AI: when fabricated training data creates real bias
Photo by BoliviaInteligente on Unsplash

Synthetic data in healthcare AI: when fabricated training data creates real bias

By Jason Alan Snyder·May 7, 2026

Synthetic health data inherits, amplifies, and launders the biases present in real-world clinical datasets. When healthcare AI models train on fabricated records that were never scored for provenance or consent, the resulting bias becomes invisible to standard audits. The only defense is scoring data trust before any model trains on it.

A synthetic patient record looks clean. It has no missing fields, no duplicated entries, no obvious formatting errors. It also has no provenance, no consent trail, and no connection to a real person whose biology behaved in a specific way at a specific time. That cleanliness is the problem.

The promise and the trap of synthetic health data

Synthetic data generation in healthcare follows a simple logic: real patient data is scarce, expensive, and privacy-constrained, so algorithms generate statistically similar records to fill the gap. Generative adversarial networks (GANs), variational autoencoders, and large language models can produce millions of synthetic patient records in hours.

The appeal is obvious. A 2023 study in Nature Medicine estimated that synthetic data could reduce dataset acquisition costs by up to 60% for clinical AI development. Hospitals avoid HIPAA exposure. Researchers get larger sample sizes. Everyone moves faster.

But synthetic records are derivatives. They inherit every distributional assumption baked into the source data. If the original dataset underrepresents Black women with heart failure, the synthetic version will too. Worse, it will do so without any flag, footnote, or metadata indicating the gap exists.

How synthetic health data bias enters the pipeline

Bias in healthcare AI synthetic training does not arrive through a single failure. It compounds across three stages.

First, source data selection. Most synthetic generators train on EHR exports from large academic medical centers. These centers serve populations that skew whiter, more insured, and more urban than the national average. The synthetic output mirrors that skew.

Second, generation assumptions. GANs optimize for statistical fidelity to the training distribution. They do not optimize for equity. If 4% of the source records represent Native American patients, the generator will treat 4% as the correct proportion, even when the target population is 12%.

Third, downstream validation. Teams evaluate synthetic data on completeness, format consistency, and distributional similarity to the source. None of these checks measure whether the data represents the population the model will serve. A synthetic dataset can score perfectly on every quality metric and still encode dangerous gaps.

Why standard data quality checks miss synthetic bias

DTI dimension weights: where synthetic data fails
DTI dimension weights: where synthetic data fails

Data quality and data trust are not the same thing. Quality measures whether a field is populated and formatted correctly. Trust measures whether a record should be believed, used, and relied upon for a specific purpose.

Synthetic records pass quality checks by design. They are generated to be complete. But they fail on provenance (there is no original source), consent (no patient agreed to this record's creation), and recency (the temporal dynamics of the source data may be years old). These are three of the eight dimensions SuperTruth's Data Trust Index (DTI) scores on every record. Without them, a healthcare AI system is training on records that cannot be audited, traced, or validated against reality.

This is not a theoretical concern. When the FDA issued its 2024 draft guidance on AI/ML-enabled devices, it explicitly called for documentation of training data provenance. Synthetic records without provenance metadata create a regulatory gap that will only widen as enforcement matures. For more on what is coming, see FDA AI guidance and data provenance: what is coming for healthcare AI developers.

Can synthetic data ever be used responsibly in healthcare AI?

Yes, but only under strict conditions. Synthetic data can augment underrepresented subgroups when the augmentation is intentional, documented, and transparent. It can stress-test models against edge cases. It can enable privacy-safe collaboration between institutions.

The problem is not synthesis itself. The problem is treating synthetic records as interchangeable with real ones. When a training pipeline mixes synthetic and real data without tagging, scoring, or separating them, synthetic data trust in healthcare collapses. The model cannot distinguish what it learned from a real patient outcome versus what it learned from a statistical approximation.

Every record that enters a training pipeline needs a trust score before the model sees it. Real records need scoring for recency, provenance, and consent. Synthetic records need an additional flag: they must be identified as synthetic, scored on the fidelity of their source, and constrained in how much weight they carry during training. The DTI Engine applies this logic at the point of ingestion, not after a model has already absorbed the bias.

Key statistics

imaware data processing: before vs after DTI scoring
imaware data processing: before vs after DTI scoring

  • Synthetic data can reduce dataset acquisition costs by up to 60%, but without provenance scoring, cost savings become bias liabilities (Nature Medicine, 2023)
  • The DTI scores every health record 0-100 across 8 dimensions; synthetic records with no provenance or consent trail score below 30 on two dimensions worth a combined 45% of the total weight
  • SuperTruth's work with imaware standardized 105,000 diagnostic records, reducing processing time from 3 weeks to 2 hours, a 95% time reduction
  • FDA's 2024 draft guidance on AI/ML devices requires training data provenance documentation, which synthetic records without metadata cannot satisfy
  • Academic medical centers, the primary source for synthetic data generators, serve populations that are on average 15-20% less racially diverse than the U.S. population they claim to represent
  • What happens when synthetic bias reaches patients

    The downstream effects are measurable. AI models trained on biased synthetic data produce risk scores that systematically underestimate disease severity in underrepresented groups. A 2019 Science study found that a widely used healthcare algorithm assigned lower risk scores to Black patients, resulting in Black patients needing to be significantly sicker than white patients to receive the same level of care. That algorithm trained on real data. Synthetic data, unchecked, would replicate and potentially amplify the same pattern with no audit trail to catch it.

    This is why health equity data work and synthetic data governance cannot be separated. If your synthetic pipeline does not account for who is missing from the source, it will generate more of the same absence.

    The fix is trust scoring, not better synthesis

    Better GANs will not solve this. More sophisticated generation techniques produce more convincing synthetic records, which makes the bias harder to detect, not easier. The fix is upstream: score every record for trust before it enters a training pipeline. Tag synthetic records explicitly. Enforce DTI floor thresholds that prevent low-provenance, low-consent data from reaching production models.

    SuperTruth built the DTI Engine for exactly this problem. When imaware needed to standardize 105,000 diagnostic records, the DTI pipeline scored each one across all eight dimensions and identified which records met the threshold for AI training use. The same logic applies to synthetic records. If a record cannot prove where it came from, it does not get to train a model that will make clinical decisions.

    The DTI Engine scores every health data record 0-100 across 8 trust dimensions before your AI model sees it. If your team is evaluating data for training, compliance, or clinical use, and especially if synthetic data is part of your pipeline, schedule a conversation with the SuperTruth commercial team or (215) 918-4140.

    Further reading:

  • DTI™ Engine
  • Health systems solution
  • AI explainability solves the wrong problem: why trusting a model output means nothing if the training data was never verified
  • Why AI models trained on unscored health data will fail in production
  • Glass box vs black box: why health AI needs explainable data provenance
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0 to 100. Travels with every record permanently.

    See the DTI Engine
    Share
    Synthetic data in healthcare AI: when fabricated training data creates real bias | SuperTruth