You can now score a health record live on supertruth.ai — what the DTI pipeline actually does and why it matters
Photo by 灿雄 邱 on Unsplash
insight

You can now score a health record live on supertruth.ai — what the DTI pipeline actually does and why it matters

By Jason Alan Snyder·April 26, 2026

SuperTruth now lets you score a health data record live on supertruth.ai. The DTI pipeline evaluates every record across 8 trust dimensions and returns a 0-100 score before any AI model touches it. This post explains what happens inside that pipeline, step by step, and why it changes how health data enters production.

Most health data enters AI pipelines unscored. Nobody checks whether the record has valid provenance, current consent, or consistent values across sources. The model just ingests it. That era is over.

You can now score a health data record live on supertruth.ai. Upload a record, run the DTI pipeline, and get a 0-100 trust score across eight dimensions before your model ever sees it. This is not a demo. It is a production-grade scoring engine available today.

What is SuperTruth?

SuperTruth is the trust layer underneath healthcare data and AI. We built the Data Trust Index (DTI), which functions like a FICO score for health data. Every record that enters the pipeline receives a composite score from 0 to 100 based on eight weighted dimensions: Provenance (25%), Consent (20%), Recency (15%), Quality (10%), Concordance (10%), Validation (10%), Breadth (5%), and Stability (5%).

The score is not advisory. It is enforceable. You set a DTI floor, and records that fall below it never reach your model, your analytics layer, or your regulatory submission.

What the DTI pipeline actually does

The pipeline runs eight discrete checks on every record that enters it.

First, Provenance. The pipeline traces where the record originated, how many systems it passed through, and whether chain of custody is intact. This dimension carries the heaviest weight at 25% because a record with broken provenance cannot be trusted regardless of how clean its fields look. For more on why this is the hardest dimension to get right, see The chain of custody problem in health data.

Second, Consent. The engine checks whether the record carries a valid, machine-readable consent artifact that covers the intended use. Not just "was consent obtained," but "does this consent cover model training, secondary research, or commercial analytics." A record consented for care delivery that ends up in an AI training set scores poorly here.

Third, Recency. A provider address from 2019 or a lab result from 14 months ago carries less weight than data captured this quarter. The pipeline timestamps every field and penalizes staleness.

Fourth through eighth: Quality checks for completeness and formatting. Concordance compares the record against other sources to flag contradictions. Validation confirms the record against authoritative references. Breadth measures how many relevant data elements the record contains. Stability tracks whether the record's values have changed erratically over time.

Each check runs independently. The weighted composite becomes the DTI score.

Key statistics

DTI dimension weights: how each dimension contributes to the 0-100 score
DTI dimension weights: how each dimension contributes to the 0-100 score

The numbers behind the pipeline tell the story better than any description.

  • 105,000 diagnostic records scored and standardized in the imaware partnership, the largest DTI deployment to date.
  • 95% time reduction in data preparation: what took 3 weeks manually now completes in 2 hours through the DTI pipeline.
  • 200+ hours per month saved in ongoing data operations after DTI deployment.
  • 8 trust dimensions evaluated per record, with Provenance (25%) and Consent (20%) accounting for nearly half the total score.
  • 1 segment identified through DTI scoring that drove 20% of imaware's revenue, a pattern invisible in unscored data.
  • Why scoring before ingestion changes everything

    imaware data preparation time: before and after DTI pipeline
    %22%7D%7D%5D%7D%7D%7D) imaware data preparation time: before and after DTI pipeline

    The standard workflow in health AI is: collect data, clean data, train model, discover problems in production. DTI inverts this. Problems surface at the point of ingestion, not after a model has already learned from flawed records.

    This matters for three audiences.

    For health systems deploying clinical AI, a DTI score on every incoming record means you can answer an auditor's question: "How do you know this training data was trustworthy?" You can answer it with a number, not a narrative. See Hospital system AI readiness for what infrastructure you need before deployment.

    For pharma companies building RWE packages, scored data creates an audit trail that follows the record from source to FDA submission. Every record carries its own trust receipt. We wrote about what Platinum-grade data requires in What makes health data Platinum-grade for FDA regulatory submission.

    For health plans managing provider directories, the DTI pipeline catches stale addresses, broken NPI concordance, and consent gaps before they become CMS CRUSH findings. One scored record is worth more than a hundred unverified ones.

    What you see when you score a record live

    The live scoring tool on supertruth.ai accepts a health data record and returns a visual breakdown. You see the composite score, the individual dimension scores, and the specific flags that pulled the score down.

    A record might score 82 overall but carry a Consent score of 41 because the consent artifact does not cover secondary use. That granularity is the point. You do not just know a record is weak. You know exactly where it is weak and what to fix.

    Brodie Flanders, CEO of imaware, put it directly: "The lab industry has never had a trust standard. DTI created one."

    What this is not

    This is not data cleaning. Data cleaning fixes formatting. DTI scores trustworthiness. A record can be perfectly formatted and still score 30 because its provenance is broken and its consent expired two years ago.

    This is not a one-time audit. The pipeline runs continuously. Records get rescored as conditions change. A consent revocation drops the score in real time. Temporal drift in field values triggers a Stability flag. For a deeper look at how drift destroys model accuracy, see How temporal drift destroys AI model accuracy in healthcare.

    Try it now

    The DTI Engine scores every health data record 0-100 across 8 trust dimensions before your AI model sees it. If your team is evaluating data for training, compliance, or clinical use, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140. You can also score a synthetic patient record live at supertruth.ai/dti right now — no login required, synthetic data only, results in under five seconds.

    Further reading:

  • Try the live DTI™ demo
  • DTI™ Engine
  • Health systems solution
  • The eight dimensions of health data trust: a practical guide
  • Why AI models trained on unscored health data will fail in production
  • The Data Trust Index: SuperTruth publishes the first formal framework for health data integrity scoring
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0–100. Travels with every record permanently.

    See the DTI Engine
    Share
    You can now score a health record live on supertruth.ai — what the DTI pipeline actually does and why it matters | SuperTruth