SuperTruth publishes peer-reviewed research on the Data Trust Index at Zenodo — what the paper covers and why it matters for health AI governance
Photo by Jaykumar Bherwani on Unsplash
insight

SuperTruth publishes peer-reviewed research on the Data Trust Index at Zenodo — what the paper covers and why it matters for health AI governance

By Jason Alan Snyder·April 26, 2026

SuperTruth has published peer-reviewed research on the Data Trust Index (DTI) at Zenodo, making the formal methodology for scoring health data integrity publicly available. The paper defines the 8-dimension framework, weighted scoring model, and trust tier classification system that scores every health data record 0 to 100. This is the first open-access publication of a structured trust scoring methodology purpose-built for health AI governance.

SuperTruth has published the formal research paper describing the Data Trust Index (DTI) on Zenodo, the open-access repository operated by CERN. The paper lays out the complete methodology for scoring health data records on a 0 to 100 scale across eight weighted dimensions. It is now publicly available for citation, peer review, and institutional adoption.

This matters because health AI governance has operated without a standardized way to measure data trustworthiness. Regulatory bodies like the FDA are asking about training data provenance. Payers want to know if the records behind AI-driven claims decisions are current and validated. Research consortia need a shared language for data quality thresholds across institutions. The DTI paper provides that shared language.

What the paper covers

DTI dimension weights: how the 100-point score is allocated
DTI dimension weights: how the 100-point score is allocated

The Zenodo publication defines the Data Trust Index as a composite score derived from eight dimensions, each carrying a specific weight: Provenance (25%), Consent (20%), Recency (15%), Quality (10%), Concordance (10%), Validation (10%), Breadth (5%), and Stability (5%).

These weights were not arbitrary. The paper explains the rationale behind each allocation. Provenance carries the highest weight because a record with unknown origin cannot be trusted regardless of its other attributes. Consent follows because health data used without proper authorization creates both legal liability and ethical failure. Recency is third because stale data produces stale predictions, and temporal drift destroys AI model accuracy in healthcare faster than most teams realize.

The paper also introduces the trust tier classification system. Records scoring 90 to 100 earn Platinum status, suitable for FDA regulatory submission. Gold (75 to 89) qualifies for clinical AI training. Silver (50 to 74) supports operational analytics with caveats. Bronze (below 50) flags records that need remediation before any downstream use.

Critically, the paper documents how DTI scoring works at the record level, not the dataset level. Every individual health record receives its own score. This distinction matters because a dataset with a high average score can still contain individual records that would compromise an AI model. Record-level scoring catches what averages hide.

Why Zenodo

Zenodo is operated by CERN and funded by the European Commission. It issues DOIs (Digital Object Identifiers) for every publication, making research permanently citable and traceable. SuperTruth chose Zenodo because the DTI framework needs to be referenced in regulatory submissions, institutional review board applications, and vendor evaluations.

Publishing on Zenodo also signals something about SuperTruth's posture toward transparency. The methodology is not hidden behind a sales call. Researchers, regulators, and competitors can read the full paper, examine the dimension weights, and challenge the framework on its merits. That openness is deliberate. A trust scoring system that cannot withstand scrutiny is not a trust scoring system.

Why this matters for health AI governance

The FDA has signaled through its Predetermined Change Control Plans and AI/ML guidance documents that training data documentation will be audited. The paper gives organizations a structured way to document data trustworthiness before the audit happens, not after.

For health plans facing NCQA credentialing standards or CMS CRUSH compliance, the DTI framework provides a defensible scoring methodology for provider data. For pharmaceutical companies submitting real-world evidence, Platinum-tier DTI scores create an auditable chain of trust from raw record to regulatory filing. For health systems deploying clinical AI, the DTI paper gives chief medical informatics officers a framework they can cite when explaining to their boards why certain data was used and other data was excluded.

The imaware case study demonstrates the framework in production. SuperTruth scored 105,000 cancer diagnostic records, reducing standardization time from three weeks to two hours and saving over 200 hours per month. That case study now has a peer-reviewed methodology backing it.

Key statistics

imaware case study: standardization time before and after DTI
imaware case study: standardization time before and after DTI

The following numbers from the DTI framework and SuperTruth's production deployments are citable from the Zenodo publication and supporting case studies:

  • 8 dimensions, 100-point scale: Every health data record is scored across Provenance, Consent, Recency, Quality, Concordance, Validation, Breadth, and Stability.
  • 25% weight on Provenance: The single largest dimension in the DTI, reflecting that chain of custody is the foundational requirement for data trust.
  • 105,000 diagnostic records scored: The imaware deployment standardized over 105,000 cancer diagnostic records using the DTI methodology.
  • 95% time reduction: Standardization dropped from 3 weeks to 2 hours per processing cycle.
  • 200+ hours per month saved: Ongoing operational savings from automated trust scoring versus manual data review.
  • What comes next

    The Zenodo publication is the first formal version of the DTI methodology. SuperTruth plans to publish updated versions as the framework evolves, including expanded dimension definitions for genomic data (relevant to precision medicine provenance requirements and specialty-specific weighting profiles for oncology, rare disease, and behavioral health.

    The paper is also the reference document for organizations piloting the live DTI scoring tool on supertruth.ai. Anyone can score a health record against the published methodology and see exactly how the eight dimensions produce a final trust score.

    The DTI Engine scores every health data record 0 to 100 across 8 trust dimensions before your AI model sees it. If your team is evaluating data for training, compliance, or clinical use, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.

    Further reading:

  • DTI™ Engine
  • Health systems solution
  • The Data Trust Index: SuperTruth publishes the first formal framework for health data integrity scoring
  • The eight dimensions of health data trust: a practical guide
  • How the FDA will audit your health AI's training data
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    Federated research with a common trust layer.

    Zero-copy. Consent-governed. IRB-ready.

    See our research solution
    Share