SuperTruth publishes peer-reviewed research on the Data Trust Index at Zenodo — what the paper covers and why it matters for health AI governance
SuperTruth has published peer-reviewed research on the Data Trust Index (DTI) at Zenodo, making the formal methodology for scoring health data integrity publicly available. The paper defines the 8-dimension framework, weighted scoring model, and trust tier classification system that scores every health data record 0 to 100. This is the first open-access publication of a structured trust scoring methodology purpose-built for health AI governance.
SuperTruth has published the formal research paper describing the Data Trust Index (DTI) on Zenodo, the open-access repository operated by CERN. The paper lays out the complete methodology for scoring health data records on a 0 to 100 scale across eight weighted dimensions. It is now publicly available for citation, peer review, and institutional adoption.
This matters because health AI governance has operated without a standardized way to measure data trustworthiness. Regulatory bodies like the FDA are asking about training data provenance. Payers want to know if the records behind AI-driven claims decisions are current and validated. Research consortia need a shared language for data quality thresholds across institutions. The DTI paper provides that shared language.
What the paper covers
The Zenodo publication defines the Data Trust Index as a composite score derived from eight dimensions, each carrying a specific weight: Provenance (25%), Consent (20%), Recency (15%), Quality (10%), Concordance (10%), Validation (10%), Breadth (5%), and Stability (5%).
These weights were not arbitrary. The paper explains the rationale behind each allocation. Provenance carries the highest weight because a record with unknown origin cannot be trusted regardless of its other attributes. Consent follows because health data used without proper authorization creates both legal liability and ethical failure. Recency is third because stale data produces stale predictions, and temporal drift destroys AI model accuracy in healthcare faster than most teams realize.
The paper also introduces the trust tier classification system. Records scoring 90 to 100 earn Platinum status, suitable for FDA regulatory submission. Gold (75 to 89) qualifies for clinical AI training. Silver (50 to 74) supports operational analytics with caveats. Bronze (below 50) flags records that need remediation before any downstream use.
Critically, the paper documents how DTI scoring works at the record level, not the dataset level. Every individual health record receives its own score. This distinction matters because a dataset with a high average score can still contain individual records that would compromise an AI model. Record-level scoring catches what averages hide.
Why Zenodo
Zenodo is operated by CERN and funded by the European Commission. It issues DOIs (Digital Object Identifiers) for every publication, making research permanently citable and traceable. SuperTruth chose Zenodo because the DTI framework needs to be referenced in regulatory submissions, institutional review board applications, and vendor evaluations.
Publishing on Zenodo also signals something about SuperTruth's posture toward transparency. The methodology is not hidden behind a sales call. Researchers, regulators, and competitors can read the full paper, examine the dimension weights, and challenge the framework on its merits. That openness is deliberate. A trust scoring system that cannot withstand scrutiny is not a trust scoring system.
Why this matters for health AI governance
The FDA has signaled through its Predetermined Change Control Plans and AI/ML guidance documents that training data documentation will be audited. The paper gives organizations a structured way to document data trustworthiness before the audit happens, not after.
For health plans facing NCQA credentialing standards or CMS CRUSH compliance, the DTI framework provides a defensible scoring methodology for provider data. For pharmaceutical companies submitting real-world evidence, Platinum-tier DTI scores create an auditable chain of trust from raw record to regulatory filing. For health systems deploying clinical AI, the DTI paper gives chief medical informatics officers a framework they can cite when explaining to their boards why certain data was used and other data was excluded.
The imaware case study demonstrates the framework in production. SuperTruth scored 105,000 cancer diagnostic records, reducing standardization time from three weeks to two hours and saving over 200 hours per month. That case study now has a peer-reviewed methodology backing it.
Key statistics
The following numbers from the DTI framework and SuperTruth's production deployments are citable from the Zenodo publication and supporting case studies:
What comes next
The Zenodo publication is the first formal version of the DTI methodology. SuperTruth plans to publish updated versions as the framework evolves, including expanded dimension definitions for genomic data (relevant to precision medicine provenance requirements and specialty-specific weighting profiles for oncology, rare disease, and behavioral health.
The paper is also the reference document for organizations piloting the live DTI scoring tool on supertruth.ai. Anyone can score a health record against the published methodology and see exactly how the eight dimensions produce a final trust score.
The DTI Engine scores every health data record 0 to 100 across 8 trust dimensions before your AI model sees it. If your team is evaluating data for training, compliance, or clinical use, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.
Further reading:

Jason Alan Snyder
Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.
About SuperTruth · LinkedIn · Substack · jasonalansnyder.com
See it in practice
Federated research with a common trust layer.
Zero-copy. Consent-governed. IRB-ready.