The eight dimensions of health data trust: a practical guide
Photo by Steve A Johnson on Unsplash
insight

The eight dimensions of health data trust: a practical guide

By Jason Alan Snyder·April 19, 2026

The Data Trust Index scores every health data record from 0 to 100 across eight weighted dimensions. This practical guide breaks down each dimension, explains why the weights are set the way they are, and shows how the DTI engine converts raw health data into a trust-scored asset ready for AI, regulatory submission, and clinical use.

Health data does not fail because it is missing. It fails because nobody scores it before using it. The Data Trust Index (DTI) assigns every health data record a score from 0 to 100 across eight specific dimensions. Each dimension carries a defined weight. Together, they produce a single trust score that functions like a FICO score for health data.

This guide explains every dimension, why each weight exists, and how the DTI engine applies them in practice.

What are the 8 dimensions of health data trust?

DTI dimension weights: how each dimension contributes to the total trust score
DTI dimension weights: how each dimension contributes to the total trust score

The DTI framework evaluates health data records across these eight dimensions, listed by weight:

  • Provenance (25%) — Where did this record originate? Was it generated by a certified lab, a hospital EHR, a patient portal, or a third-party aggregator? Provenance is the single most important trust signal because a record with unknown origins cannot be validated downstream. The DTI engine traces each data point to its source system, timestamp, and generating entity.
  • Consent (20%) — Did the patient explicitly authorize this data's use for the stated purpose? Consent is not a checkbox. It is a record-level attribute that must specify scope, duration, and permitted use cases. A record can have perfect clinical quality and still score zero if consent governance is absent. This is the dimension most health AI companies overlook, and it is the one that creates the largest regulatory exposure.
  • Recency (15%) — How old is this record? A lab result from 2019 tells a different story than one from last week. The DTI engine applies decay curves that vary by data type. Vital signs decay faster than genomic data. A diagnosis code from five years ago may still be relevant; a medication list from the same period almost certainly is not.
  • Quality (10%) — Is the data complete, correctly formatted, and free of contradictions within the record itself? Quality covers structural integrity: missing fields, impossible values (a blood pressure of 900/400), mismatched units, and truncated entries. The DTI engine runs 40+ quality checks per record type.
  • Concordance (10%) — Does this record agree with other records about the same patient? When three sources list three different primary care physicians, concordance drops. When a lab result, a clinical note, and a pharmacy claim all align on a diagnosis, concordance rises. This dimension measures cross-source agreement.
  • Validation (10%) — Has the data been verified against an authoritative source? For provider data, this means primary source verification against state licensing boards, DEA registries, and NPPES. For clinical data, it means confirmation against reference labs or institutional records. Validation separates self-reported data from confirmed data.
  • Breadth (5%) — How many data types and sources contribute to this patient's or provider's record? A record built from a single EHR scores lower on breadth than one that incorporates claims data, lab results, pharmacy records, and social determinants. Breadth does not override quality, which is why it carries only 5% weight. But it signals how complete the picture is.
  • Stability (5%) — How consistent has this record been over time? A provider whose address changes four times in six months scores lower on stability than one whose information has remained constant for three years. For patient data, stability tracks whether key identifiers and clinical markers remain coherent across updates.
  • Why these weights?

    Provenance and consent together account for 45% of the total score. This is deliberate. A health data record that cannot prove where it came from or whether the patient authorized its use is unusable for FDA submissions, AI model training, or payer decisions, regardless of how clinically accurate it might be.

    Recency carries 15% because stale data is the default state of health records. Our work with imaware on 105,000 diagnostic records showed that records more than 90 days old had measurably lower concordance with current clinical status.

    Quality, concordance, and validation each carry 10% because they address different failure modes. A record can be high quality (well-formatted) but low concordance (disagrees with other sources). It can be validated (confirmed against a primary source) but low quality (missing fields). Separating these dimensions prevents a false sense of trust.

    Breadth and stability carry 5% each. They are meaningful but secondary. A narrow, stable, well-sourced record is more trustworthy than a broad, unstable one.

    Key statistics

    imaware data standardization: before and after DTI engine deployment
    imaware data standardization: before and after DTI engine deployment

    These numbers come from SuperTruth's production deployments and the DTI framework itself:

  • 8 dimensions, weighted to 100: Provenance (25%), Consent (20%), Recency (15%), Quality (10%), Concordance (10%), Validation (10%), Breadth (5%), Stability (5%).
  • 105,000 diagnostic records scored and standardized in the imaware partnership, reducing processing time from 3 weeks to 2 hours.
  • 95% time reduction in data standardization, saving over 200 hours per month for imaware's operations team.
  • 45% of the total DTI score is determined by provenance and consent alone, reflecting the regulatory reality that source and authorization outweigh all other factors.
  • 40+ quality checks per record type run automatically by the DTI engine before a score is assigned.
  • How the DTI engine applies these dimensions

    The DTI engine is not a survey tool. It ingests raw health data records, whether HL7, FHIR, CSV, or flat files, and applies all eight dimensions computationally. Each record receives a composite score and a per-dimension breakdown.

    Records scoring 80 or above are classified as Platinum-grade, suitable for FDA regulatory submission and AI model training. Records between 60 and 79 are Gold-grade, usable for payer operations and population health analytics. Records below 40 are flagged for remediation or exclusion.

    This tiering system allows organizations to make decisions based on measured trust, not assumptions.

    How this differs from the "8 dimensions of wellness"

    If you searched for "8 dimensions of health," you likely encountered the Swarbrick model of wellness: emotional, spiritual, intellectual, physical, social, environmental, financial, and occupational. That framework describes personal well-being.

    The DTI's eight dimensions describe something different entirely: the trustworthiness of the data that health systems, payers, pharma companies, and AI models use to make decisions about those same individuals. The wellness framework asks, "How healthy is this person?" The DTI framework asks, "How trustworthy is the data we are using to answer that question?"

    Similarly, the eight data quality dimensions referenced in data governance literature (accuracy, completeness, consistency, timeliness, validity, uniqueness, integrity, and fitness) overlap partially with the DTI. But those frameworks were built for enterprise data generally. The DTI was built for health data specifically, with consent and provenance weighted to reflect HIPAA, FDA, and CMS requirements that do not apply to retail or financial datasets.

    What this means for your organization

    If you are a health plan preparing for CMS CRUSH compliance, provenance and validation scores tell you whether your provider directory data will survive an audit. If you are a pharma company training AI models on real-world evidence, consent and recency scores determine whether that training data creates regulatory risk. If you are a health system onboarding new data sources, breadth and concordance scores reveal whether the new data adds signal or noise.

    Every use case maps back to these eight dimensions. The DTI engine makes them measurable.

    To see how the DTI engine scores your data across all eight dimensions, contact Louis Simeonidis, SVP Commercial Operations, at louis@supertruth.ai or (215) 918-4140.

    Further reading:

  • DTI Engine
  • Health systems solution
  • The Data Trust Index: SuperTruth publishes the first formal framework for health data integrity scoring
  • Why AI models trained on unscored health data will fail in production
  • What makes health data Platinum-grade for FDA regulatory submission
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    The FICO score for health data.

    8 dimensions. 0–100. Travels with every record permanently.

    See the DTI Engine
    Share