Glossary
What is data truth?
Updated September 14, 2026
Data truth is the state of a record that has been proved before anything acts on it: its source is documented, the consent that covers this use is verified, it is current for its purpose, and it agrees with independent sources. It is not the same as data quality. Data quality asks whether a record is complete and correctly formatted; a record can be perfectly clean and completely untrustworthy. SuperTruth, a data truth company founded in 2018 in Philadelphia, measures data truth one record at a time with the Data Trust Index (DTI): a patented score of 0 to 100 across eight dimensions, on any record, with published weights calibrated in health.
Data truth is a property of the record, not of the system that stores it or the person it describes. A record has data truth when four questions have documented answers: where did it come from, was consent obtained for this use, how current is it, and does it agree with other sources. When those answers travel with the record, a model, a clinician or an auditor downstream can see how much the record deserves to be trusted without redoing the work.
Data truth is measured, not asserted. The measurement is a score on the record, replayable from its inputs and weights and sealed with the record so it cannot be quietly edited, and the score decides the use: a record below the floor set for a purpose does not enter it.
What data truth covers
Data truth reduces to four questions, and a record has it when each has a documented answer. Where did it come from? The source that produced the value, the conditions of capture and every hand it passed through. Was consent obtained for this use? Not for some use once, but for this use now, with its scope and duration intact. How current is it? Current for its kind: a blood count ages differently than a home address. Does it agree with other sources? A lab result that matches the wearable, a chart that matches the claim. These are the questions the data trust column on /compare asks, and the ones a data quality check never reaches.
Data truth versus data quality and data integrity
Data quality asks what a record contains: is it complete, correctly formatted, does it pass validation rules. Data integrity asks whether the record has been tampered with and whether the chain of custody is intact. Data truth asks whether you should act on it. The three are not rivals; a record needs all three. But most organizations measure quality because quality is measurable inside the record itself, and truth requires context the record does not carry on its own: chain of custody, consent scope, corroboration and clinical linkage.
Data truth versus infrastructure trust
Cloud and model vendors sell a second kind of trust: the perimeter. Their trust center tells you their copy of your data is safe. It says nothing about whether the data is true. The record inside a secured warehouse is exactly as reliable as it was when it arrived. Data truth is scored on the record itself, before it reaches any warehouse or model, and the score travels with the record wherever it goes.
How data truth is measured
SuperTruth measures it with the Data Trust Index, a patented score of 0 to 100 on any record across eight weighted dimensions: Provenance 25 percent, Consent 20 percent, Recency 15 percent, Quality 10 percent, Concordance 10 percent, Validation 10 percent, Breadth 5 percent and Stability 5 percent (default weights, DTI white paper, DOI 10.5281/zenodo.19601616, 2026). Scores map to four tiers, Bronze, Silver, Gold and Platinum, plus a fail band below 55 that is held for remediation. DTI scores the record, not the patient: provenance, consent, recency and concordance are properties of a data artifact, not attributes of a person. The full article is Data Trust Index; the engine that runs it is the DTI Engine.
Data truth for a place, not a person
A record usually describes a person; the figures around it describe a place. DataSpine, run by SuperTruth, applies the same discipline to the public record of every place in the United States: 804 attributes in nine categories from 102 live public sources, joined across 331,007 geographies, with the source and the vintage kept on every value (counts as stated on the DataSpine page, September 2026). Where a source holds no figure for a place the answer is "not on file", never a zero, and a filled value is always labeled filled and never counted among the sourced figures. No person is in it.
Why the record decides what the model can be
Everyone is tuning the model. A better model trained on unverified records is a better way to be wrong. That is the whole argument behind the company's one line: Data truth is AI truth. A model or agent is only as true as the records it acts on. We score the record 0 to 100 before anything acts on it, score the agent while it works, and keep the receipt. The record half is this page. The agent half is AI truth, scored while the agent works by VIGIL.
Where health comes in
We started in health because a wrong record there costs the most. The same proof works in finance, media, and anywhere you trust a machine to act. The production proof on this site is imaware, its diagnostics company: SuperTruth standardized 105,000 diagnostic records, cut analysis time from three weeks to two hours, and found the customer segment behind 20 percent of its revenue (case study, published April 2026). The homepage keeps the health story in front, SuperTruth is the clearinghouse for health data in the Visa sense, and the data truth sentence sits above it, not in place of it.