How the Data Reservoir turns scored data into queryable intelligence
Photo by Anton Savinov on Unsplash
insight

How the Data Reservoir turns scored data into queryable intelligence

By Jason Alan Snyder·May 3, 2026

Most health data platforms store records. Few make those records queryable by trust score, consent status, and provenance chain simultaneously. The Data Reservoir is the layer that converts DTI-scored health data into queryable intelligence, letting researchers, health systems, and pharma teams run queries that return not just answers but auditable confidence levels.

A data lake stores everything. A data warehouse structures some of it. Neither tells you whether the records inside are trustworthy enough to query for clinical, regulatory, or commercial decisions.

That gap is where most healthcare AI projects fail. Not at the model layer. At the data layer, before a single query runs.

The Data Reservoir exists to close that gap. It takes health data that has already been scored by the DTI Engine, organized by trust tier, and governed by ConsentOS, and makes it queryable with full provenance, consent status, and quality metadata attached to every result.

What a health data reservoir actually is

DTI scoring dimensions and their weight in every reservoir record
DTI scoring dimensions and their weight in every reservoir record

A dataset in a traditional data lake is a collection of structured or semi-structured files stored in a common repository. It may include EHR exports, claims feeds, lab results, wearable data, or SDOH records. The data lake holds all of it without discrimination. That is both its strength and its fatal weakness.

A health data reservoir is different. Every record inside has been scored across eight trust dimensions before it enters the queryable layer. Provenance (25%), Consent (20%), Recency (15%), Quality (10%), Concordance (10%), Validation (10%), Breadth (5%), and Stability (5%) are evaluated and attached as metadata. The reservoir does not just hold data. It holds scored data with queryable trust attributes.

This means a researcher can run a query that says: "Return all colorectal cancer screening records from the last 18 months with a DTI score above 80 and active Tier 3 consent for research use." A traditional data lake cannot answer that query. The Data Reservoir can.

When converting operational data into insights

The process of converting operational data into insights requires more than SQL and dashboards. It requires knowing which records are safe to use, which have degraded over time, and which carry consent that covers the intended use case.

When health systems convert operational EHR data into population health insights, they typically run aggregate queries against whatever is in the warehouse. They rarely check whether the underlying records have consistent provenance, whether consent covers the analytic use case, or whether the data has drifted since it was last validated. The result is insights built on an unknown foundation.

The Data Reservoir inverts this. Every query result includes the trust score of the records that produced it. If a population health query returns 10,000 records but 3,200 of them score below Bronze tier (DTI below 60), the query interface flags that. The analyst sees not just the answer but the confidence level of the answer.

The process where intelligent methods extract data patterns

Data mining, machine learning, and statistical analysis are the standard methods applied to extract patterns from datasets. But in healthcare, the pattern is only as reliable as the data underneath it.

The Data Reservoir applies the DTI score as a pre-filter before pattern extraction begins. This is not post-hoc quality checking. It is a structural constraint. If you set a DTI floor of 75 for a pharmacovigilance study, the reservoir excludes records below that threshold before your model touches them. The patterns you extract come from data that has already been verified for provenance, consent, recency, and concordance.

This is what separates a scored health data query from a standard database query. The trust metadata travels with the data through every stage of analysis.

What data analysis gets you from scored versus unscored data

Data analysis in healthcare typically uses descriptive, diagnostic, predictive, and prescriptive methods. All four depend on input quality. When the input is unscored, every output carries hidden risk.

With SuperTruth's imaware partnership, we processed 105,000 diagnostic records. Before scoring, the data took three weeks to standardize manually. After DTI scoring, the same process took two hours. That 95% time reduction was not just an efficiency gain. It was a trust gain. Every record that entered the queryable layer carried a verified score. The analysis that followed identified a customer segment driving 20% of revenue, a finding that was invisible in the unscored dataset.

The Data Reservoir made that discovery possible because it did not just store the records. It made them queryable by trust attributes that exposed patterns hidden by noise and inconsistency.

Key statistics

imaware data preparation: before and after DTI scoring
imaware data preparation: before and after DTI scoring

  • 105,000 diagnostic records scored and standardized through the DTI Engine in the imaware case study
  • 95% reduction in data preparation time: from 3 weeks to 2 hours
  • 200+ hours per month saved in ongoing data operations
  • 8 trust dimensions scored per record, with Provenance (25%) and Consent (20%) carrying the highest weight
  • DTI floor enforcement means queries can exclude records below any specified trust tier (Bronze 60, Silver 70, Gold 80, Platinum 90)
  • How the data trust query interface works in practice

    The Data Reservoir exposes a query interface that treats trust as a first-class parameter. This is not a filter added after the fact. Trust dimensions are indexed alongside clinical and demographic attributes.

    A pharma company running a real-world evidence study can query: "Return all lung cancer diagnosis records from participating health systems, DTI above 80, with Tier 4 consent for commercial research, updated within the last 90 days." That query runs against the reservoir and returns records that meet every criterion, with full audit trails.

    A health plan running a credentialing compliance check can query provider records by recency and concordance scores, the two dimensions that most frequently break NCQA audits.

    A clinical trial team can identify eligible patient populations using behavioral signals from VIOLET combined with DTI-scored clinical records, ensuring both the signal and the underlying data meet regulatory standards.

    Why data lakes and data warehouses cannot do this

    Data lakes optimize for storage cost and schema flexibility. Data warehouses optimize for structured query performance. Neither optimizes for trust.

    The top-ranking content on this topic describes data lakes as inexpensive, horizontally scalable repositories. That is accurate. But scalable storage of unverified data does not produce queryable intelligence. It produces queryable risk.

    The Data Reservoir sits downstream of the DTI Engine and ConsentOS. By the time data reaches the reservoir, it has been scored, consent-tiered, and provenance-verified. The reservoir is not a replacement for a data lake. It is the trust-scored, query-ready layer that makes the data lake useful for regulated healthcare decisions.

    Without this layer, organizations build AI models on data they cannot defend to the FDA, cannot audit for consent compliance, and cannot verify for temporal drift. The reservoir eliminates those risks at the query layer, before the data reaches a model.

    The DTI Engine scores every health data record 0-100 across 8 trust dimensions before your AI model sees it. If your team is evaluating health data for training, compliance, or clinical use, and you need a query interface that returns trust metadata alongside clinical results, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.

    Further reading:

  • DTI™ Engine
  • Health systems solution
  • The DTI score as a contract: what Platinum-grade data actually guarantees
  • Why AI models trained on unscored health data will fail in production
  • Glass Box vs Black Box: why health AI needs explainable data provenance
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0 to 100. Travels with every record permanently.

    See the DTI Engine
    Share
    How the Data Reservoir turns scored data into queryable intelligence | SuperTruth