Why recency is the most underrated dimension in health AI data scoring
Photo by Logan Voss on Unsplash
insight

Why recency is the most underrated dimension in health AI data scoring

By Jason Alan Snyder·May 7, 2026

Recency carries 15% of the Data Trust Index score, yet most health AI pipelines treat timestamps as metadata rather than a trust signal. Stale data silently degrades model accuracy, clinical decision support, and regulatory defensibility. Health data recency scoring is the single fastest way to separate actionable intelligence from archived noise.

A lab result from 2019 and a lab result from last Tuesday can contain identical values. Only one of them should inform a clinical decision today. Yet most health AI pipelines treat both records the same way, because they score for completeness and format but not for time.

This is the recency problem. It is quiet, structural, and responsible for more downstream model failures than most teams realize.

What recency actually measures

DTI dimension weights: where recency ranks
DTI dimension weights: where recency ranks

Recency is not a timestamp. It is a trust dimension. Within the Data Trust Index, recency accounts for 15% of the total score, making it the third-highest weighted dimension behind provenance (25%) and consent (20%). It answers a specific question: how close is this data point to the clinical reality it claims to represent?

A hemoglobin A1c from six months ago tells you something. The same A1c from three years ago tells you almost nothing about a patient's current glycemic control. But if your AI model ingests both records without temporal weighting, it treats them as equivalent evidence. That is not a data quality issue. It is a data trust issue.

For a full breakdown of all eight dimensions, see The eight dimensions of health data trust: a practical guide.

Why existing pipelines ignore recency

Most health data platforms focus on structure. They validate FHIR formatting, check for null fields, confirm that a record arrived. These are necessary but insufficient checks. Structure tells you whether a record is readable. Recency tells you whether it is relevant.

The reason recency gets ignored is architectural. EHR systems store records chronologically but query them flatly. When a downstream application pulls "most recent labs," it often means the most recent record in the system, not the most recent clinical event. Transfer delays, batch uploads, and reconciliation gaps mean the timestamp on a record can lag the actual clinical moment by days, weeks, or months.

This is the same structural problem that makes temporal drift destroy AI model accuracy in healthcare. The data looks current. It is not.

How stale data breaks real workflows

Consider provider directory data for health plans. CMS CRUSH compliance requires accurate, current provider information. A provider's address, network status, or specialty can change quarterly. If the data feeding a directory model was last verified 14 months ago, every downstream decision built on that record carries inherited risk.

The same applies in oncology. A patient's tumor marker panel from pre-treatment is clinically distinct from a panel drawn during cycle 3 of chemotherapy. Models that ingest both without recency weighting will generate risk scores that blend historical baselines with active treatment data. The output looks precise. It is not.

In the imaware case study, SuperTruth standardized 105,000 diagnostic records and reduced processing time from 3 weeks to 2 hours. A critical part of that work involved temporal stratification: separating records by clinical relevance windows, not just by date received.

Key statistics

imaware processing time: before and after DTI scoring
imaware processing time: before and after DTI scoring

Recency carries 15% of the total DTI score, the third-highest weight across all eight dimensions.

SuperTruth's processing of imaware's 105,000 records achieved a 95% time reduction, from 504 hours to approximately 2 hours per cycle.

Provider directory data older than 90 days fails CMS CRUSH accuracy requirements, yet industry audits routinely find 30-40% of records exceed that window.

The DTI Engine applies recency decay curves calibrated by data type: vital signs decay faster than surgical history, and genomic data decays slower than medication lists.

Models trained on data with a mean recency lag of 18+ months show measurable accuracy degradation in longitudinal outcome prediction.

Recency is not binary

The temptation is to treat recency as a cutoff. Data younger than X days is fresh; everything else is stale. This is too crude. Different data types have different half-lives.

A diagnosis of Type 1 diabetes does not expire. A fasting glucose reading is clinically irrelevant within weeks if the patient's medication regimen changed. A social determinant of health indicator, like housing status, can shift overnight or remain stable for years. The DTI Engine handles this through type-specific recency decay curves rather than a single threshold.

This approach matters especially for rural health data, which is stale by design. When clinical touchpoints are infrequent, recency scoring must account for expected data sparsity rather than penalizing it.

What health AI recency scoring looks like in practice

The DTI Engine evaluates recency at the field level, not the record level. A single patient record might contain a current address (high recency), a medication list updated 8 months ago (moderate recency), and an allergy notation from 2016 (low recency but potentially stable). Each field receives its own temporal trust contribution.

This field-level granularity means that a record with a DTI score of 78 communicates something specific about temporal reliability. Downstream consumers, whether AI models, clinical decision tools, or compliance auditors, can filter by recency floor. A research team running a retrospective study might accept records with lower recency scores. A real-time clinical alert system should not.

The scored output feeds directly into the Data Reservoir, where temporal health data trust becomes a queryable attribute rather than an assumption.

The regulatory tailwind

The FDA's emerging framework for AI/ML-based Software as a Medical Device explicitly references training data quality, and temporal representativeness is a named concern. If your training data overrepresents a clinical era that predates current treatment protocols, your model inherits that bias. Recency scoring creates an auditable record proving that temporal composition was evaluated, not assumed.

For teams preparing for FDA audits of training data, recency scoring is one of the fastest paths to defensible documentation.

The fix is scoring, not filtering

Deleting old data is wasteful. Historical records have genuine value for longitudinal analysis, trend detection, and baseline comparison. The problem is not that old data exists. The problem is that old data circulates without a trust signal indicating its temporal position.

Scoring solves this. Every record stays in the pipeline. Every record carries a recency score. Every downstream consumer decides its own threshold. That is what Platinum-grade data actually guarantees: not perfection, but transparency about exactly what you are working with.

The DTI Engine scores every health data record 0-100 across 8 trust dimensions before your AI model sees it. Recency is weighted at 15% because temporal relevance determines whether a clinically accurate record is still clinically useful. If your team is evaluating data for training, compliance, or clinical use, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.

Further reading:

  • DTI™ Engine
  • Health systems solution
  • How temporal drift destroys AI model accuracy in healthcare
  • The eight dimensions of health data trust: a practical guide
  • Rural health data is stale by design. SDOH scoring is the fix.
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0–100. Travels with every record permanently.

    See the DTI Engine
    Share