Health literacy and data quality: how patient-entered data degrades over time
Photo by Justin Schwartfigure on Unsplash
insight

Health literacy and data quality: how patient-entered data degrades over time

By Jason Alan Snyder·August 3, 2026

Patient-entered health data loses clinical reliability at a measurable rate. Within 90 days of collection, self-reported medication lists, symptom logs, and health histories show concordance drops of 20-50% against verified clinical records. Understanding this decay curve is essential before any AI model trains on patient-reported inputs.

A patient fills out a health history form in a waiting room. Three months later, that same data sits in an EHR, feeding a risk model. Nobody asks whether the patient understood the questions. Nobody checks whether the answers were accurate at the time they were entered. And nobody measures how much those answers have drifted from clinical reality since.

This is the patient self-reported data decay problem. It is measurable, predictable, and almost universally ignored.

The decay curve nobody tracks

Patient-reported medication list concordance decay over time
Patient-reported medication list concordance decay over time

Patient-entered data does not degrade randomly. It follows a pattern. A 2023 study in the Journal of the American Medical Informatics Association found that self-reported medication lists matched pharmacy records only 60% of the time at the point of entry. By 90 days, concordance dropped to 42%. By six months, fewer than one in three self-reported medication entries aligned with what the patient was actually taking.

This is not a documentation problem. It is a compounding accuracy problem. Every downstream system that consumes this data inherits the error without knowing it.

Symptom severity scores show a similar trajectory. Patient-reported outcome measures (PROMs) collected during a clinic visit reflect a snapshot. But patients with limited health literacy often conflate symptom duration with symptom severity, or report what they think the clinician wants to hear rather than what they actually experience. The data looks clean in the system. It is structurally compromised from the moment it enters.

How does health literacy affect patients?

Health literacy is the capacity to obtain, process, and understand basic health information needed to make appropriate health decisions. The National Assessment of Adult Literacy found that 36% of U.S. adults have basic or below-basic health literacy. That is roughly 87 million people.

People with limited health literacy are less likely to understand prescription drug labels, follow treatment instructions, or accurately describe their symptoms using medical terminology. They are more likely to be hospitalized, less likely to use preventive services, and more likely to report inaccurate health histories.

The clinical consequences are well documented. A systematic review published in BMC Public Health found that low health literacy is independently associated with higher mortality rates, more emergency department visits, and lower use of mammography screening. But the data quality consequences receive almost no attention.

When a patient with limited health literacy enters data into a portal, a survey, or an intake form, the resulting record carries the same structural weight as data entered by a clinician. No flag. No literacy adjustment. No trust differential. The system treats both inputs as equivalent, and every model trained on that data inherits the assumption.

What four factors affect people's level of health literacy?

Four primary factors drive variation in health literacy across populations.

Education level. Adults without a high school diploma are five times more likely to have below-basic health literacy than college graduates. But education alone does not determine health literacy; a PhD in physics does not prepare someone to interpret a hemoglobin A1c result.

Age. Adults over 65 have the lowest average health literacy scores of any age group. Cognitive decline, sensory impairment, and unfamiliarity with digital interfaces all contribute. This population also generates some of the highest volumes of patient-entered data through Medicare wellness visits and chronic disease management programs.

Language and cultural background. Non-native English speakers score significantly lower on health literacy assessments, even when controlling for education. Cultural differences in how symptoms are described, how pain is rated, and how medication adherence is reported introduce systematic variance that standard intake forms do not capture.

Chronic disease burden. Paradoxically, patients managing multiple chronic conditions often have lower functional health literacy despite frequent healthcare contact. Information overload, medication complexity, and conflicting instructions from multiple providers erode comprehension over time rather than building it.

Each of these factors introduces a distinct pattern of data quality degradation. None of them are captured in standard EHR metadata.

What is an example of poor data quality in healthcare?

Consider a 68-year-old patient with type 2 diabetes, hypertension, and moderate hearing loss who completes a pre-visit questionnaire through a patient portal. The form asks: "List all current medications including dosages."

The patient enters "metformin" but omits the dosage because he does not remember whether he takes 500mg or 1000mg. He lists "blood pressure pill" instead of lisinopril because he has never learned the generic name. He omits the aspirin his cardiologist recommended because he considers it a supplement, not a medication. He does not mention that he stopped taking his statin three weeks ago because of muscle pain.

The EHR now contains a medication list with one correct entry, one unidentifiable entry, one omission, and one phantom medication the patient no longer takes. This record will be used for medication reconciliation, drug interaction checking, and potentially as training data for clinical AI.

This is not a hypothetical. A 2022 analysis in the Annals of Internal Medicine found that 52% of patient-reported medication lists contained at least one clinically significant discrepancy when compared against pharmacy dispensing records. The error rate was 71% among patients with limited health literacy.

Key statistics

DTI scoring dimensions and their weights
DTI scoring dimensions and their weights

  • 36% of U.S. adults (approximately 87 million people) have basic or below-basic health literacy, according to the National Assessment of Adult Literacy.
  • Patient-reported medication lists match pharmacy records only 60% of the time at entry, dropping to 42% concordance at 90 days.
  • 52% of self-reported medication lists contain at least one clinically significant discrepancy against dispensing records.
  • 71% error rate in medication self-reporting among patients with limited health literacy.
  • SuperTruth's DTI Engine reduced data standardization time from 3 weeks to 2 hours across 105,000 diagnostic records in its imaware partnership, demonstrating that trust scoring at scale is operationally feasible.
  • What are the 5 C's of medical record entries?

    The 5 C's provide a framework for evaluating the quality of any medical record entry, whether clinician-authored or patient-entered.

    Clear. The entry must be unambiguous. "Blood pressure pill" is not clear. "Lisinopril 10mg daily" is.

    Concise. Entries should contain relevant information without unnecessary narrative. Patient-entered data often fails in both directions: too sparse (single-word answers) or too verbose (paragraph-length free-text responses that bury clinical details).

    Complete. Every relevant data element should be present. Omissions are the most common failure mode in patient-entered data. Patients do not know what is clinically relevant, so they leave out information that seems obvious to them but is critical to a provider.

    Correct. The information must be factually accurate at the time of entry. Self-reported diagnoses are frequently imprecise. A patient may report "heart disease" when the actual diagnosis is mitral valve prolapse, or "arthritis" when the diagnosis is gout.

    Chronological. Entries must reflect accurate timing. Patients routinely misreport when symptoms began, when medications were started, or when a prior procedure occurred. A 2021 study found that patient-reported surgical dates were inaccurate by more than 12 months in 23% of cases.

    Patient-entered data routinely fails on three or more of these five dimensions. Yet it enters the same database, feeds the same models, and carries the same implicit trust score as clinician-verified entries.

    The decay mechanism: why patient data gets worse over time

    Patient self-reported data decay is not simply a matter of medical records going stale. All health data loses recency. The specific problem with patient-entered data is that it was often partially inaccurate at collection, and the inaccuracies compound as clinical reality changes.

    Three mechanisms drive the decay.

    Medication changes without record updates. Patients change medications, adjust doses, and stop treatments without updating their records. Unlike pharmacy dispensing data, which reflects actual fills, patient-reported medication lists reflect memory and intention. The gap widens with every unreported change.

    Symptom reinterpretation. Patients reinterpret past symptoms in light of new information. A patient who learns they have a herniated disc may retroactively attribute all back pain to the disc, overwriting the original symptom reports with a narrative that aligns with the diagnosis. This is called recall bias, and it is well documented in clinical research but rarely accounted for in EHR data quality frameworks.

    Progressive health literacy drift. As patients learn more about their conditions, their understanding of medical terminology shifts. A patient who initially reported "chest tightness" may begin reporting "angina" after reading about their condition online. The symptom description changes without any change in the underlying symptom. NLP systems processing these records may interpret the terminology shift as a clinical change.

    We wrote extensively about the recency dimension of data trust and why it remains the most underrated factor in health AI data scoring. The decay problem with patient-entered data is that recency alone does not capture the degradation pattern. A record can be recent and still be wrong if the patient lacked the health literacy to enter accurate data in the first place. This is why the DTI framework scores across eight dimensions, not just one.

    The AI training problem: garbage in, clinical risk out

    When patient-entered data feeds AI training pipelines, the decay problem becomes a safety problem.

    Consider a readmission prediction model trained on data that includes patient-reported symptom surveys collected at discharge. If 36% of the population has limited health literacy, and those patients systematically underreport or misreport symptoms, the model learns that certain symptom patterns predict readmission. But it is actually learning which patients can accurately describe their symptoms, not which patients are clinically at risk.

    This is not a theoretical concern. Research on readmission prediction model bias has shown that training data trust directly affects whether clinical AI produces equitable outputs. Patient-entered data with unscored health literacy introduces a confounding variable that no amount of model tuning can correct.

    The same problem applies to patient-generated health data (PGHD) from wearables and apps, where the patient's ability to configure devices, interpret readings, and report context varies enormously by health literacy level.

    What clinical teams actually need: trust scoring at the point of entry

    The solution is not to exclude patient-entered data. Patients are the only source for many critical data elements: symptom onset, functional status, quality of life, treatment preferences, and social determinants of health. Excluding their input would create different and arguably worse data gaps.

    The solution is to score patient-entered data differently than clinician-verified data, with explicit adjustments for the factors that drive decay.

    A trust scoring framework needs to account for:

  • Provenance. Was this data entered by the patient, a caregiver, a clinician, or auto-populated from a connected device? Each source carries a different baseline accuracy profile.
  • Recency. How old is the entry? Patient-reported data degrades faster than clinician-documented data, so recency penalties should be steeper.
  • Concordance. Does the patient-entered data align with other verified sources? A self-reported medication list that matches pharmacy claims data is more trustworthy than one that does not.
  • Validation. Has any clinician reviewed and confirmed the patient-entered data? Unvalidated patient entries should carry a lower trust score than entries that have been clinically reconciled.
  • These are four of the eight dimensions in SuperTruth's Data Trust Index. The DTI does not reject patient-entered data. It scores it, so downstream systems can make informed decisions about how much weight to assign.

    When should nurses assess health literacy?

    A question circulating in clinical education asks when nurses should assess health literacy. The answer, based on current evidence, is at every point of data collection. Not just at intake. Not just during patient education. Every time a patient enters data that will persist in their record.

    This matters for data quality because the moment of collection determines the ceiling of accuracy. A nurse who identifies limited health literacy at intake can assist with form completion, verify medication names against pharmacy records, and flag entries that need clinician review. A nurse who does not assess literacy allows unverified, potentially inaccurate data to enter the system with no quality marker.

    Recent discussions in the nursing community, including the MedPage Today conversation about federal recognition of nursing as a profession, highlight the gap between what clinical staff are expected to do and what they are resourced to do. Health literacy assessment at every data collection point requires time, training, and workflow integration. Most health systems have none of these in place.

    The cost of ignoring decay

    Patient self-reported data decay is not an abstract data governance concern. It has measurable downstream costs.

    Medication reconciliation errors driven by inaccurate patient-reported lists contribute to an estimated 1.5 million adverse drug events per year in the United States, according to the Institute of Medicine. Each preventable adverse drug event costs an average of $8,750 in additional hospital charges.

    AI models trained on decayed patient data produce outputs that clinicians learn to distrust, leading to alert fatigue and reduced adoption of clinical decision support tools. The cost is not just financial. It is the erosion of clinician trust in AI systems that could genuinely improve care if they were trained on scored, validated data.

    The fragmentation problem multiplies the decay effect. When patient-entered data from one system is shared across a health information exchange without trust metadata, every receiving system inherits the original accuracy problems. We have written about the $3.5 trillion cost of bad health data and how fragmentation amplifies quality failures. Patient-entered data is where many of those failures begin.

    Fixing the problem before the model sees the data

    The pattern across the healthcare data landscape is consistent: data quality problems are cheapest to fix at the point of collection and most expensive to fix after they have propagated through training pipelines, model outputs, and clinical decisions.

    For patient-entered data, this means three things.

    First, score every patient-entered record at ingestion. Do not wait for a downstream audit to discover that a self-reported medication list has zero concordance with pharmacy claims. Flag it immediately.

    Second, apply literacy-adjusted decay curves. Patient-entered data should not carry the same recency profile as clinician-documented data. A self-reported symptom log from six months ago is not equivalent to a clinician-documented assessment from six months ago. The trust score should reflect the difference.

    Third, create feedback loops. When a clinician reviews and corrects patient-entered data during a visit, that correction should update the trust score of the original entry. Most EHR systems overwrite the original data. A trust-scored system preserves the provenance chain: original patient entry, clinician correction, updated trust score.

    The DTI Engine scores every health data record 0-100 across 8 trust dimensions before your AI model sees it. If your team is building on patient-reported data for clinical AI, population health analytics, or regulatory submission, and you need to know which records are trustworthy before they reach a model, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.

    Further reading:

  • DTI™ Engine
  • Health systems solution
  • Patient-reported outcome measure (PROM) data trust: collection quality requirements
  • Patient-generated health data (PGHD) and the trust threshold for clinical use
  • Why recency is the most underrated dimension in health AI data scoring
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0 to 100. Travels with every record permanently.

    See the DTI Engine
    Share