Patient-reported outcome measure (PROM) data trust: collection quality requirements
Photo by Mika Baumeister on Unsplash
insight

Patient-reported outcome measure (PROM) data trust: collection quality requirements

By Jason Alan Snyder·August 2, 2026

Patient-reported outcome measures generate some of the most clinically valuable data in healthcare, yet PROM collection quality failures render up to 40% of submissions unusable for regulatory or AI purposes. Meeting PROM data trust requirements means enforcing provenance, temporal precision, and validation standards at the point of collection, not after the fact.

The patient voice is the most valuable signal you are probably corrupting

Patient-reported outcome measures capture something no EHR field, lab value, or imaging study can: the patient's own assessment of their health status, functional capacity, and quality of life. FDA has accepted PROMs as primary or secondary endpoints in over 100 drug approvals since 2009. CMS ties reimbursement to PROM collection rates in joint replacement, cardiac surgery, and oncology care pathways.

But the collection process itself introduces systematic quality failures that most organizations never detect. A 2023 analysis in Quality of Life Research found that 38% of PROM submissions in multicenter trials contained at least one critical data quality issue: missing timestamps, ambiguous recall periods, proxy completion without documentation, or instrument version mismatches. These are not minor metadata problems. They are structural failures that make the resulting data untrustworthy for regulatory submission, AI model training, or clinical decision support.

PROM data trust starts at the moment a patient picks up a pen or opens an app. Everything downstream depends on what happens in that moment.

What makes PROM data different from other clinical data

Most clinical data originates from professionals operating within controlled environments. A lab tech runs an assay on calibrated equipment. A radiologist reads a scan on a PACS workstation with known specifications. A clinician enters a diagnosis code at a documented encounter.

PROM data originates from patients. They complete instruments at home, in waiting rooms, on phones, on paper, sometimes with a family member reading questions aloud. The environment is uncontrolled. The respondent's literacy level, emotional state, language proficiency, and understanding of the recall period all vary. No two PROM collection events share identical conditions.

This variability is not a flaw. It is the nature of the data source. But it means that PROM data trust requirements must account for factors that other clinical data types do not face. You cannot apply the same quality framework you use for lab results or billing claims and expect trustworthy PRO data to emerge.

The seven collection quality requirements for PROM data trust

After scoring thousands of health data records across clinical, claims, and patient-generated sources, we have identified seven requirements that determine whether PROM data achieves a trust score sufficient for clinical AI, regulatory submission, or research use.

1. Instrument provenance and version control

Every PROM submission must be traceable to a specific, validated instrument version. The PROMIS-29 v2.1 is not the PROMIS-29 v2.0. The PHQ-9 is not the PHQ-2. The EORTC QLQ-C30 version 3.0 contains different scoring algorithms than version 2.0.

Yet in practice, organizations routinely collect PROMs without recording which instrument version was used. A 2022 survey of 47 oncology practices found that 29% could not identify the exact version of the FACT-G they were administering. When instrument versions are ambiguous, longitudinal comparisons become invalid. Scores collected at baseline cannot be reliably compared to scores collected six months later if the instrument changed between measurements.

PROM data trust requires that every submission record the instrument name, version number, licensing status, and language/cultural adaptation identifier.

2. Temporal precision and recall period enforcement

Many PROM instruments specify a recall period. The SF-36 asks about the past four weeks. The EQ-5D asks about health "today." The FACT-B asks about the past seven days.

When patients complete instruments late, the recall period shifts. A patient who was supposed to complete the EQ-5D on day 30 post-surgery but completes it on day 45 is reporting on a different clinical window. If the system records only the submission date without flagging the protocol deviation, the data appears valid but measures the wrong thing.

PROM data trust requires millisecond-precision timestamps for both the intended collection window and the actual completion time. Any deviation beyond the instrument's specified recall period must be flagged, not silently accepted.

3. Mode of administration documentation

The same instrument administered by phone interview, paper form, electronic tablet, or smartphone app can produce systematically different scores. A 2021 meta-analysis in the Journal of Clinical Epidemiology found that mode effects for the PHQ-9 ranged from 0.5 to 2.3 points depending on the comparison, with electronic self-administration producing consistently lower scores than interviewer-administered versions.

PROM data trust requires that every submission document the mode of administration: paper, electronic (device type specified), phone interview, in-person interview, or video interview. Mixing modes without adjustment is a concordance failure that degrades any analysis built on the data.

4. Proxy completion identification

When a caregiver, family member, or clinical staff member completes a PROM on behalf of a patient, the data represents a different construct. Proxy-reported outcomes systematically differ from patient-reported outcomes, particularly for subjective domains like pain, fatigue, and emotional well-being. Studies in palliative care populations show proxy-patient agreement rates as low as 60% for symptom severity items.

Despite this, many collection systems have no mechanism to flag proxy completion. The data enters the same pipeline as self-reported data, with no indicator that the "patient" perspective came from someone else.

PROM data trust requires a mandatory proxy indicator field. When proxy completion is documented, downstream systems must treat that record differently in scoring, aggregation, and model training.

5. Completeness thresholds and missing data handling

Most validated PROM instruments specify scoring rules for missing items. The PROMIS system uses item response theory and can generate valid scores with some missing items. The SF-36 requires at least 50% of items within each scale. The FACT-G allows prorated scoring if more than 80% of items in a subscale are completed.

But collection systems often accept any submission regardless of completeness. A patient who answers 3 of 27 items generates a record that looks complete to the database but produces an invalid score. Worse, some systems impute missing values using mean substitution or last-observation-carried-forward without documenting the imputation method.

PROM data trust requires that every submission include an item-level completeness percentage, a validity flag based on the instrument's own scoring rules, and transparent documentation of any imputation applied.

6. Clinical context linkage

A PROMIS pain interference score of 65 means something very different for a patient two days after surgery than for a patient six months into remission. Without clinical context, PROM data floats in isolation, disconnected from the events that give it meaning.

PROM data trust requires linkage to at least three clinical context elements: the triggering clinical event (surgery date, treatment start, diagnosis), the collection protocol (which visit in the care pathway), and the patient's concurrent treatment status. Without these anchors, PROM data cannot support longitudinal outcome measurement or comparative effectiveness research.

7. Consent specificity for secondary use

Patients who agree to complete a PROM questionnaire as part of their clinical care have not necessarily consented to that data being used for AI model training, real-world evidence generation, or commercial research. The consent architecture around PROM data is often the weakest link in the trust chain.

A patient completing a knee replacement PROM at their surgeon's office may reasonably expect that data to inform their care. They may not expect it to appear in a pharmaceutical company's label expansion submission or an AI vendor's training dataset. PROM data trust requires granular consent documentation that specifies permitted uses: clinical care, quality improvement, research, AI training, commercial analytics. ConsentOS provides the five-tier architecture needed to manage this complexity.

Key statistics

PROM data quality failure rates by category
PROM data quality failure rates by category

  • 38% of PROM submissions in multicenter trials contain at least one critical data quality issue (Quality of Life Research, 2023)
  • 29% of oncology practices cannot identify the exact version of the FACT-G instrument they administer (2022 survey of 47 practices)
  • Proxy-patient agreement on symptom severity items drops to 60% in palliative care populations
  • Mode effects for the PHQ-9 range from 0.5 to 2.3 points depending on administration method (Journal of Clinical Epidemiology, 2021)
  • FDA has accepted PROMs as primary or secondary endpoints in over 100 drug approvals since 2009
  • How PROM failures cascade into AI and regulatory risk

    When PROM data with undetected quality failures enters downstream systems, the damage compounds. Consider three scenarios.

    First, a pharmaceutical company submits a label expansion application using real-world PROM data to demonstrate quality-of-life improvements. FDA reviewers discover that 22% of the PROM submissions lack mode-of-administration documentation and 15% have recall period deviations exceeding five days. The submission receives a Refuse to File letter. Twelve months of data collection are wasted. This pattern aligns with what FDA expects from real-world evidence submissions.

    Second, a health system trains a readmission prediction model that includes PROM scores as features. The training data contains proxy-completed PROMs that were never flagged, systematically biasing the model toward patients with engaged caregivers. The model underperforms for isolated patients, precisely the population at highest readmission risk. The problem described in how training data trust affects clinical AI applies directly.

    Third, a medical device company uses post-market PROM data to monitor device performance. Instrument version changes midway through the surveillance period go undocumented. Apparent improvements in patient-reported function are actually artifacts of the version change, not real clinical gains. The device continues to market with misleading performance claims. This is the kind of failure that post-market surveillance data trust standards are designed to prevent.

    The relationship between PROMs and patient-generated health data

    PROM data occupies a unique position between clinical data and patient-generated health data (PGHD). Unlike passively collected PGHD from wearables or apps, PROMs are structured instruments with validated psychometric properties. Unlike clinical data, they originate entirely from the patient.

    This hybrid nature creates specific trust challenges. PROM data inherits the provenance complexity of patient-generated health data while demanding the validation rigor of clinical trial endpoints. Organizations that treat PROMs as "just another patient survey" or "just another clinical data point" will fail at both.

    The trust threshold for PROM data depends on its intended use. Clinical care decisions require a minimum viable trust score. Regulatory submissions demand near-perfect provenance and validation. AI training requires documented consent and completeness. Each use case has a different floor, and the collection process must be designed to meet the highest anticipated use.

    Scoring PROM data trust with the DTI framework

    DTI dimension weights applied to PROM data trust scoring
    DTI dimension weights applied to PROM data trust scoring

    The Data Trust Index evaluates every health data record across eight dimensions. For PROM data, each dimension maps to specific collection requirements.

    Provenance (25% weight): instrument version, licensing status, language adaptation, and the chain from patient completion to database entry. Did the data pass through an intermediary system? Was it transcribed from paper? Each handoff point must be documented.

    Consent (20% weight): did the patient consent specifically to the uses this data will serve? Clinical care consent does not cover research or AI training.

    Recency (15% weight): when was this PROM collected relative to the clinical event it references? A six-month-old PROM score for a rapidly progressing disease is clinically stale. Recency is the most underrated dimension in data trust scoring, and PROM data makes this obvious.

    Quality (10% weight): item-level completeness, scoring validity per the instrument's rules, and absence of impossible response patterns (e.g., all items answered identically in under 30 seconds).

    Concordance (10% weight): does this PROM score align with other clinical indicators? A patient reporting zero pain while receiving escalating opioid doses warrants a concordance flag, not automatic rejection, but documentation.

    Validation (10% weight): has the data passed instrument-specific validation rules? Has mode-of-administration bias been assessed?

    Breadth (5% weight): does this PROM capture sufficient domains for its intended use? A single-domain instrument used to support a multi-domain outcome claim is a breadth failure.

    Stability (5% weight): are collection protocols consistent over time? Mid-study instrument changes, mode switches, or recall period modifications all degrade stability.

    What good PROM collection infrastructure looks like

    Organizations that achieve high PROM data trust scores share common infrastructure characteristics.

    They enforce instrument version locking at the protocol level, preventing ad hoc switches. They timestamp both the intended and actual collection windows. They require mode-of-administration selection before any items display. They include a mandatory proxy indicator with free-text relationship documentation.

    They apply real-time completeness validation, preventing submission of records that fall below the instrument's scoring threshold. They link every PROM to a clinical context record. And they capture consent at a granularity that distinguishes clinical care from research from AI training.

    None of this requires exotic technology. It requires intentional design of the collection workflow and enforcement of metadata requirements at the point of data creation.

    The cost of retrofitting trust into existing PROM data

    Organizations that have collected PROM data without these requirements face a painful reality. Retrofitting trust metadata onto existing records is expensive, often impossible, and always incomplete.

    You cannot reconstruct which instrument version was used if it was never recorded. You cannot determine whether a caregiver completed the form if no proxy field existed. You cannot recover the intended collection date if only the submission date was stored.

    The imaware case study illustrates the economics at a smaller scale: standardizing 105,000 diagnostic records took SuperTruth's DTI Engine just 2 hours, replacing a process that previously took 3 weeks. But diagnostic records have structured metadata that PROM data often lacks entirely. Retrofitting PROM trust is harder and more expensive per record.

    The lesson is straightforward. Build trust requirements into PROM collection systems from the start. Every day of collection without proper metadata creates records that may never reach a usable trust score.

    Where PROM data trust is heading

    FDA's 2023 guidance on patient-focused drug development explicitly calls for documentation of PROM collection methods, instrument selection rationale, and missing data handling in regulatory submissions. EMA's qualification process for novel PRO instruments now requires digital provenance documentation. CMS is expanding PROM collection requirements in value-based care programs, with quality payment adjustments tied to both collection rates and collection quality.

    The direction is clear: regulators and payers will increasingly treat PROM data trust as a prerequisite, not a nice-to-have. Organizations collecting PROMs today without meeting these requirements are building data assets that will not survive future scrutiny.

    The DTI Engine scores every health data record, including PROMs, from 0 to 100 across 8 trust dimensions before your AI model sees it. If your team is collecting patient-reported outcomes for regulatory submission, value-based care, or clinical AI and needs to know whether that data meets trust requirements, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.

    Further reading:

  • DTI™ Engine
  • Health systems solution
  • Patient-generated health data (PGHD) and the trust threshold for clinical use
  • Real-world evidence data quality: FDA requirements for RWE regulatory submission
  • Post-market surveillance data trust: what FDA expects from device performance data
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    DTI scores the record, not the patient.

    8 dimensions. 0 to 100. Travels with every record permanently.

    See the DTI Engine
    Share