Post-market surveillance data trust: what FDA expects from device performance data
FDA expects post-market surveillance data to meet specific standards for completeness, timeliness, and traceability. Most device manufacturers collect the data but fail to prove it is trustworthy. Here is what the agency actually requires, what the ISO standards add, and why AI-enabled devices face a fundamentally different surveillance burden.
The FDA has cleared or approved over 1,000 AI/ML-enabled medical devices. Every single one of them carries a post-market surveillance obligation. But the gap between collecting device performance data and proving that data is trustworthy enough for regulatory review is where most manufacturers fail.
Post-market surveillance is not optional. It is not a best practice. It is a legal requirement under federal law. And for AI-enabled devices, the data burden is heavier than for any device class that came before, because the device itself changes over time.
What is the FDA guidance for post-market surveillance?
FDA's post-market surveillance framework operates across several regulatory mechanisms. The broadest is the Medical Device Reporting (MDR) requirement under 21 CFR Part 803, which mandates that manufacturers, importers, and device user facilities report deaths, serious injuries, and malfunctions to the FDA.
Beyond MDR, the FDA issues post-market surveillance orders under Section 522 of the Federal Food, Drug, and Cosmetic Act. These orders require manufacturers to actively collect clinical data on specific devices after they reach the market. The FDA maintains a public database of all Section 522 orders, and as of 2024, hundreds of active surveillance studies are underway.
The FDA's 2023 guidance document, "Postmarket Surveillance Under Section 522 of the Federal Food, Drug, and Cosmetic Act," specifies what the agency expects from these studies. Key requirements include a surveillance plan that identifies the clinical questions being addressed, the study population, the data collection methods, and the analysis plan. The FDA also expects manufacturers to submit interim reports and a final study report.
For AI/ML-enabled devices specifically, the FDA's 2021 Action Plan for AI/ML-Based Software as a Medical Device (SaMD) introduced the concept of a "predetermined change control plan." This plan requires manufacturers to describe anticipated modifications to their algorithms and to monitor real-world performance against defined benchmarks. The Total Product Life Cycle (TPLC) approach means post-market data is not just about safety events; it is about ongoing algorithm performance validation.
What regulation drives FDA-ordered post-market surveillance activities?
Section 522 of the Federal Food, Drug, and Cosmetic Act (FD&C Act) is the primary regulation that authorizes the FDA to order post-market surveillance. The statute applies to Class II and Class III devices that meet specific criteria:
The FDA can order a manufacturer to conduct surveillance for a period of up to 36 months, though extensions are possible. The regulation is codified at 21 CFR Part 822.
For AI-enabled devices, FDA has increasingly relied on the 522 pathway alongside the newer Real-World Performance (RWP) monitoring expectations. The agency's January 2025 draft guidance on Marketing Submission Recommendations for a Predetermined Change Control Plan explicitly connects pre-market change plans to post-market data collection requirements. If your device learns, the FDA expects you to prove it learns correctly, and that proof lives in post-market surveillance data.
A recent MedPage Today opinion piece titled "Easing AI and Wearables Regulation Is a Risky Move" highlights the tension between regulatory streamlining and patient safety. The article describes a patient arriving with AI-generated health data on their phone, illustrating exactly why post-market surveillance data needs to be trustworthy: clinicians and patients are already acting on device outputs. When those outputs change because the algorithm updated, the surveillance data must capture whether performance held.
What is the ISO standard for post-market surveillance?
ISO 13485:2016, the quality management system standard for medical devices, requires post-market surveillance as part of its Section 8.2.1 (feedback) and Section 8.2.3 (monitoring and measurement of processes). But the more specific standard is ISO 14971:2019, which governs risk management for medical devices and requires manufacturers to collect post-production information to update risk analyses.
The most directly applicable standard is the EU's MDR 2017/745, which codifies post-market surveillance requirements in Articles 83 through 86. While this is a European regulation, its influence on global device manufacturers is substantial. The MDR requires a Post-Market Surveillance Plan (PMS Plan), Periodic Safety Update Reports (PSURs) for Class IIa and higher devices, and Post-Market Clinical Follow-up (PMCF) studies.
ISO/TR 20416:2020 provides additional guidance specifically on post-market surveillance for medical devices, offering a technical report that maps surveillance activities to the product lifecycle.
For AI-enabled devices, IMDRF (International Medical Device Regulators Forum) has published guidance on Software as a Medical Device that addresses ongoing monitoring. The convergence of ISO risk management requirements with FDA's TPLC approach creates a dual mandate: you must collect post-market data, and you must prove that data meets quality standards sufficient for risk analysis.
What is the post-market surveillance requirement?
The core requirement is deceptively simple: manufacturers must systematically collect, analyze, and report data about device performance after commercial distribution. The complexity lies in what "systematically" means in practice.
FDA expects post-market surveillance data to address five categories:
The requirement is not just to collect this data. The requirement is to collect data that the FDA can trust. And that is where most manufacturers have a problem.
The data trust gap in post-market surveillance
Most post-market surveillance programs collect data from multiple sources: electronic health records, claims databases, patient registries, complaint databases, and increasingly, real-world data from connected devices. Each source carries its own provenance challenges.
A complaint record entered by a call center agent six weeks after the event occurred has different trust characteristics than a real-time telemetry log from the device itself. A registry entry manually abstracted from a hospital chart has different accuracy characteristics than a structured data feed from an EHR.
The FDA does not treat all post-market data equally. The agency's 2023 guidance on "Use of Real-World Evidence to Support Regulatory Decision-Making for Medical Devices" explicitly states that real-world data must be "fit for use," meaning it must have sufficient relevance and reliability to answer the regulatory question at hand.
Reliability, in FDA's framework, includes data accrual (completeness), data assurance (accuracy and integrity), and data linkage (the ability to connect records across sources). These map directly to measurable trust dimensions: provenance, quality, concordance, recency, and validation.
The problem is that most manufacturers assess data fitness qualitatively. They write protocols describing their data sources and analysis plans. They do not score every record for trustworthiness before it enters the surveillance dataset. That gap between protocol-level assurance and record-level assurance is where regulatory risk concentrates.
Key statistics
FDA has cleared or approved over 1,000 AI/ML-enabled medical devices as of early 2025, each carrying post-market data obligations.
Section 522 surveillance orders can span up to 36 months of active data collection, with extensions possible for devices showing safety signals.
FDA's Medical Device Reporting system receives over 2 million individual adverse event reports per year, and the agency has repeatedly flagged data quality issues in these submissions.
SuperTruth's DTI Engine reduced data standardization time for imaware from 3 weeks to 2 hours across 105,000 diagnostic records, demonstrating the scale of efficiency gains possible when trust scoring is applied at the record level.
Manufacturers who fail to comply with 522 orders face Warning Letters, consent decrees, and potential civil monetary penalties; FDA issued 12 Warning Letters for post-market surveillance deficiencies in fiscal year 2023 alone.
Why AI-enabled devices face a different surveillance burden
Traditional medical devices have fixed performance characteristics. A hip implant performs the same way on day 1 as it does on day 1,000 (assuming no material degradation). The post-market question for traditional devices is primarily about long-term safety.
AI-enabled devices are different. Their performance characteristics can change for three reasons:
Each of these changes can alter device outputs without any modification to the device software itself. This means post-market surveillance for AI devices must monitor not just outcomes but input data characteristics, model behavior over time, and the trustworthiness of the data feeding the model.
FDA's predetermined change control plan framework acknowledges this reality. But the framework places the burden on manufacturers to define performance monitoring methods and thresholds. If the monitoring data itself is untrustworthy, the entire change control framework collapses.
Consider a diagnostic AI that uses EHR data to predict sepsis risk. If the underlying EHR data has inconsistent LOINC coding, undocumented time zone shifts in vital sign timestamps, or missing social determinants of health fields, the surveillance data will show apparent performance degradation that may be a data quality problem rather than an algorithm problem. Without record-level trust scoring, manufacturers cannot distinguish between the two.
What FDA reviewers actually look for in surveillance submissions
Based on FDA guidance documents, Warning Letters, and public advisory committee discussions, reviewers evaluate post-market surveillance data across several specific dimensions:
Completeness of follow-up: What percentage of enrolled subjects have complete outcome data? Surveillance studies with follow-up rates below 80% receive heightened scrutiny.
Source data verification: Can the manufacturer trace a reported outcome back to the original clinical record? FDA expects audit trails, not just summary statistics.
Timeliness of reporting: MDR regulations require manufacturers to report deaths within 30 days and serious injuries within 30 days. But beyond MDR, surveillance study protocols should specify data collection intervals that match the clinical question.
Consistency across sites: Multi-site surveillance studies must demonstrate that data collection methods are consistent. Variation in how sites capture adverse events, define clinical endpoints, or code diagnoses undermines the entire dataset.
Subgroup analysis capability: FDA increasingly expects manufacturers to report device performance across racial, ethnic, age, and sex subgroups. This requires demographic data that is complete and consistently coded, a requirement that many EHR-derived datasets fail to meet.
For AI-enabled devices, reviewers also look for:
Performance monitoring dashboards: Evidence that the manufacturer is actively tracking model performance metrics (AUC, sensitivity, specificity, calibration) against pre-specified thresholds.
Drift detection documentation: Evidence that the manufacturer has methods to detect distributional shifts in input data or output predictions.
Retraining documentation: If the device has been retrained, complete documentation of training data provenance, validation results, and any changes in performance.
How trust scoring closes the surveillance gap
The fundamental problem with post-market surveillance data is that manufacturers treat data collection and data trustworthiness as separate activities. They build surveillance systems that capture data, then rely on statistical analyses to detect problems in aggregate.
This approach misses the root cause. A surveillance dataset with 100,000 records where 30% have provenance gaps, 15% have stale timestamps, and 10% have coding inconsistencies will produce aggregate statistics that look reasonable but hide systematic bias.
Trust scoring at the record level changes the paradigm. When every record entering a surveillance dataset receives a score across dimensions like provenance, recency, quality, concordance, and validation, manufacturers can:
This is not theoretical. SuperTruth's Data Trust Index scores every health data record 0 to 100 across 8 dimensions: Provenance (25%), Consent (20%), Recency (15%), Quality (10%), Concordance (10%), Validation (10%), Breadth (5%), and Stability (5%). Applied to post-market surveillance data, the DTI provides exactly the kind of record-level assurance that FDA's "fit for use" standard demands but that no existing surveillance infrastructure delivers.
The regulatory trajectory: where post-market surveillance is heading
FDA's Total Product Life Cycle Advisory Program (TAP) and the increasing emphasis on real-world evidence signal a future where post-market data carries more regulatory weight, not less. Several trends are converging:
Continuous learning systems: FDA has signaled openness to AI devices that update their algorithms based on real-world data. This requires post-market surveillance data that is trustworthy enough to serve as training data, a standard far higher than traditional adverse event reporting.
Patient registries as surveillance infrastructure: FDA is increasingly relying on existing patient registries (like the National Cardiovascular Data Registry or the American Joint Replacement Registry) as post-market surveillance data sources. Registry data quality varies enormously, and the agency knows it.
International harmonization: The EU MDR's more prescriptive post-market surveillance requirements are influencing FDA expectations. Manufacturers selling in both markets need surveillance data that satisfies the more demanding standard.
Real-world performance requirements: FDA's guidance on predetermined change control plans effectively creates a continuous post-market performance monitoring obligation. This obligation requires data infrastructure, not just data collection.
Device manufacturers who build trust scoring into their surveillance data pipelines now will have a structural advantage as these requirements tighten. Those who wait will face the choice between retrofitting their data infrastructure under regulatory pressure or accepting the risk of non-compliance.
The DTI Engine scores every health data record 0 to 100 across 8 trust dimensions before your AI model sees it. If your team is building post-market surveillance infrastructure for AI-enabled devices and needs record-level trust assurance that satisfies FDA's fit-for-use standard, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.
Further reading:

Jason Alan Snyder
Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.
About SuperTruth · LinkedIn · Substack · jasonalansnyder.com
See it in practice
DTI scores the record, not the patient.
8 dimensions. 0 to 100. Travels with every record permanently.