Registry data quality for rare disease: what FDA expects from patient registries
Photo by Markus Spiske on Unsplash
insight

Registry data quality for rare disease: what FDA expects from patient registries

By Jason Alan Snyder·July 26, 2026

The FDA has issued specific guidance on what patient registries must deliver before rare disease data can support regulatory decisions. Fewer than 7,000 rare diseases have FDA-approved treatments, and registry data quality is a primary reason. This post maps the gap between what registries collect and what the FDA actually requires for accelerated approval, natural history studies, and post-market surveillance.

Roughly 25 to 30 million Americans live with a rare disease, yet fewer than 5% of the approximately 7,000 recognized rare diseases have an FDA-approved treatment. One of the most persistent bottlenecks is not the absence of patients or the lack of scientific interest. It is the quality of the data those patients generate inside registries that were never built to meet regulatory standards.

The FDA has been explicit about what it expects from patient registries used in rare disease drug development. Most registries do not meet those expectations. The gap is not theoretical. It affects which drugs reach patients, how quickly accelerated approvals convert to full approvals, and whether post-market safety commitments produce usable evidence.

What the FDA considers a rare disease

The Orphan Drug Act of 1983 defines a rare disease as any condition affecting fewer than 200,000 people in the United States. The FDA's Office of Orphan Products Development (OOPD) administers this designation and grants orphan drug status to therapies targeting these populations.

That 200,000-person threshold creates an immediate data problem. Small patient populations mean small registries, limited natural history data, and fragmented clinical records spread across dozens of institutions. When the FDA evaluates a rare disease therapy, it is often working with datasets that would be considered statistically underpowered in any other therapeutic area.

This is exactly why registry data quality matters more, not less, for rare diseases. Every record carries disproportionate weight.

What is the rare disease registry program?

Rare disease registry programs are organized efforts to collect longitudinal clinical, demographic, and outcome data from patients with a specific rare condition. These registries are typically sponsored by patient advocacy organizations, academic medical centers, pharmaceutical companies, or government agencies like the National Institutes of Health (NIH).

The NIH's National Center for Advancing Translational Sciences (NCATS) maintains the Global Rare Diseases Patient Registry Data Repository (GRDR), which provides a standardized infrastructure for rare disease registries. The Rare Diseases Clinical Research Network (RDCRN) supports more than 20 consortia conducting studies across over 200 rare diseases.

But the existence of a registry does not equal the existence of usable data. A 2019 analysis found that fewer than half of rare disease registries used standardized data elements, and only a minority had governance structures that would satisfy FDA requirements for regulatory-grade evidence.

What the FDA actually expects from patient registries

The FDA's 2016 guidance document, "Rare Diseases: Common Issues in Drug Development," and its 2016 draft guidance on "Use of Real-World Evidence to Support Regulatory Decision-Making for Medical Devices" lay out several expectations for registry data. The agency's 2020 guidance on "Real-World Data: Assessing Electronic Health Records and Medical Claims Data To Support Regulatory Decision-Making for Drug and Biological Products" reinforces these standards.

Here is what the FDA requires, distilled to operational specifics.

Prospective data collection with defined protocols

The FDA distinguishes sharply between registries that collect data prospectively with a written protocol and those that aggregate data retrospectively from existing records. For regulatory submissions, prospective registries with pre-specified endpoints carry far more weight.

The protocol must define inclusion and exclusion criteria, data elements, collection intervals, follow-up schedules, and procedures for handling missing data. Registries that lack a formal protocol are treated as convenience samples, not as evidence sources.

Standardized data elements

The FDA expects registries to use common data elements (CDEs) wherever possible. For rare diseases, the NIH's GRDR program publishes recommended CDEs that include demographics, diagnosis confirmation, disease progression markers, treatment history, and patient-reported outcomes.

When registries use non-standard terminology, data harmonization becomes a manual process that introduces errors and delays. The FDA has flagged inconsistent outcome definitions as a specific barrier to using registry data for accelerated approval confirmatory studies.

Data provenance and audit trails

Every data point submitted to the FDA must have a traceable origin. The agency expects registries to maintain audit trails that document who entered data, when, from what source, and whether any corrections were made.

This is where most academic and advocacy-sponsored registries fall short. A registry that cannot demonstrate chain of custody for its data will not pass FDA scrutiny, regardless of how many patients it enrolls. As we have written before, audit trails are the foundation of health AI accountability, and the same principle applies to regulatory submissions.

Missing data handling

Rare disease registries commonly have missing data rates exceeding 30% for key variables. The FDA requires sponsors to document the mechanism of missingness (missing completely at random, missing at random, or missing not at random) and to use appropriate statistical methods to handle gaps.

Simply excluding patients with incomplete records is not acceptable when your total population might be fewer than 500 patients. The FDA has rejected supplementary evidence from registries where missing data patterns suggested systematic bias.

Patient consent and re-consent

Registry consent must cover the specific uses the data will support. If a registry was originally designed for natural history research and a sponsor later wants to use the data for a regulatory submission, the consent framework must explicitly permit that use. The FDA has raised concerns about registries that lack re-consent mechanisms when data use expands beyond original scope.

This consent governance challenge mirrors what we see across healthcare data more broadly. Consent governance fails when it treats consent as a one-time checkbox rather than an ongoing relationship.

What are at least three things to know about clinical data registries and why they matter?

Clinical data registries serve functions that clinical trials cannot replicate. Here are the three most critical:

1. Registries capture natural history data that defines what "normal progression" looks like. For rare diseases without approved treatments, there is often no published natural history. Without this baseline, the FDA cannot evaluate whether a new therapy actually changes the disease trajectory. Registries are frequently the only source of this information.

2. Registries enable external control arms when randomized trials are infeasible. When a disease affects 500 people worldwide, randomizing half to placebo is ethically and practically impossible. The FDA has accepted well-constructed registry data as external comparators in rare disease submissions, but only when the registry meets the data quality standards described above.

3. Registries support post-market surveillance that determines whether accelerated approvals survive. Many rare disease drugs receive accelerated approval based on surrogate endpoints. The FDA then requires confirmatory evidence, often from registries, to convert to full approval. If the registry data is not regulatory-grade, the confirmatory study fails and the drug risks withdrawal.

Recent examples illustrate the stakes. The FDA has withdrawn or threatened to withdraw accelerated approvals when post-market commitments, including registry-based studies, failed to deliver confirmatory evidence on schedule.

What is the FDA accelerated approval for rare disease?

Accelerated approval under 21 CFR 314.510 (drugs) and 21 CFR 601.41 (biologics) allows the FDA to approve therapies based on a surrogate endpoint that is reasonably likely to predict clinical benefit. This pathway is disproportionately used for rare diseases because traditional endpoint-driven trials are often not feasible with small populations.

Between 1992 and 2023, the FDA granted more than 300 accelerated approvals, with rare disease indications representing a growing share. The Accelerated Approval Program Integrity Act of 2023 gave the FDA new authority to require post-approval studies and withdraw approvals more efficiently when confirmatory evidence is not produced.

Here is what matters for registry sponsors: the FDA increasingly expects post-market confirmatory evidence to come from registries rather than from new trials. This means your registry data must be collected with the same rigor you would apply to a pivotal trial. If it is not, you are building a data asset that cannot fulfill its intended regulatory purpose.

Key statistics

Data Trust Index: 8 dimensions and their weights for registry scoring
Data Trust Index: 8 dimensions and their weights for registry scoring

  • 7,000+ rare diseases recognized globally, with FDA-approved treatments available for fewer than 5% of them
  • 25-30 million Americans affected by rare diseases, a population comparable to the entire state of Texas
  • 30%+ missing data rate commonly observed in rare disease registry key variables, above the threshold most FDA reviewers consider acceptable
  • 4.8 years average diagnostic delay for rare disease patients in the U.S., meaning registries often capture patients years after symptom onset
  • 95% time reduction achieved by SuperTruth's DTI Engine when applied to diagnostic record standardization (105,000 records, 3 weeks reduced to 2 hours), demonstrating what structured quality scoring can do at scale
  • Where registries fail the FDA's trust test

    The top-ranking content on this topic focuses on the value proposition of registries and best practices in the abstract. What that content does not address is the specific failure modes that cause the FDA to reject or discount registry data. Here are the most common.

    Inconsistent diagnostic confirmation

    Rare disease registries frequently enroll patients based on physician diagnosis without requiring confirmation through genetic testing, biomarker assays, or standardized diagnostic criteria. The FDA has flagged this as a critical weakness, particularly for diseases with phenotypic overlap (such as hereditary transthyretin amyloidosis and other forms of amyloidosis, or the various subtypes of Ehlers-Danlos syndrome).

    When 15% of your registry population may be misdiagnosed, your efficacy signal is diluted before analysis even begins.

    Cross-site data variability

    Multi-site registries commonly show significant variation in how sites interpret data collection forms. One site may record "disease onset" as the date of first symptom, while another records the date of diagnosis. When your total enrollment is 300 patients across 12 sites, this kind of variability destroys the statistical power of your dataset.

    The FDA expects sponsors to demonstrate inter-site data concordance, a dimension that maps directly to the Concordance component of our Data Trust Index framework.

    Temporal gaps in longitudinal data

    Rare disease patients often travel long distances to specialty centers, leading to irregular follow-up intervals. A registry that collects data every 6 months at one site and every 18 months at another cannot produce reliable disease progression curves.

    Recency, the third-highest weighted dimension in the DTI framework at 15%, matters enormously here. Stale data in a rare disease registry is not just an inconvenience. It is a regulatory liability.

    Lack of patient-reported outcome (PRO) integration

    The FDA has increasingly emphasized patient-focused drug development, and its 2023 guidance on incorporating patient experience data reinforces the expectation that registries capture validated PRO measures. Registries that collect only clinician-reported data miss the patient experience dimension that the FDA now considers essential.

    As recent MedPageToday coverage of conditions like NMOSD demonstrates, the patient-facing reality of rare diseases, including how conditions manifest across different racial and ethnic groups, often diverges from what clinician-reported data captures alone.

    How data trust scoring applies to rare disease registries

    The challenge with rare disease registry data is not simply "is this data clean?" It is: "Can this data withstand FDA scrutiny across multiple dimensions simultaneously?"

    A record might have perfect data quality (no missing fields, correct formats) but fail on provenance (no audit trail showing where the data originated). It might have strong provenance but fail on recency (last updated 3 years ago). It might be recent and well-sourced but fail on consent (the original consent form did not cover regulatory submissions).

    This is why a multidimensional trust score matters. The DTI scores every record across 8 dimensions: Provenance (25%), Consent (20%), Recency (15%), Quality (10%), Concordance (10%), Validation (10%), Breadth (5%), and Stability (5%). For rare disease registries heading toward FDA submission, every dimension maps to a specific regulatory expectation.

    Provenance at 25% weight reflects the FDA's insistence on chain of custody. Consent at 20% reflects the re-consent challenges unique to registries that expand their data use over time. Recency at 15% captures the temporal integrity that longitudinal rare disease data demands.

    What registry sponsors should do now

    If your rare disease registry aims to support regulatory submissions, either directly or as external control data, here is the operational checklist.

    Adopt standardized common data elements. Start with the NIH GRDR CDEs and supplement with disease-specific elements from relevant clinical research networks.

    Implement audit trail infrastructure from day one. Retroactively reconstructing data provenance is expensive and often impossible. Every data entry, modification, and correction should be logged with timestamps and user identification.

    Build re-consent capability into your consent architecture. Your consent framework should anticipate that data uses will expand. ConsentOS-style tiered consent is not optional for registries with regulatory ambitions.

    Score your data before submission. Do not wait for an FDA reviewer to discover that 35% of your progression endpoints are missing or that three sites define "response" differently. Identify and remediate data quality gaps proactively.

    Plan for post-market from the beginning. If your registry will need to serve as a confirmatory data source for an accelerated approval, the data collection protocol should be designed for that purpose from enrollment forward.

    The cost of getting this wrong

    Registry data processing: before and after DTI scoring (imaware case study)
    Registry data processing: before and after DTI scoring (imaware case study)

    When a rare disease registry fails to meet FDA data quality standards, the consequences cascade. The sponsor cannot use registry data to support a regulatory submission. A new prospective study must be designed, funded, and enrolled, adding years and millions of dollars. Patients with no approved treatment wait longer. In the worst case, an accelerated approval is withdrawn because the confirmatory study, built on inadequate registry data, fails to produce the evidence the FDA requires.

    The imaware case study illustrates what happens when you apply structured data quality scoring to diagnostic records. SuperTruth standardized 105,000 records, reduced processing time from 3 weeks to 2 hours, and saved more than 200 hours per month. The same approach, applied to rare disease registries, can identify data quality gaps before they become regulatory rejection letters.

    For rare diseases, where every patient's data carries outsized weight, the difference between scored and unscored data is the difference between a therapy reaching patients and a therapy stalling in regulatory review.

    SuperTruth's trust layer turns RWE from a compliance liability into a competitive asset. If your team needs audit-ready provenance for FDA submission, post-market registry compliance, or rare disease data standardization, contact Louis Simeonidis at louis@supertruth.ai or (215) 918-4140.

    Further reading:

  • Pharma solution
  • DTI™ Engine
  • Rare disease registries and data trust requirements for research use
  • What makes health data Platinum-grade for FDA regulatory submission
  • How the FDA will audit your health AI's training data
  • Jason Alan Snyder

    Jason Alan Snyder

    Co-founder of SuperTruth and Artists & Robots, and an inventor on the Data Trust Index patents. Twenty-plus years building technology inside Interpublic Group. He writes here nearly every day on data trust, provenance, and what AI should be allowed to act on, and publishes essays on his Substack.

    About SuperTruth · LinkedIn · Substack · jasonalansnyder.com

    See it in practice

    750+ cancer search terms. Live in production.

    VIOLET maps behavioral signals 12–18 months before clinical presentation.

    See VIOLET in action
    Share