BAM Dataset All articles
Medical Research & Policy

Trained to Miss: How Rare Disease Patients Are Being Systematically Excluded From Medical AI

BAM Dataset
Trained to Miss: How Rare Disease Patients Are Being Systematically Excluded From Medical AI

Photo: National Human Genome Research Institute (NHGRI) Web site genome.gov is in the public domain., CC0, via Wikimedia Commons

When a machine learning model is asked to identify a rare genetic disorder, it reaches into a training dataset and finds, in many cases, almost nothing. The absence is not accidental. It is the predictable outcome of how medical AI systems are built, funded, and deployed in the United States — and it carries consequences that fall most heavily on the patients least equipped to absorb them.

There are approximately 7,000 recognized rare diseases affecting an estimated 30 million Americans. For most of those conditions, the genomic data required to train a diagnostic model simply does not exist in sufficient volume within any open-access repository. The datasets that do exist are concentrated in a small number of academic medical centers and, more significantly, within proprietary pharmaceutical databases that are never made publicly available.

The Data That Doesn't Exist — and the Data That's Hidden

The foundational problem is one of scale and incentive. Medical AI models require large, well-annotated datasets to achieve clinically meaningful accuracy. Common conditions — type 2 diabetes, coronary artery disease, certain cancers — generate millions of patient records and are well-represented in publicly accessible repositories like the National Institutes of Health's dbGaP or the UK Biobank, which American researchers frequently access. Rare diseases, by definition, generate far fewer data points.

But scarcity of patients is only part of the story. The data that does exist is often locked away. Pharmaceutical companies conducting natural history studies on rare conditions accumulate genomic and phenotypic data over years of clinical trial enrollment. That information, once trials conclude, is retained as proprietary intellectual property. It is not submitted to open repositories. It is not made available for independent model training. It sits in internal databases, occasionally mined for follow-on drug development, but effectively invisible to the broader research community.

Research institutions present a parallel problem. Academic medical centers with rare disease clinics have access to patient cohorts that could meaningfully expand open genomic datasets. Yet institutional data governance frameworks, concerns about patient re-identification, and the competitive pressures of grant-funded research all create friction against open contribution. The result is that the genomic record of rare disease in America is fragmented, incomplete, and largely inaccessible to the researchers trying to build tools that could benefit patients.

When the Algorithm Doesn't Recognize You

The clinical consequences of this data gap are not abstract. Consider the diagnostic odyssey — the years-long process, averaging five to seven years in the United States, that many rare disease patients endure before receiving an accurate diagnosis. AI-assisted diagnostic tools have been positioned as a potential solution, capable of cross-referencing symptom profiles, lab values, and imaging data against vast medical knowledge bases to surface unlikely diagnoses that a clinician might not consider.

But a diagnostic AI trained predominantly on common-condition data will not reliably surface a rare genetic disorder. It will pattern-match against what it knows. Patients presenting with atypical symptom clusters may receive confident algorithmic suggestions pointing toward more common conditions — suggestions that carry the weight of technological authority even when they are wrong. Clinicians, already under time pressure, may reasonably defer to a well-validated tool without recognizing that the tool was never validated for the patient in front of them.

This is not a hypothetical failure mode. Researchers examining the performance of commercially deployed diagnostic AI systems have documented significant accuracy disparities across patient subpopulations. Rare disease cohorts, when tested against these systems, consistently underperform relative to the general population benchmarks that vendors use to market their products.

The Incentive Architecture That Perpetuates the Gap

Understanding why this problem persists requires examining the economic logic that governs medical data in the United States. Pharmaceutical companies operating in the rare disease space — a sector that has expanded significantly since the passage of the Orphan Drug Act — have strong financial incentives to develop proprietary datasets. Rare disease drugs can command extraordinary prices, and the genomic and biomarker data underlying them represents a competitive moat. Sharing that data openly would, in the view of most corporate legal and strategy teams, erode that advantage.

For academic researchers, the incentive structure is different but produces similar outcomes. Grant funding in rare disease research is competitive and often tied to institutional publication records. Assembling a unique patient cohort and the associated genomic data is a years-long investment. Contributing that data to an open repository before publication — or even shortly after — risks being scooped by better-resourced competitors. The rational individual response to these pressures is to hold data close, even when the systemic cost of doing so is significant.

Federal policy has attempted to address this dynamic with limited success. The NIH's data sharing mandate, strengthened in 2023, requires that research funded by the agency produce a data management and sharing plan. But the mandate contains exceptions, and compliance mechanisms remain inconsistent. Rare disease data generated through industry-sponsored trials falls almost entirely outside the policy's reach.

Building a More Complete Record

Some efforts are underway to close the gap. The Undiagnosed Diseases Network, a federally funded consortium, has committed to depositing case data into open repositories and has made meaningful progress in characterizing previously undocumented conditions. Patient advocacy organizations, particularly those representing specific rare disease communities, have become increasingly sophisticated in advocating for data sharing as a condition of research participation — a form of patient-driven data governance that has produced results in conditions like Duchenne muscular dystrophy and certain pediatric cancers.

Open-access repositories with rare disease-specific infrastructure, including Orphanet and the European Genome-phenome Archive, have demonstrated that collection at scale is achievable when institutional will and funding are aligned. The scientific community's challenge is to build equivalent infrastructure within the United States and to create the policy and financial incentives that would bring pharmaceutical and academic data into it.

The machine learning models that will shape the next generation of medical diagnosis are being trained right now. The datasets used to train them will determine, with considerable precision, which patients those models can help and which patients they will miss. Rare disease patients are currently on the wrong side of that determination — not because their conditions are inherently unrecognizable, but because the data infrastructure that would make them recognizable has not been built, shared, or opened.

That is a policy failure as much as a technical one. And like most policy failures, it is reversible.

All Articles

Related Articles

Replication as Resistance: Independent Researchers Are Auditing Oncology Science — and Finding It Wants

Replication as Resistance: Independent Researchers Are Auditing Oncology Science — and Finding It Wants

Taxpayer-Funded, Publicly Unavailable: The Institutional Barriers Keeping NIH Research Data Out of Reach

Taxpayer-Funded, Publicly Unavailable: The Institutional Barriers Keeping NIH Research Data Out of Reach

Invisible Evidence: How Drug Approval Data Submitted to the FDA Disappears From Scientific Scrutiny

Invisible Evidence: How Drug Approval Data Submitted to the FDA Disappears From Scientific Scrutiny