Built on Vapor: The Orphaned AI Models Running on Datasets No One Can Find
There is a particular kind of scientific anxiety that emerges not from what a model gets wrong, but from the impossibility of determining whether it is right. Across the United States, research laboratories, hospital systems, and federal agencies are deploying artificial intelligence tools whose training data—the very substance from which these systems derived their predictive capacities—has quietly ceased to exist. The models keep running. The datasets do not.
This is the problem of the orphaned AI model: a system whose origin story has been erased, whose foundations cannot be inspected, and whose behavior can be observed but never fully explained. It is, in practical terms, a reproducibility crisis wearing a different face.
The Data Beneath the Model
To understand why this matters, it helps to understand what training data actually does. When a machine learning model is developed, it is exposed to vast quantities of structured information—patient records, imaging scans, annotated research outputs, genomic sequences—from which it learns statistical patterns. Those patterns become the model. The data, in a meaningful sense, is the model's reasoning.
When that data is unavailable, researchers lose the ability to ask foundational questions: Was the training population representative? Were there systematic biases baked into the labeling process? Did the dataset include samples from the demographic groups the model is now being applied to? These are not abstract methodological concerns. In medical research contexts, they carry direct implications for patient safety and scientific validity.
Yet the infrastructure for preserving AI training datasets has lagged severely behind the pace of model development. A 2023 analysis of machine learning papers published in high-impact biomedical journals found that fewer than a third provided persistent, accessible links to the datasets used in model training. Of those that did, a significant portion resolved to dead URLs, restricted institutional repositories, or data portals that had been quietly decommissioned.
Speed as the Enemy of Stewardship
The commercial and competitive pressures surrounding AI development have not been kind to data preservation practices. Research teams—whether at universities, private companies, or government-affiliated laboratories—routinely prioritize model performance benchmarks over documentation rigor. Datasets are assembled from multiple sources, sometimes under temporary licensing arrangements, and then processed, filtered, and transformed in ways that are rarely recorded with sufficient granularity.
By the time a model is published, presented at a conference, or integrated into a clinical workflow, the exact composition of its training data may already be partially inaccessible. Licensing agreements expire. Cloud storage contracts lapse. Graduate students graduate. The institutional memory that once connected a model to its origins disperses.
This is not always the result of negligence. Some training datasets incorporate proprietary or patient-derived data that cannot be publicly released for legitimate legal and ethical reasons. The problem is that even in these cases, the metadata—the documentation describing what the data contained, how it was collected, and how it was processed—often disappears along with the data itself. Researchers are left with a model and no map.
When Accountability Becomes Impossible
The consequences of this gap are most acute in high-stakes domains. Consider the proliferation of AI diagnostic tools in radiology and pathology, several of which have received FDA clearance based on validation studies that cited training datasets now inaccessible to independent reviewers. When clinicians or researchers raise questions about a model's performance on underrepresented populations—Black patients, rural communities, elderly adults—the mechanism for investigating those concerns through the training data simply does not exist.
The same dynamic appears in federally funded research contexts. Models developed under NIH grants, intended to advance reproducible science, are frequently published without any durable data archiving plan. The grant ends, the data storage budget expires, and the model persists in the literature—and sometimes in active use—as a kind of scientific artifact whose provenance has been severed.
Open science principles demand that findings be verifiable. A model whose training data cannot be examined is, by that standard, an unverifiable finding. It is a conclusion without auditable evidence.
What Responsible Practice Would Require
The open data research community has identified several practical interventions that could substantially reduce the orphaning problem, though none have been adopted at scale.
First, dataset registration at the point of model development—rather than at the point of publication—would create a timestamped record of what data existed and in what form. Repositories such as those maintained by the NIH National Library of Medicine or the Inter-university Consortium for Political and Social Research offer models for how this might work, but they were not designed with AI training data in mind and lack the metadata schemas necessary to capture the full complexity of how modern datasets are constructed.
Second, persistent identifiers for training datasets—analogous to the Digital Object Identifiers used for published papers—would allow citations to remain resolvable even when data moves or access conditions change. Several European research consortia have begun implementing this practice; American institutions have been slower to follow.
Third, and most structurally significant, funding agencies could require data archiving plans for AI components of research grants with the same seriousness currently applied to human subjects protocols. At present, data management plans submitted to federal agencies rarely address training data with specificity, and compliance is inconsistently enforced.
The Audit That Cannot Happen
There is a useful thought experiment for grasping the full weight of this problem. Imagine a pharmaceutical compound whose synthesis pathway had been documented only in the memory of researchers who have since left the field, whose laboratory notebooks have been discarded, and whose raw experimental data exists nowhere in retrievable form. The compound is on the market. It works, apparently. But no one can reconstruct how it was made, and no independent scientist can verify the claims made about its development.
Regulators would not accept this. The scientific community would not accept this. Yet for a growing number of AI models embedded in research pipelines and clinical systems, this is precisely the situation that currently obtains.
The phrase "open science" implies more than open access to published findings. It implies that the foundations of those findings—the data, the methods, the conditions under which knowledge was produced—remain accessible and challengeable over time. When training datasets vanish, that principle is violated not through any single act of misconduct, but through the accumulated weight of institutional indifference to data stewardship.
The models keep running. The datasets are gone. And the scientific record, built partly on their outputs, carries the weight of foundations no one can examine.