BAM Dataset All articles
Medical Research & Policy

Lost Before They Can Be Found: The Silent Erosion of Scientific Datasets After Publication

BAM Dataset
Lost Before They Can Be Found: The Silent Erosion of Scientific Datasets After Publication

Photo: San Diego Air & Space Museum Archives, Public domain, via Wikimedia Commons

Scientific knowledge is supposed to accumulate. Each published study is meant to build on what came before, adding a verified layer to the edifice of human understanding. But what happens when the foundation itself begins to crumble? Across disciplines—from epidemiology to agricultural genomics—researchers attempting to reproduce, extend, or simply verify prior findings are encountering the same frustrating obstacle: the data no longer exists where it was said to exist.

This is not a minor inconvenience. It is a structural failure with measurable consequences.

The Numbers Behind the Disappearance

A frequently cited analysis published in the journal Current Biology found that the odds of successfully accessing a dataset decline by approximately 17 percent with each year following publication. By the five-year mark, a substantial portion of datasets referenced in journal articles are effectively unreachable. More recent audits of biomedical literature have found broken or unresponsive data links in anywhere from 20 to 80 percent of papers examined, depending on the field and the age of the publication.

For a repository like BAM Dataset—built on the premise that verified, accessible data is the prerequisite for genuine discovery—these figures represent more than an academic concern. They represent a direct threat to the reproducibility of the scientific record itself.

The causes are varied, but several patterns emerge with uncomfortable regularity. Institutional repositories frequently operate on grant-funded timelines. When funding lapses, so does maintenance. A researcher deposits data with a university system in 2017, accepts a position at a different institution in 2020, and by 2023 the original server has been decommissioned without any formal data handoff protocol. The dataset does not disappear dramatically. It simply becomes unreachable—a ghost citation.

The Human Cost of Institutional Gaps

Consider the downstream effects on clinical and translational research. A graduate student attempting to conduct a meta-analysis of early-stage diabetes intervention trials contacts corresponding authors from fifteen studies published between 2012 and 2018. Of the fifteen, she reaches nine. Of those nine, three no longer have access to the original data files. Two others have data stored in formats that are no longer readable by current software. Four provide usable datasets. The meta-analysis proceeds on a fraction of the available evidence base—not because researchers were secretive or uncooperative, but because no durable infrastructure existed to preserve what was created.

This pattern repeats across fields. In agricultural science, soil microbiome datasets collected during multi-year federal grants have been lost when principal investigators retired. In neuroscience, imaging datasets that required months of participant recruitment and tens of thousands of dollars in equipment time have been rendered inaccessible by the discontinuation of proprietary storage platforms.

The scientific community is, in effect, conducting a slow-motion book burning—not out of malice, but out of structural indifference.

Why Existing Solutions Have Fallen Short

The instinct to address this problem through individual researcher responsibility has largely failed. Data management plans, required by funding agencies including the NIH and NSF, are frequently treated as compliance checkboxes rather than genuine preservation strategies. Researchers under pressure to publish, secure the next grant, and manage laboratory operations rarely have the bandwidth to become archivists.

Commercially operated repositories present a different set of risks. Platforms that offer free hosting at launch may pivot their business models, introduce paywalls, or simply cease operations—taking hosted data with them or rendering it inaccessible to users without institutional subscriptions. The consolidation of academic infrastructure under a small number of large commercial publishers has, in several documented cases, resulted in data previously available under open licenses being absorbed into subscription-gated environments.

Persistent identifiers such as DOIs and ORCIDs were designed to address link rot, but they function only as well as the infrastructure they point to. A DOI that resolves to a decommissioned server is, for practical purposes, as useless as a broken URL.

Toward Durable Preservation: What the Evidence Supports

The archival science community has long grappled with analogous challenges in the preservation of digital cultural heritage, and several of its principles translate directly to scientific data. Redundant storage across geographically distributed nodes—a model employed by the Internet Archive—significantly reduces the risk of total loss from any single institutional failure. Standardized, non-proprietary file formats ensure that data remains readable across decades and across software generations.

Federal intervention offers another lever. Legislation mandating that federally funded research data be deposited in perpetually maintained public repositories—with explicit transfer protocols when researchers change institutions—would address the single largest source of dataset loss. The NIH's 2023 Data Management and Sharing Policy moves in this direction, but enforcement mechanisms and long-term storage funding remain areas requiring further development.

Domain-specific repositories with dedicated curatorial staff, rather than general-purpose institutional systems, have demonstrated stronger preservation outcomes. When data deposit is treated as a professional curatorial act rather than an administrative afterthought, datasets survive longer and remain more usable.

For the broader open science community, the lesson is clear: transparency at the moment of publication is necessary but not sufficient. Data that is open today and inaccessible in five years has contributed little to the cumulative scientific enterprise. The goal is not merely to share data—it is to preserve it.

A Call for Systemic Accountability

The scientific community has invested enormous effort in addressing fabrication, selective reporting, and methodological opacity. These are legitimate concerns. But the quiet disappearance of valid, well-collected data through infrastructural neglect deserves equivalent attention. A dataset that cannot be found cannot be verified, cannot be extended, and cannot inform the next generation of inquiry.

Open-access data repositories have a particular responsibility here—not merely to host data at the moment of submission, but to maintain it, migrate it across format changes, and guarantee its availability to researchers who may not yet know they need it. That commitment is not a technical challenge alone. It is an ethical one.

All Articles

Related Articles

Rewiring the Incentives: How a New Generation of Universities Is Making Data Sharing a Career Asset

Rewiring the Incentives: How a New Generation of Universities Is Making Data Sharing a Career Asset

Replication as Resistance: Independent Researchers Are Auditing Oncology Science — and Finding It Wants

Replication as Resistance: Independent Researchers Are Auditing Oncology Science — and Finding It Wants

Trained to Miss: How Rare Disease Patients Are Being Systematically Excluded From Medical AI

Trained to Miss: How Rare Disease Patients Are Being Systematically Excluded From Medical AI