BAM Dataset All articles
Medical Research & Policy

Declared but Unreachable: The Quiet Collapse of Data Availability Statements in Published Research

BAM Dataset
Declared but Unreachable: The Quiet Collapse of Data Availability Statements in Published Research

Somewhere in the supplementary materials of a 2019 clinical nutrition study, there is a sentence that reads: "Raw data are available upon reasonable request to the corresponding author." The corresponding author left her university position in 2021. Her institutional email address bounces. The university's data management office has no record of the dataset. The study has been cited forty-seven times.

This is not an isolated anecdote. It is, by several measures, a representative one.

Across the biomedical and social sciences, the data availability statement — a short declaration appended to published papers specifying where underlying research data can be found — has become one of the most routinely ignored formalities in academic publishing. Journals require them. Funders mandate them. And in a significant proportion of cases, they describe data that is, for all practical purposes, gone.

The Architecture of a Broken Promise

The mechanics of data inaccessibility are varied, but they cluster into a handful of recurring failure modes.

The most common involves what researchers informally call the "upon request" declaration — a statement that raw data will be shared if a qualified researcher contacts the corresponding author. Studies examining this practice have found that actual response rates to such requests hover somewhere between poor and catastrophic, with some analyses placing successful data retrieval from "upon request" statements at well below fifty percent even within a few years of publication. As time passes, that number falls further. Corresponding authors retire, change institutions, or simply do not respond. There is no mechanism for enforcement, and journals rarely follow up after a paper is accepted.

A second failure mode involves links to institutional repositories or personal lab servers that were functional at the time of publication but have since moved or been decommissioned. Universities migrate their digital infrastructure. Departments reorganize. Faculty websites are archived or deleted when researchers leave. The links embedded in published papers do not update themselves, and the journals that published those papers have no obligation — and often no practical means — to monitor whether cited data locations remain accessible.

A third, arguably more troubling pattern involves datasets that were listed in published papers but never fully deposited in the first place. Some researchers cite repositories they intended to use, or deposit partial files that omit the raw observations needed to reproduce their analysis. The data availability statement, in these cases, functions less as a factual disclosure than as a gesture toward compliance — a box checked for the benefit of peer reviewers who rarely verify the contents.

The Replication Bottleneck

For researchers attempting to build on existing work, these failures are not abstract. They represent concrete dead ends.

Consider the position of a graduate student or early-career investigator trying to replicate the findings of a published meta-analysis. The paper may cite ten or fifteen component studies, each with its own data availability statement. If three of those links are broken, two researchers have left academia, and one dataset was only partially deposited, the replication attempt is compromised before it begins. The researcher cannot know whether the original findings hold, whether they contain errors, or whether they might have been interpreted differently under different analytical assumptions. The scientific record, at that point, has effectively closed itself off.

This problem carries particular weight in medical research, where the downstream consequences of unreproducible findings can extend well beyond academic inconvenience. Treatment guidelines, systematic reviews, and clinical decision tools are all built, at least in part, on the foundation of published studies. When the underlying data for those studies cannot be retrieved or verified, the evidentiary basis for those guidelines becomes opaque in ways that are difficult to communicate to practitioners or patients.

What Policy Has and Has Not Accomplished

The past decade has seen a significant expansion of formal data-sharing requirements from both public funders and major journals. The National Institutes of Health implemented a revised data management and sharing policy in 2023, requiring grant recipients to develop and adhere to data management plans. Many high-impact journals now require data availability statements as a condition of submission. Some have gone further, mandating deposit in recognized public repositories rather than accepting "available upon request" as a satisfactory response.

These are meaningful developments. They have increased the proportion of studies that deposit data in structured repositories, and they have made it marginally harder for authors to avoid the question of data sharing entirely. But policy adoption and policy enforcement are different things.

Journals that require data availability statements rarely audit them before publication. Peer reviewers, already burdened by the demands of evaluating a manuscript's scientific content, are not systematically asked to verify that cited datasets are accessible. Post-publication complaints about inaccessible data have no standard adjudication process at most journals, and corrections or expressions of concern are rarely issued on data-sharing grounds alone. The NIH's updated policy includes provisions for monitoring compliance, but the infrastructure for doing so at scale remains underdeveloped.

The result is a compliance culture that is wide but shallow — one in which the formal requirements exist, are acknowledged, and are frequently not met in any substantive sense.

Structural Conditions That Sustain the Problem

It would be a mistake to attribute this failure entirely to individual researchers acting in bad faith. The structural conditions that produce data inaccessibility are deeply embedded in how academic science is organized and rewarded.

Data management is time-consuming work that generates no direct professional credit. Preparing a dataset for public deposit — cleaning files, writing documentation, creating codebooks, selecting an appropriate repository, obtaining any necessary consent amendments — can represent a substantial investment of hours that produce nothing publishable. For researchers operating under pressure to produce manuscripts, secure grant renewals, and advance toward tenure, this investment is difficult to justify within existing incentive structures.

Institutional support for data management is also inconsistent. Large research universities with dedicated data librarians and established repository infrastructure are better positioned to assist researchers in meeting their obligations. Smaller institutions, underfunded departments, and independent research centers may lack these resources entirely. The burden of compliance, in these environments, falls entirely on individual investigators who may have received little training in data management practices.

Succession planning for research data is almost nonexistent as a formal institutional practice. When a faculty member departs — whether through retirement, a move to industry, or a career change — there is rarely a mechanism for transferring custodianship of their datasets to a colleague or institutional archive. The data simply goes with them, or it doesn't go anywhere at all.

Toward Accountability With Infrastructure

Addressing the data availability gap requires more than stronger language in journal policies. It requires the construction of systems that make compliance the path of least resistance rather than an additional burden.

Automated link-checking at the time of publication and at regular intervals thereafter is technically feasible and would at minimum surface broken citations before they propagate further into the literature. Expanded investment in domain-specific public repositories — with the staffing and storage capacity to actually onboard datasets from working researchers — would reduce the friction involved in proper deposit. Institutions that receive federal research funding could be required to maintain data succession plans as a condition of grant eligibility, ensuring that datasets do not become inaccessible simply because a principal investigator has moved on.

None of these solutions is costless. All of them require a level of institutional commitment that has, to date, been largely absent. But the alternative — a scientific literature increasingly populated by findings that cannot be examined, challenged, or built upon — carries costs of its own, costs that are borne not by the researchers who generated the data, but by everyone who relies on the integrity of the scientific record.

All Articles

Related Articles

Fabricated Foundations: When Synthetic Data Enters the Scientific Record Unannounced

Fabricated Foundations: When Synthetic Data Enters the Scientific Record Unannounced

Ink That Outlasted the Server: The Quiet Revival of Forgotten Laboratory Notebooks

Ink That Outlasted the Server: The Quiet Revival of Forgotten Laboratory Notebooks

Cited Into Thin Air: The Growing Problem of Scientific Datasets That Exist Only on Paper

Cited Into Thin Air: The Growing Problem of Scientific Datasets That Exist Only on Paper