BAM Dataset All articles
Medical Research & Policy

Funding the Same Discovery Twice: The Measurable Cost of Undiscoverable Public Research Data

BAM Dataset
Funding the Same Discovery Twice: The Measurable Cost of Undiscoverable Public Research Data

Federal agencies distribute billions of dollars annually to researchers who, in a meaningful fraction of cases, are collecting data that already exists somewhere in the public record — labeled poorly, formatted inconsistently, or deposited in a repository no search engine indexes effectively. The resulting duplication is not a minor inefficiency. It is a structural tax on American scientific productivity, and it compounds silently across thousands of grant cycles without appearing as a line item in any budget.

Estimating the precise magnitude of this waste is difficult, which is itself part of the problem. Studies examining specific research domains have found duplication rates ranging from concerning to alarming. A 2022 analysis of clinical biomarker research found that a substantial proportion of studies collecting baseline physiological data were unaware of existing, publicly deposited datasets that covered the same patient populations and measurement protocols. In environmental health research, parallel findings have emerged: investigators writing new data collection into their grant proposals frequently cannot find, through standard search procedures, the datasets that previous funded work has already produced.

The Discoverability Gap

The failure is rarely one of nonexistence. In many cases, the relevant dataset has been deposited — the researcher who generated it fulfilled the letter of their funding agency's data-sharing requirement. The dataset exists in a public repository. It is, in a technical sense, available. It simply cannot be found by someone who does not already know it exists.

The reasons for this are multiple and mutually reinforcing. Metadata standards across the research data ecosystem are inconsistent. A dataset deposited in one repository may use controlled vocabulary terms that differ entirely from those used by a comparable dataset in another repository. A researcher searching for longitudinal cardiovascular data collected from a specific demographic in a specific geographic region may find no results not because no such data exists but because the depositing researcher used different terminology, omitted geographic tagging, or described the study population in ways that do not surface under standard search queries.

Repository fragmentation compounds the problem. The United States does not have a unified national research data infrastructure. Data generated under NIH funding may be deposited in dbGaP, or the NCBI, or a university institutional repository, or a journal's supplementary materials system, or a domain-specific archive — or, in a meaningful number of cases, in a personal or lab website that will cease to exist when the PI moves institutions. A researcher conducting a thorough literature review before writing a grant proposal would need to search dozens of repositories using different interfaces, vocabularies, and access protocols to be confident they had not missed relevant existing data.

In practice, most researchers do not conduct that search. They conduct a literature search — which surfaces papers, not datasets — and proceed on the assumption that if relevant data existed and was accessible, they would have encountered a reference to it in the literature. This assumption is reasonable given the available tools. It is also frequently wrong.

Counting the Cost

Translating discoverability failures into dollar figures requires assumptions, but the available estimates are striking. Research examining duplication in federally funded data collection efforts has suggested that between ten and thirty percent of new data collection in certain biomedical fields replicates information already present in accessible but undiscoverable public datasets. Applied to the NIH's annual extramural research budget — which has exceeded thirty billion dollars in recent years — even the conservative end of that range implies billions of dollars annually spent collecting data that already exists.

The fifty million dollar figure that serves as a threshold for this analysis is not a single documented case but a representative order of magnitude: the scale at which duplication waste accumulates within a single research domain over a grant cycle. Individual researchers rarely waste that sum in isolation. The figure emerges from aggregating smaller redundancies — a two-hundred-thousand-dollar data collection component in a grant here, a four-hundred-thousand-dollar longitudinal survey there — across dozens of projects that would have been designed differently had their investigators been able to find and access the data that preceded them.

Researchers on the Ground

The experience of encountering this problem at the individual level is one of quiet frustration rather than dramatic failure. A principal investigator at a research university in the Mid-Atlantic describes spending weeks, over the course of several months, attempting to locate a specific type of patient outcome data that she had seen referenced in multiple papers. The papers cited the data. The citations led to journals. The journals pointed to supplementary material links that no longer resolved. Emails to corresponding authors went unanswered or produced responses indicating the data was no longer accessible. She ultimately wrote new data collection into her grant proposal, received funding, and collected the data herself — at which point a colleague mentioned, almost in passing, that a nearly identical dataset had been deposited in a domain repository three years earlier under a title that bore no obvious relationship to the subject matter.

This pattern — data that exists but cannot be found through reasonable effort — is reported consistently by researchers across fields. The cost is not only financial. It is temporal: careers, particularly early-career trajectories, are shaped by the projects that receive funding, and projects built on redundant data collection are projects that could have addressed different questions.

What Would Actually Help

Several interventions have been proposed with varying degrees of institutional traction. Cross-repository search infrastructure — a unified index that harvests metadata from major public repositories and exposes it through a common search interface — would address the fragmentation problem directly. Efforts in this direction exist, including initiatives supported by federal agencies and research library consortia, but none has achieved the comprehensiveness or usability that would make it a reliable first stop for grant preparation.

Mandatory metadata standards, enforced at the point of deposit rather than recommended as a best practice, would address the vocabulary inconsistency problem. Funders requiring that deposited datasets conform to a specified metadata schema — including controlled terms for subject matter, geographic scope, study population, and measurement instruments — would substantially improve the precision of cross-repository search.

Pre-award data discovery requirements represent a more direct intervention: requiring grant applicants to document, as a condition of application, that they have searched specified repositories for relevant existing data and to explain why existing data is insufficient for their proposed research. This would not eliminate all duplication, but it would create an institutional moment at which the question is asked — a moment that currently does not reliably exist in most funding agency review processes.

The Return on Investment

The case for investing in data discoverability infrastructure is, in purely economic terms, straightforward. If improved search tools, metadata standards, and pre-award review processes redirected even five percent of currently duplicated data collection toward novel research questions, the return on that infrastructure investment would be measurable in tens of millions of dollars annually — and that figure does not account for the scientific value of the new knowledge that redirected funding would generate.

Publicly funded research data is a public asset. Making that asset findable is not a technical luxury. It is a basic condition of the return on investment that taxpayers have already made.

All Articles

Related Articles

Rebuilding From Scratch: The Volunteer Scientists Reconstructing Research Data That Never Should Have Disappeared

Rebuilding From Scratch: The Volunteer Scientists Reconstructing Research Data That Never Should Have Disappeared

Preserved but Unreachable: The University Vaults Where Decades of Research Data Quietly Disappear

Preserved but Unreachable: The University Vaults Where Decades of Research Data Quietly Disappear

Training Data Is the Science: Why Deleting It After Deployment Undermines Every AI Model Built on It

Training Data Is the Science: Why Deleting It After Deployment Undermines Every AI Model Built on It