BAM Dataset All articles
Medical Research & Policy

Taxpayer-Funded, Publicly Unavailable: The Institutional Barriers Keeping NIH Research Data Out of Reach

BAM Dataset
Taxpayer-Funded, Publicly Unavailable: The Institutional Barriers Keeping NIH Research Data Out of Reach

Photo: DutchTreat, CC BY-SA 4.0, via Wikimedia Commons

Every year, the National Institutes of Health distributes approximately $4 billion in extramural research grants to universities, medical schools, and research hospitals across the United States. The science those grants produce — clinical observations, genomic sequences, epidemiological records, imaging datasets — is, at least in principle, a public asset. In practice, accessing it often requires an institutional affiliation, a subscription fee, or a formal data-sharing agreement that can take months to negotiate.

The gap between what federal funding promises and what the public actually receives has become one of the more consequential fault lines in contemporary research policy.

The Architecture of Inaccessibility

Understanding why publicly funded data remains privately held requires examining the layered system of incentives that governs academic research. Universities do not simply receive NIH grants and forward the resulting data to a central repository. Instead, institutions negotiate overhead recovery rates — sometimes exceeding 50 percent of direct research costs — that help fund everything from administrative salaries to building maintenance. The data generated under these arrangements is treated, in many cases, as a proprietary institutional asset.

Publishers compound the problem. When researchers submit findings to peer-reviewed journals, copyright transfer agreements frequently assign control of the underlying data to the journal or its parent company. A physician in rural Montana seeking access to a dataset that directly informs treatment decisions for her patients may find herself confronted by a paywall demanding $35 per article — for research her own tax dollars helped produce.

Legal frameworks have struggled to keep pace. The Bayh-Dole Act of 1980, originally designed to encourage commercialization of federally funded inventions, has been interpreted broadly enough to allow universities to assert intellectual property claims over research outputs in ways Congress almost certainly did not anticipate. Meanwhile, NIH's own data-sharing policies, while strengthened in recent years, have historically lacked robust enforcement mechanisms.

The 2023 Policy Shift and Its Limits

In January 2023, the NIH finalized its updated Data Management and Sharing Policy, requiring researchers who receive NIH funding to submit a data management plan alongside their grant applications. On paper, this represents a meaningful step toward open science. In practice, compliance remains uneven.

The policy permits researchers to cite "legitimate privacy, security, or ethical concerns" as grounds for restricting data sharing. These are real considerations, particularly in clinical research involving protected health information. But critics argue that the exemptions are drawn broadly enough to accommodate institutional reluctance that has little to do with patient protection and much to do with competitive advantage.

Dr. Jessica Polka, executive director of ASAPbio, a nonprofit that advocates for open research practices, has noted publicly that policy mandates without enforcement teeth tend to produce compliance theater rather than genuine openness. Institutions file data management plans. Whether the data actually becomes accessible to independent researchers is a separate question — and one that federal agencies have been slow to audit.

Who Bears the Cost of Closure

The consequences of restricted data access are not evenly distributed. Researchers at well-resourced R1 universities can typically negotiate access through institutional agreements and professional networks. Independent scientists, faculty at smaller teaching institutions, public health practitioners, and policy analysts working outside academia face far steeper barriers.

This asymmetry matters enormously for reproducibility. When a dataset underpinning a major clinical finding is available only to researchers affiliated with the institution that generated it, independent replication becomes structurally impossible. The reproducibility crisis that has attracted considerable attention in psychology and social science has a quieter but equally serious counterpart in biomedical research — and restricted data access is among its root causes.

There are also direct public health implications. During the early months of the COVID-19 pandemic, the inability of independent researchers to rapidly access clinical datasets from multiple institutions meaningfully delayed the synthesis of evidence on treatment protocols, risk factors, and population-level transmission dynamics. The fragmentation of data across institutional silos, each governed by its own access policies, imposed a measurable cost in human terms.

Institutions Breaking the Pattern

The picture is not uniformly discouraging. Several institutions have made deliberate commitments to open data that stand in notable contrast to the prevailing norm.

The University of California system has negotiated transformative open-access agreements with major publishers that include provisions for data availability. The Broad Institute at MIT and Harvard has built a widely used model for genomic data sharing that balances open access with appropriate privacy protections. The National Cancer Institute's Genomic Data Commons provides controlled-access infrastructure that allows qualified researchers to work with sensitive cancer genomics data without requiring them to negotiate bespoke agreements with individual institutions.

These examples demonstrate that open data and responsible data governance are not mutually exclusive. They require investment, institutional commitment, and governance frameworks that take both scientific openness and participant protection seriously — but they are achievable.

What Meaningful Reform Requires

Advocates for open science generally converge on several reform priorities. First, NIH should move from requiring data management plans to auditing actual data availability, with consequences for non-compliance tied to future funding eligibility. Second, federal grant agreements should include explicit provisions preventing copyright transfer arrangements that restrict data access. Third, Congress should revisit the interpretation of Bayh-Dole as it applies to data — as distinct from patentable inventions — to clarify that publicly funded datasets cannot be treated as proprietary institutional assets.

None of these changes are technically complex. Each faces significant resistance from the universities, publishers, and professional societies that benefit from the current arrangement.

The fundamental question is straightforward: when the American public finances the production of scientific knowledge, does it have a right to access that knowledge? The answer, embedded in the current system's design, is effectively no. Whether that answer is acceptable is a policy choice — and one that receives far less public scrutiny than it deserves.

At BAM Dataset, our work proceeds from a different premise: that verified, openly accessible data is not a courtesy extended to the research community but a precondition for the kind of cumulative, reproducible science that actually advances human understanding. The infrastructure for that kind of science exists. The will to require it, consistently, is still being built.

All Articles

Related Articles

Invisible Evidence: How Drug Approval Data Submitted to the FDA Disappears From Scientific Scrutiny

Invisible Evidence: How Drug Approval Data Submitted to the FDA Disappears From Scientific Scrutiny

When the Black Box Wins: Algorithmic Secrecy and the Reproducibility Crisis Reshaping Academic Science

When the Black Box Wins: Algorithmic Secrecy and the Reproducibility Crisis Reshaping Academic Science

Fragmented by Design: How Siloed Oncology Data Is Slowing the Fight Against Cancer

Fragmented by Design: How Siloed Oncology Data Is Slowing the Fight Against Cancer