Paper Promises: Why Federal Data-Sharing Mandates Rarely Survive Contact With Reality
Photo: U.S. Government Accountability Office from Washington, DC, United States, Public domain, via Wikimedia Commons
On paper, the architecture looks solid. Before a single dollar of federal research funding changes hands, applicants submitting proposals to the National Institutes of Health or the National Science Foundation must include a data management plan — a formal document outlining how the datasets their research produces will be stored, organized, and made accessible to other scientists. The policy, which has been incrementally strengthened over the past decade, represents one of the most concrete institutional commitments to open science that the federal government has ever made.
In practice, the architecture has a significant structural flaw: almost no one is checking whether the building actually gets built.
A Mandate Without a Mechanism
Data management plans, or DMPs, have become a routine feature of the grant application process. Researchers, grant administrators, and program officers at federal agencies all interact with these documents at the proposal stage. What happens to them after an award is made is a different matter entirely.
Program officers at both NIH and NSF — the individuals most directly responsible for overseeing funded projects — frequently acknowledge that monitoring data-sharing compliance falls well outside their practical bandwidth. A typical program officer manages dozens of active awards simultaneously, each generating its own reporting requirements, budget modifications, and scientific updates. Verifying whether a dataset has been deposited in an appropriate repository, formatted correctly, and documented with sufficient metadata is a task for which most agencies have allocated neither dedicated staff nor automated systems.
The result is a compliance environment that researchers describe, with remarkable candor, as essentially voluntary. Grant administrators at major research universities confirm that their institutions rarely receive inquiries from federal agencies about whether data commitments have been honored. When project periods close and final reports are submitted, the question of whether datasets were actually shared is seldom raised.
What Non-Compliance Looks Like — and Doesn't
The absence of enforcement does not mean researchers are acting in bad faith. Many face genuine obstacles: insufficient institutional infrastructure for data archiving, uncertainty about which repository meets federal standards, concerns about proprietary data generated through industry partnerships, and the simple reality that data curation takes time that grant budgets rarely account for.
Others, however, make no attempt to comply — and face no consequences for it. Under current policy, federal agencies retain the theoretical authority to withhold future funding from researchers who fail to meet data-sharing obligations. In documented practice, that authority is almost never exercised. Researchers who have openly acknowledged non-compliance in conversations with program officers report receiving guidance, suggestions, or expressions of concern — but not penalties.
This creates a predictable dynamic. When compliance is optional in effect, it becomes optional in behavior. Researchers operating under competing pressures — publication deadlines, laboratory management, teaching obligations, the relentless pursuit of the next grant — rationally deprioritize tasks that carry no enforcement weight.
The Review Stage Illusion
One of the most persistent misconceptions about DMPs is that their quality is meaningfully evaluated during the peer review process. In many grant competitions, data management plans are reviewed as a separate administrative component rather than scored as part of the scientific merit evaluation. Reviewers assessing a proposal's intellectual contribution are not always the same individuals examining the data plan, and in some mechanisms, the DMP receives only a pass-or-fail determination rather than substantive critique.
This separation matters. A DMP that commits to depositing data in a named repository within six months of publication may satisfy the formal requirement while remaining entirely unenforceable — because no one in the review chain is tasked with verifying whether that commitment is realistic, adequately resourced, or ever fulfilled.
Some program officers argue that the review stage is simply the wrong place to solve a compliance problem. Data sharing is a post-award behavior, they note, and pre-award review can only evaluate intentions. That observation is accurate. It also underscores why the absence of post-award monitoring is so consequential.
What Genuine Accountability Would Require
Researchers and administrators who study open science policy have proposed several concrete mechanisms that could meaningfully close the gap between mandate and practice.
The most frequently cited is conditional renewal. Under this model, researchers seeking subsequent federal funding would be required to demonstrate that data from previously funded projects has been shared in accordance with their stated plans before a new award is processed. This approach shifts enforcement to a moment when agencies already exercise leverage — the funding decision — without requiring the creation of new oversight infrastructure.
A second proposal involves integrating data deposit verification into existing reporting workflows. NIH's Research Performance Progress Reports and NSF's equivalent mechanisms already require periodic updates on project activities. Adding a mandatory, verifiable data deposit confirmation to these reports — with links to publicly accessible repository records — would create a paper trail that program officers could audit without dramatically expanding their workload.
A third approach would direct a portion of indirect cost recovery toward institutional data management infrastructure, with compliance rates tied to continued eligibility. This would create incentives at the university level, where grant administrators have both the administrative capacity and the institutional interest to monitor researcher behavior.
None of these solutions is costless, and each carries implementation challenges. Conditional renewal, for instance, raises questions about how to handle legitimate delays caused by data sensitivity, embargo periods, or technical complications. Any workable system would need to distinguish between researchers who failed to share data and those who encountered genuine barriers — a distinction that itself requires administrative capacity.
The Data That Doesn't Exist
Perhaps the most striking feature of this policy landscape is how little systematic data exists on the problem itself. Federal agencies do not publish comprehensive statistics on DMP compliance rates. Researchers have attempted to measure data availability by searching for datasets linked to published papers from federally funded projects, and the results are consistently discouraging — availability rates for federally funded research datasets are estimated to be well below fifty percent in most fields, and substantially lower in some.
For a policy domain nominally committed to evidence-based decision-making, the absence of reliable compliance metrics is both ironic and consequential. Agencies cannot fix a problem they have not measured, and they cannot measure a problem they have not defined with sufficient precision to track.
BAM Dataset exists, in part, because the gap between data that should be publicly available and data that actually is publicly available has real scientific costs. Findings that cannot be reproduced, meta-analyses built on incomplete evidence bases, and junior researchers who cannot access the datasets they need to test new hypotheses — these are not abstract harms. They accumulate, quietly, every time a federally funded dataset fails to find its way into a repository.
The mandate exists. The intention, in most cases, is genuine. What remains missing is the institutional will to treat data sharing as a condition of public funding rather than a courtesy extended to the scientific community when circumstances permit.