BAM Dataset All articles
Medical Research & Policy

Broken Findings: Inside the Movement to Rebuild American Medical Research on a Foundation of Transparent Data

BAM Dataset
Broken Findings: Inside the Movement to Rebuild American Medical Research on a Foundation of Transparent Data

Photo: Unknown authorUnknown author, CC BY-SA 4.0, via Wikimedia Commons

The Study That Could Not Be Repeated

In 2011, Bayer HealthCare researchers attempted to replicate the findings of 67 preclinical oncology studies — the kind of foundational research that informs decisions about which drug candidates advance toward human trials. They succeeded with fewer than 25 percent of them. A separate effort by scientists at Amgen reached a similarly unsettling conclusion: of 53 landmark cancer biology studies reviewed, only 6 could be confirmed.

These were not fringe publications. Many had appeared in respected journals, generated significant citation counts, and directly influenced research investment decisions worth hundreds of millions of dollars. The implications were not merely academic. Patients enrolled in trials downstream of unreplicable findings were exposed to experimental interventions built on a foundation that had never been independently verified.

This is the reproducibility crisis — and in American medical research, it remains an unresolved structural problem rather than a corrected historical episode.

Why Studies Fail to Replicate

The causes are multiple and interconnected, which is precisely what makes the problem so resistant to simple fixes. Publication bias — the well-documented tendency of journals to accept positive findings over null results — creates a literature that systematically overrepresents effects that may not be real. When a study demonstrating a promising drug interaction is published and a dozen failed replications are not, the scientific record becomes a curated illusion.

Statistical practices compound the problem. P-hacking, the manipulation of analytical choices to push results below the conventional 0.05 significance threshold, is widespread enough to have generated its own scholarly literature. Underpowered studies — those with sample sizes too small to reliably detect the effects they claim to measure — produce findings that are mathematically unlikely to survive independent scrutiny.

Institutional incentives reinforce these tendencies. Academic promotion structures in the United States have historically rewarded novel, positive findings published in high-impact journals. A researcher who spends two years carefully replicating a prior study and finds nothing produces work that is scientifically valuable but professionally unrewarding under conventional metrics. The system, in effect, penalizes rigor.

Underlying all of these dynamics is a more fundamental problem: opacity. When the raw data underlying a published study is unavailable for inspection, independent researchers cannot determine whether the analytical choices made were appropriate, whether the sample was representative, or whether the reported statistics accurately reflect the underlying numbers. Without access to the data, replication is not merely difficult — it is, in many cases, impossible by design.

The Economic and Ethical Ledger

A 2015 analysis published in PLOS Biology estimated that the annual cost of irreproducible preclinical research in the United States exceeded $28 billion. That figure encompasses wasted laboratory expenditure, failed drug development pipelines, and the downstream costs of clinical trials built on preclinical evidence that does not hold. It does not capture the subtler costs: the erosion of public trust in biomedical institutions, the careers of researchers who pursued false leads in good faith, or the opportunity cost of scientific attention directed toward findings that were never real.

The ethical dimension is equally serious. When clinical guidelines are informed by studies that cannot be replicated, patients receive treatments whose evidence base is weaker than the published record suggests. In fields such as psychiatry, pain management, and nutrition science — areas where the reproducibility record is particularly troubled — the gap between published findings and reproducible knowledge has direct consequences for clinical decision-making across the American healthcare system.

Open Data as Structural Remedy

The researchers who have most visibly shifted toward open-data methodologies describe the transition as both principled and practical. When all analytical inputs — raw datasets, preprocessing code, statistical scripts — are made publicly available through verified repositories, the research community gains the ability to scrutinize not just conclusions but the process by which those conclusions were reached.

This level of transparency does not eliminate error. What it does is make error detectable and correctable. A statistical mistake buried in a proprietary dataset may persist in the literature for years; the same mistake embedded in a publicly accessible, version-controlled repository is far more likely to be identified and flagged by an independent reviewer.

Several institutions have begun formalizing this approach. The Center for Open Science, based in Charlottesville, Virginia, has registered more than 100,000 preregistered studies — research plans submitted to a public repository before data collection begins, making post-hoc analytical manipulation far more difficult to conceal. The NIH has progressively strengthened its data sharing policies, now requiring most funded clinical trials to deposit results in publicly accessible repositories within specified timeframes.

Researchers who have adopted these practices consistently report an initial friction — the process of preparing data for public deposit requires documentation and organization that proprietary workflows do not demand — followed by tangible professional benefits. Studies with openly available data are cited more frequently. They attract collaborators who can build directly on verified foundations. And they tend to survive scrutiny in ways that opaque research, however impressive its initial reception, often does not.

What Institutional Transition Actually Requires

For research institutions seeking to move toward open-data practices, the path forward involves several concrete commitments. Data management plans must be treated as scientific documents rather than administrative formalities, with genuine investment in the infrastructure needed to prepare, deposit, and maintain research data in accessible, well-documented formats.

Journal policies matter substantially. Peer reviewers who have access to underlying datasets can evaluate analytical choices in ways that are impossible when reviewing a manuscript alone. Several journals have introduced data availability requirements; broader adoption of these standards would meaningfully alter the incentive structure surrounding reproducibility.

Training is perhaps the most underappreciated element. Graduate programs in the biomedical sciences have historically devoted limited attention to data management, statistical power analysis, and research transparency practices. Building these competencies into standard scientific education would address the reproducibility problem at its source rather than attempting remediation after the fact.

Restoring Credibility Through Verification

The reproducibility crisis is not evidence that science is broken. It is evidence that science, practiced without adequate transparency and verification mechanisms, is vulnerable to the same institutional pressures and cognitive biases that affect every human enterprise. The corrective is not cynicism — it is structure.

Open-access, verified datasets represent that structure. They do not guarantee correct findings, but they create the conditions under which incorrect findings can be identified, challenged, and ultimately replaced by more reliable knowledge. In a healthcare system that serves more than 330 million Americans, the difference between findings that can be trusted and findings that merely appear credible is not a matter of academic interest. It is a matter of clinical consequence — and the case for building research on a foundation of transparent, reproducible data has never been more urgent.

All Articles

Related Articles

From Soil to Satellite: How Public Climate Data Is Leveling the Playing Field for American Farmers

From Soil to Satellite: How Public Climate Data Is Leveling the Playing Field for American Farmers