BAM Dataset All articles
Medical Research & Policy

Invisible Evidence: How Drug Approval Data Submitted to the FDA Disappears From Scientific Scrutiny

BAM Dataset
Invisible Evidence: How Drug Approval Data Submitted to the FDA Disappears From Scientific Scrutiny

Photo: The U.S. Food and Drug Administration, Public domain, via Wikimedia Commons

When a pharmaceutical company seeks approval for a new drug in the United States, it submits a New Drug Application to the Food and Drug Administration that can run to hundreds of thousands of pages. Within those documents lie the clinical trial results that determine whether a medication reaches the market — the raw safety data, the efficacy endpoints, the adverse event records, the statistical analyses. This is, in the most direct sense, the evidentiary foundation of American drug policy.

Most of it will never be seen by an independent scientist.

A Regulatory System Built on Private Evidence

The FDA's drug approval process is, by design, a confidential one. Companies submit proprietary data under trade secret protections and federal regulations that restrict what the agency can disclose. The FDA's reviewers examine the submissions internally, and the agency publishes summary documents — known as drug approval packages — that provide some visibility into the review process. But the underlying datasets, the patient-level trial records that would allow independent researchers to verify the agency's conclusions, remain the property of the sponsoring pharmaceutical company.

This arrangement has a certain institutional logic. Drug development is expensive, and the promise of data exclusivity is one of the mechanisms through which the regulatory system attempts to sustain investment in new therapies. The counterargument — that evidence used to justify public health decisions affecting hundreds of millions of people should itself be publicly verifiable — has never fully penetrated the regulatory framework.

The consequences are more serious than they might initially appear. Independent reanalysis of clinical trial data has, on multiple occasions, produced conclusions that diverge substantially from those of the original sponsor. A landmark 2012 reanalysis of Tamiflu (oseltamivir) data by the Cochrane Collaboration found that the drug's benefits had been materially overstated in published literature — a finding that had significant implications for the billions of dollars governments worldwide had spent stockpiling the medication. That reanalysis was possible only because the researchers spent years in a protracted dispute with Roche to obtain the underlying data. It should not have taken years. It should not have required a dispute.

What the FDA Actually Publishes

The FDA does make certain information publicly available. Drug approval packages, accessible through the agency's Drugs@FDA database, include medical officer reviews, statistical analyses, and in some cases detailed clinical study reports. For researchers willing to invest significant time, these documents can be informative.

But there is a meaningful difference between a regulatory agency's internal summary of trial data and the underlying patient-level datasets that would allow independent statistical reanalysis. The former tells you what the FDA concluded. The latter would allow the scientific community to determine whether those conclusions were warranted — to check the work, in other words, with the same rigor we would apply to any other empirical claim.

The European Medicines Agency has moved further in this direction than its American counterpart. Since 2016, the EMA has operated a clinical data publication policy that provides access to clinical study reports for approved products. The policy has limitations and has faced legal challenges from pharmaceutical companies, but it represents a substantially more open posture than the FDA currently maintains.

The Ghost Data Problem

Perhaps the most troubling dimension of this issue is not the data that exists but is withheld — it is the data that may exist but has never been reported at all.

Clinical trial registries, including ClinicalTrials.gov, were established in part to address the problem of selective publication: the well-documented tendency for trials with positive results to be published at higher rates than trials with null or negative findings. A drug that appears to work in three published studies might look considerably less impressive if the four unpublished trials showing no effect were factored into the analysis.

Compliance with trial registration and results reporting requirements has improved but remains incomplete. A 2020 analysis published in The BMJ found that a substantial proportion of trials registered on ClinicalTrials.gov had not reported results within the legally required timeframe. The FDA has authority to levy fines for non-compliance, but enforcement has been inconsistent.

The result is a scientific literature on drug efficacy and safety that is systematically skewed toward positive findings — not because the underlying science is uniformly positive, but because the data infrastructure governing what gets reported, and to whom, has never been designed with completeness as its primary value.

What Reform Would Actually Require

Meaningful transparency reform in this domain would require changes at multiple levels simultaneously.

At the regulatory level, the FDA should adopt and enforce a requirement for the public release of patient-level clinical trial data — appropriately de-identified — for all products submitted for approval. The agency has the statutory authority to condition approval on data sharing commitments; it has not exercised that authority in a systematic way.

At the legislative level, Congress should strengthen and enforce the results reporting requirements of the Food and Drug Administration Amendments Act of 2007, including robust penalties for non-compliance that are actually assessed. It should also revisit the trade secret protections that currently shield clinical trial data from disclosure, drawing a clearer distinction between genuinely proprietary manufacturing processes and the evidentiary record underlying a public health decision.

At the institutional level, academic journals that publish pharmaceutical research should require data availability as a condition of publication — not as a stated policy aspirationally applied, but as a hard requirement with teeth.

None of this is a radical proposition. The scientific community's ability to verify, replicate, and build upon research findings is foundational to the enterprise of evidence-based medicine. Applying that standard to the data that determines which drugs Americans take is not an extraordinary demand. It is the minimum condition for a functional system of scientific accountability.

The Public Health Stakes

The argument for transparency in drug approval data is sometimes framed as a matter of abstract scientific principle. It is, more urgently, a matter of concrete public health consequence. Prescribing physicians rely on the published literature to make treatment decisions. Patients rely on their physicians. Policymakers rely on the evidentiary record to determine what the public insurance system should cover.

If that evidentiary record is incomplete — if the data underlying it is inaccessible to independent scrutiny — then the entire chain of inference on which clinical medicine depends is built on a foundation that cannot be verified. That is not a tolerable condition for a system that affects the health of every American.

Open, verified data is not merely a methodological preference. In this context, it is a prerequisite for trustworthy medicine.

All Articles

Related Articles

Taxpayer-Funded, Publicly Unavailable: The Institutional Barriers Keeping NIH Research Data Out of Reach

Taxpayer-Funded, Publicly Unavailable: The Institutional Barriers Keeping NIH Research Data Out of Reach

When the Black Box Wins: Algorithmic Secrecy and the Reproducibility Crisis Reshaping Academic Science

When the Black Box Wins: Algorithmic Secrecy and the Reproducibility Crisis Reshaping Academic Science

Fragmented by Design: How Siloed Oncology Data Is Slowing the Fight Against Cancer

Fragmented by Design: How Siloed Oncology Data Is Slowing the Fight Against Cancer