Fragmented by Design: How Siloed Oncology Data Is Slowing the Fight Against Cancer
Photo: Unknown photographer, Public domain, via Wikimedia Commons
Consider the paradox at the center of modern cancer research. Clinical trials—the gold standard for evaluating new therapies—generate some of the richest, most detailed patient data in all of medicine. Genomic profiles, treatment response curves, adverse event logs, imaging sequences, and longitudinal survival records accumulate across thousands of trial participants. The information exists. The scientific questions that could be answered with it are well-defined. And yet, for most researchers, that data remains effectively out of reach.
The problem is not volume. It is fragmentation. Oncology trial data sits distributed across hundreds of hospital systems, dozens of pharmaceutical sponsors, government registries, and private contract research organizations—each operating under different data standards, different privacy frameworks, and different institutional incentives. The result is a landscape where the signal researchers need to identify drug interactions, predict treatment resistance, or flag emerging safety concerns is present in the aggregate but invisible in any single slice.
The Architecture of Inaccessibility
To understand why oncology data is so thoroughly scattered, it helps to trace how clinical trial information is generated and controlled.
When a pharmaceutical company sponsors a Phase III cancer trial, it retains ownership of the trial data. Participating hospitals collect patient information under protocols approved by their own institutional review boards, but the data flows upward to the sponsor, not outward to the research community. When the trial concludes—whether it succeeds or fails—the sponsor is legally required to report summary results to ClinicalTrials.gov, the federal registry maintained by the National Library of Medicine. But summary results are not patient-level data. They are aggregate statistics that tell researchers what happened on average, not what happened to specific subgroups, in specific treatment sequences, or under specific comorbidity conditions.
Patient-level data—the granular records that make sophisticated secondary analyses possible—typically remains with the sponsor indefinitely. Some companies have implemented voluntary data-sharing programs. Pfizer's Clinical Trial Data Sharing initiative, GlaxoSmithKline's participation in the Yale Open Data Access project, and the multi-sponsor Project Data Sphere platform represent genuine efforts to broaden access. But participation is discretionary, the data that gets shared is curated rather than comprehensive, and researchers seeking access must navigate application processes that can take months.
Publicly funded trials present a different but equally frustrating picture. The National Cancer Institute sponsors an extensive portfolio of cooperative group trials through mechanisms like the Alliance for Clinical Trials in Oncology and NRG Oncology. Data from these trials is nominally available through the NCI's Cancer Data Access System. In practice, researchers report that the application process is cumbersome, turnaround times are long, and the data, once obtained, frequently requires extensive harmonization before it can be analyzed alongside data from other sources.
What Gets Lost in the Gaps
The practical consequences of this fragmentation surface most acutely in the study of drug interactions and treatment sequencing—areas where individual trials are structurally unable to provide answers.
A standard Phase III oncology trial enrolls patients according to carefully defined eligibility criteria. Patients with certain prior treatments, organ function thresholds outside a specified range, or concurrent medications are frequently excluded. This is methodologically appropriate for establishing a clean causal estimate of a drug's effect. But it means the trial population is systematically different from the real-world patient population who will ultimately receive the drug. The questions that matter most for practicing oncologists—how does this agent perform in patients who have already received two prior lines of therapy? what happens when it is combined with a common diabetes medication that many cancer patients take?—are precisely the questions no single trial is designed to answer.
Answering them requires pooling data across multiple trials, across multiple institutions, across multiple sponsors. And that is where the siloed architecture becomes a genuine barrier to discovery.
Researchers at Memorial Sloan Kettering Cancer Center have documented cases where meaningful drug interaction signals were identifiable only after combining patient-level data from three separate trial datasets—a process that required over a year of data-access negotiations and produced a harmonized dataset that no individual institution could have assembled independently. The interaction they identified had clinical implications for a patient population numbering in the tens of thousands annually in the United States alone.
Regulatory Frameworks and Their Limits
The regulatory environment governing clinical trial data sharing has evolved, but unevenly.
The Food and Drug Administration Amendments Act of 2007 established the legal basis for the ClinicalTrials.gov reporting requirements that now cover most federally funded and many industry-sponsored trials. The 2016 Final Rule expanded those requirements significantly, mandating results reporting within 12 months of trial completion for a broader set of studies. Compliance, however, has been inconsistent. A 2020 analysis published in The BMJ found that a substantial proportion of trials subject to the Final Rule had not reported results within the required window, with pharmaceutical sponsors and academic medical centers both appearing in the non-compliant cohort.
The European Medicines Agency has moved more aggressively than its US counterpart on patient-level data sharing, requiring sponsors of approved drugs to make clinical study reports—detailed documents that include more granular data than summary results—publicly available through its clinical data portal. US researchers frequently express frustration that data on drugs approved by both the FDA and EMA is more accessible through European channels than domestic ones.
Open Data Initiatives Gaining Ground
Against this backdrop, several initiatives are demonstrating that interoperable, accessible oncology data is achievable—and that it accelerates discovery in measurable ways.
Project Data Sphere, a not-for-profit initiative launched in 2014, hosts patient-level data from completed cancer trials contributed voluntarily by member companies. As of 2024, the platform holds data on over 100,000 patients across more than 150 trials. Researchers who have used the platform report that the ability to conduct secondary analyses across multiple datasets has surfaced findings that would have been statistically underpowered—or simply invisible—in any constituent trial alone.
The Genomic Data Commons, maintained by the NCI, provides a unified repository for genomic and clinical data from NCI-funded studies including The Cancer Genome Atlas and the Clinical Proteomic Tumor Analysis Consortium. The platform has become a critical resource for translational researchers, demonstrating that large-scale, standardized oncology data sharing is operationally feasible when institutional commitment exists.
At the standards level, the CDISC (Clinical Data Interchange Standards Consortium) Oncology Therapeutic Area User Guide provides a common data model for representing trial data—a prerequisite for meaningful cross-dataset analysis. Broader adoption of these standards by trial sponsors and academic medical centers would dramatically reduce the harmonization burden that currently consumes so much researcher time.
The Path Forward
No single policy change will resolve the fragmentation of oncology trial data. The problem is structural, reflecting the distributed and competitive nature of pharmaceutical development, the legitimate privacy interests of trial participants, and the genuine complexity of harmonizing data collected under different protocols.
But the direction of necessary change is clear. Mandatory patient-level data sharing for federally funded trials, with standardized formats and reasonable de-identification protections, would substantially expand the research community's analytical surface area. Stronger FDA enforcement of existing results-reporting requirements would close gaps in the public record. And sustained investment in shared infrastructure—platforms, standards, and the human expertise required to maintain them—would make the interoperable oncology data ecosystem that researchers need a reality rather than an aspiration.
The data to accelerate cancer drug discovery already exists. At BAM Dataset, we believe that verified, accessible data is not a luxury—it is a prerequisite for the kind of discovery that saves lives. The question is whether the institutions that control that data will treat sharing it as an obligation rather than a concession.