BAM Dataset All articles
Medical Research & Policy

Haunted by Citation: The Broken Reference Chains Corrupting the Scientific Record

BAM Dataset
Haunted by Citation: The Broken Reference Chains Corrupting the Scientific Record

In academic science, a citation is meant to function as a receipt — a verifiable record of evidence that any reader can retrieve and examine. That assumption quietly breaks down when the underlying dataset no longer exists. What remains is something more troubling than a missing source: it is a reference that continues to generate trust, accumulate downstream citations, and shape research directions long after the data it describes has become permanently inaccessible.

This phenomenon, increasingly recognized among data librarians and reproducibility researchers, has no universally agreed-upon name. Some call it the ghost citation problem. Others describe it as citation laundering — a process by which broken links to vanished data gradually acquire the appearance of legitimate evidence simply through repetition. Whatever the terminology, the consequences for scientific integrity are significant and largely unaddressed.

The Mechanics of Disappearance

Datasets vanish for reasons that are rarely dramatic. A university server is decommissioned. A federal contract expires and the hosting arrangement ends. A principal investigator retires, and the institutional memory of where files were stored retires with them. A journal changes ownership and legacy supplementary materials are not migrated. In each case, the disappearance is administrative rather than scientific — but the effect on the literature is identical.

Once a dataset is gone, the papers that cited it remain. Those papers, in turn, are cited by subsequent studies. Each layer of citation adds apparent credibility to the original claim, even though the foundational evidence is no longer available for inspection. A 2021 review of biomedical literature found that a substantial proportion of dataset citations in high-impact journals led to broken or inaccessible links, yet those papers continued to accumulate forward citations at rates comparable to papers with fully accessible data.

The problem compounds in fields where a small number of large, foundational datasets anchor entire research programs. In epidemiology, nutrition science, and behavioral medicine — areas where replication is already difficult and contested — a handful of widely cited datasets can define the evidentiary basis for clinical guidelines, public health recommendations, and drug approvals. When those datasets become inaccessible, the authority they conferred does not diminish. It simply becomes unverifiable.

Tracing the Chain

Researchers who have attempted to audit citation chains in specific subfields describe a process that is equal parts archival detective work and institutional negotiation. A team studying cardiovascular risk factors at a large public university recently spent several months attempting to locate the primary dataset underlying a frequently cited meta-analysis published in the early 2000s. The journal listed a corresponding author whose institutional email address had been inactive for years. The university's data management office had no record of the dataset. A request to the journal's current editorial team produced a form response directing the researchers to a supplementary materials page that returned a 404 error.

The meta-analysis itself had been cited more than 800 times. None of those citing papers noted that the underlying data was unavailable. Several had described the original dataset as a "robust and publicly accessible" resource.

This kind of audit is painstaking precisely because the infrastructure for tracking dataset availability over time has never been systematically built. Digital object identifiers, or DOIs, were designed to provide persistent links to research outputs, but DOI maintenance depends on the continued participation of the issuing institution. When institutions stop paying maintenance fees or simply cease to exist, DOIs resolve to nothing — yet the citations referencing them remain formally intact in databases like PubMed and Web of Science.

Institutional Failures and Structural Gaps

The persistence of phantom citations reflects a series of overlapping institutional failures. Journals bear some responsibility: most do not verify data availability at the time of submission, and virtually none conduct post-publication audits to confirm that cited datasets remain accessible. The burden of verification falls entirely on readers, who typically lack the time, resources, or institutional access to perform systematic checks.

Funding agencies have begun to require data management plans as a condition of grant awards, a policy shift that has produced modest improvements in initial data deposit rates. But deposit is not the same as preservation. A dataset uploaded to a repository at the time of publication may be deleted, migrated without notice, or rendered unreadable by format obsolescence within a decade. Federal mandates have not yet established enforceable standards for long-term accessibility, and the agencies responsible for oversight rarely have mechanisms to detect or respond to post-publication data loss.

Universities, for their part, have historically treated research data as the intellectual property of individual investigators rather than as institutional assets requiring active stewardship. That model creates obvious vulnerabilities when investigators leave, retire, or die. Some research universities have begun to establish centralized data repositories with explicit preservation commitments, but adoption is uneven and coverage is far from comprehensive.

The Downstream Consequences

The practical consequences of ghost citations extend beyond abstract concerns about reproducibility. In medical research and public health policy, findings that rest on inaccessible data cannot be subjected to the independent scrutiny that scientific consensus requires. When regulatory bodies, clinical guideline committees, or systematic review authors incorporate studies whose underlying data cannot be examined, they are building on foundations they cannot inspect.

This matters particularly in areas where original findings have been contested or where subsequent research has produced conflicting results. If the dataset underlying a contested finding cannot be retrieved and reanalyzed, the dispute cannot be resolved through evidence. The original citation simply continues to accumulate authority by default.

Open science advocates argue that mandatory data deposit in certified, long-term repositories — combined with persistent identifier standards and post-publication availability monitoring — would substantially reduce the rate at which datasets disappear. Platforms committed to verified, open-access data infrastructure represent one model for how the research community might begin to address this gap systematically rather than case by case.

A Record That Remembers What No Longer Exists

There is something structurally perverse about a citation system that preserves the appearance of evidence long after the evidence itself has ceased to exist. The scientific literature is supposed to function as a cumulative, self-correcting record — one in which claims are traceable, verifiable, and open to challenge. Ghost citations undermine that function not through fraud or misconduct, but through neglect and institutional indifference.

Addressing the problem will require changes at multiple levels: journal policies that treat post-publication data availability as an ongoing editorial responsibility, funding mandates that extend beyond deposit to genuine long-term preservation, and infrastructure investments that make verified, persistent data access the norm rather than the exception. Until those changes are in place, the scientific record will continue to be haunted by references to evidence that exists only on paper — cited into permanence, but nowhere to be found.

All Articles

Related Articles

Stripped of Meaning: How Poor Metadata Practices Are Quietly Hollowing Out Scientific Datasets

Stripped of Meaning: How Poor Metadata Practices Are Quietly Hollowing Out Scientific Datasets

Declared but Unreachable: The Quiet Collapse of Data Availability Statements in Published Research

Declared but Unreachable: The Quiet Collapse of Data Availability Statements in Published Research

Fabricated Foundations: When Synthetic Data Enters the Scientific Record Unannounced

Fabricated Foundations: When Synthetic Data Enters the Scientific Record Unannounced