BAM Dataset Open Science. Verified Data. Real Discovery.

BAM Dataset

Open Science. Verified Data. Real Discovery.

Latest Articles

Training Data Is the Science: Why Deleting It After Deployment Undermines Every AI Model Built on It
Medical Research & Policy

Training Data Is the Science: Why Deleting It After Deployment Undermines Every AI Model Built on It

Across research institutions and commercial laboratories, the datasets used to train artificial intelligence systems are being discarded, compressed beyond utility, or sealed behind proprietary agreements shortly after the models they produced go live. Without access to training data, independent auditors, clinicians, and rival researchers cannot meaningfully evaluate what an AI system learned, what it missed, or why it fails when it does. The scientific record is accumulating AI-derived finding

When the Scientist Leaves, the Science Follows: The Institutional Failure Erasing Decades of Research Data
Medical Research & Policy

When the Scientist Leaves, the Science Follows: The Institutional Failure Erasing Decades of Research Data

Across American universities, a silent purge occurs every time a senior researcher retires or departs: decades of experimental data, methodology notes, and irreplaceable datasets quietly disappear alongside them. Institutions have long treated research data as the personal property of the investigator who generated it, a convention that is now extracting a measurable cost from the scientific record. A growing coalition of data stewards, archivists, and early-career researchers is working to reco

Vanishing Acts: Why Well-Meaning Researchers Keep Publishing Computational Science That Cannot Be Rebuilt
Medical Research & Policy

Vanishing Acts: Why Well-Meaning Researchers Keep Publishing Computational Science That Cannot Be Rebuilt

Across disciplines, researchers who genuinely intend to share their work are inadvertently producing studies whose computational cores cannot be reconstructed. The problem is not dishonesty—it is a systemic failure to treat code, parameters, and software environments as archival materials on equal footing with data itself. Understanding why this keeps happening may be the first step toward reversing it.

Still Alive, Already Forgotten: The Quiet Abandonment of Functional Research Datasets
Medical Research & Policy

Still Alive, Already Forgotten: The Quiet Abandonment of Functional Research Datasets

Thousands of research datasets remain technically accessible in institutional repositories long after the projects that created them have concluded—yet they are effectively invisible to the scientific community. Institutional memory loss, expiring grant cycles, and the absence of long-term stewardship protocols conspire to render perfectly functional data unreachable. Understanding why this happens, and what a small number of forward-thinking institutions are doing to reverse it, is among the mo

Built on Vapor: The Orphaned AI Models Running on Datasets No One Can Find
Medical Research & Policy

Built on Vapor: The Orphaned AI Models Running on Datasets No One Can Find

Across American research institutions, AI systems trained on datasets that no longer exist—or were never properly preserved—are being deployed in consequential scientific and clinical settings. When the underlying data vanishes, so does any meaningful capacity for audit, challenge, or accountability. The speed of AI development has created a generation of black-box models with phantom foundations.

Haunted by Citation: The Broken Reference Chains Corrupting the Scientific Record
Medical Research & Policy

Haunted by Citation: The Broken Reference Chains Corrupting the Scientific Record

When foundational datasets vanish from public repositories, the citations pointing to them do not disappear alongside them. Researchers tracing these phantom references are discovering that broken data links propagate through the literature for decades, lending false authority to findings that can no longer be independently verified.

Stripped of Meaning: How Poor Metadata Practices Are Quietly Hollowing Out Scientific Datasets
Medical Research & Policy

Stripped of Meaning: How Poor Metadata Practices Are Quietly Hollowing Out Scientific Datasets

Across disciplines, published datasets are arriving in public repositories with metadata so sparse or imprecise that the data itself becomes scientifically inert. When collection methods, sampling constraints, and known biases go undocumented, other researchers cannot responsibly reuse what they find. This article examines how institutional pressures are driving the problem and what concrete standards could reverse the damage.

Declared but Unreachable: The Quiet Collapse of Data Availability Statements in Published Research
Medical Research & Policy

Declared but Unreachable: The Quiet Collapse of Data Availability Statements in Published Research

Thousands of peer-reviewed studies contain data availability statements that point researchers toward files that no longer exist, links that have long since expired, or repositories that were never properly populated. The gap between what journals require authors to declare and what those authors actually deliver has grown into a systemic failure — one that quietly undermines the reproducibility of published science across nearly every discipline.

Fabricated Foundations: When Synthetic Data Enters the Scientific Record Unannounced
Medical Research & Policy

Fabricated Foundations: When Synthetic Data Enters the Scientific Record Unannounced

Researchers are increasingly supplementing experimental work with AI-generated datasets, often without disclosing the substitution to peer reviewers or the broader scientific community. The consequences for reproducibility are significant and, in some fields, already measurable. Open science repositories must now grapple with whether synthetic data can ever be treated as equivalent to empirically gathered evidence.

Ink That Outlasted the Server: The Quiet Revival of Forgotten Laboratory Notebooks
Medical Research & Policy

Ink That Outlasted the Server: The Quiet Revival of Forgotten Laboratory Notebooks

Across American research institutions, archivists and scientists are discovering that handwritten laboratory notebooks from decades past have survived more intact than the digital files that were supposed to replace them. These rediscovered records are prompting serious questions about reproducibility, institutional memory, and what it means to preserve scientific knowledge for the long term. The findings challenge foundational assumptions about which documentation formats actually serve science

Cited Into Thin Air: The Growing Problem of Scientific Datasets That Exist Only on Paper
Medical Research & Policy

Cited Into Thin Air: The Growing Problem of Scientific Datasets That Exist Only on Paper

Across academic literature, thousands of datasets are formally cited in published research but are effectively unreachable—deleted, misfiled, or never properly deposited in the first place. This phantom layer of scientific evidence quietly corrupts the citation chains that peer review depends on, leaving researchers unable to verify foundational claims. The problem is more widespread than the scientific community has publicly acknowledged.

Readable Yesterday, Gone Today: The Silent Crisis of Format Obsolescence in Public Research Data
Agricultural Science

Readable Yesterday, Gone Today: The Silent Crisis of Format Obsolescence in Public Research Data

Thousands of publicly funded research datasets sit in open-access repositories today, technically available yet practically unreachable — locked inside file formats that modern software can no longer parse. From early Excel workbooks to discontinued statistical packages, format decay is erasing scientific knowledge without a single record being deleted. The problem is structural, underappreciated, and accelerating.

Paper Promises: Why Federal Data-Sharing Mandates Rarely Survive Contact With Reality
Medical Research & Policy

Paper Promises: Why Federal Data-Sharing Mandates Rarely Survive Contact With Reality

NIH and NSF now require data management plans as a condition of federal funding, yet follow-through remains largely ceremonial. A close examination of grant administration practices reveals a system in which non-compliance carries little consequence, and the data those mandates were designed to protect quietly disappears.

Funded Once, Gone Forever: The Quiet Crisis of Expiring Scientific Data
Medical Research & Policy

Funded Once, Gone Forever: The Quiet Crisis of Expiring Scientific Data

Across American research institutions, datasets assembled over years or decades—and paid for by public dollars—are quietly disappearing due to storage costs, shifting institutional priorities, and the absence of binding preservation requirements. The consequences extend far beyond inconvenience: lost longitudinal records, vanished environmental baselines, and clinical trial data that cannot be independently verified. This article examines the structural failures driving scientific data expiratio

The Toll Gate Remains: Academic Publishers, Preprint Culture, and the Unfinished Business of Open Science
Agricultural Science

The Toll Gate Remains: Academic Publishers, Preprint Culture, and the Unfinished Business of Open Science

Despite the rapid expansion of preprint servers and a decade of open-access mandates from federal funding agencies, the major commercial academic publishers continue to extract substantial revenue from publicly funded research—often while publicly endorsing the principles of data transparency. An examination of the economics of scholarly publishing, combined with case studies from researchers who have deliberately routed around traditional journals, reveals a system in which structural incentive

Lost Before They Can Be Found: The Silent Erosion of Scientific Datasets After Publication
Medical Research & Policy

Lost Before They Can Be Found: The Silent Erosion of Scientific Datasets After Publication

Millions of dollars in publicly funded research vanish quietly each year—not through fraud or negligence, but through the mundane failures of broken hyperlinks, expired hosting contracts, and researcher departures that leave datasets stranded. A growing body of evidence suggests the scientific community is losing data faster than it can be replaced, undermining the cumulative nature of empirical inquiry. Archivists, librarians, and open-data advocates are now racing to document the scale of this

Methodology Under Lock and Key: The Proprietary Protocols Undermining Environmental Science
Agricultural Science

Methodology Under Lock and Key: The Proprietary Protocols Undermining Environmental Science

Across the United States, critical decisions about air quality, water safety, and soil contamination are being made on the basis of environmental studies that independent scientists cannot fully scrutinize. When the methods used to collect and analyze environmental data remain proprietary, the scientific foundation beneath public health policy quietly erodes. This investigation examines the structural incentives keeping environmental monitoring methodology behind closed doors—and the researchers

Rewiring the Incentives: How a New Generation of Universities Is Making Data Sharing a Career Asset
Medical Research & Policy

Rewiring the Incentives: How a New Generation of Universities Is Making Data Sharing a Career Asset

Tenure committees have long rewarded publications in prestigious journals while treating the release of underlying data as an afterthought. A small but growing cohort of American universities is dismantling that framework, embedding data transparency directly into the criteria by which researchers are hired, promoted, and funded. The institutions pioneering this shift offer an early look at what open science looks like when it is structurally incentivized rather than merely encouraged.

Trained to Miss: How Rare Disease Patients Are Being Systematically Excluded From Medical AI
Medical Research & Policy

Trained to Miss: How Rare Disease Patients Are Being Systematically Excluded From Medical AI

Machine learning models built on population-scale genomic datasets are producing diagnostic tools that work well for common conditions but fail patients with rare genetic disorders. The structural incentives that govern data contribution to open repositories help explain why this gap persists — and why the patients who most need AI-assisted diagnosis are the least likely to receive it.

Replication as Resistance: Independent Researchers Are Auditing Oncology Science — and Finding It Wants
Medical Research & Policy

Replication as Resistance: Independent Researchers Are Auditing Oncology Science — and Finding It Wants

A growing cohort of academic researchers and independent scientists is using publicly available datasets to reproduce proprietary cancer studies — and in doing so, surfacing errors, methodological inconsistencies, and overlooked findings that original publishers never corrected. The movement represents a fundamental challenge to the closed-access model that has long governed oncology research, and its results are beginning to influence how treatments are evaluated and approved.