BAM Dataset All articles
Agricultural Science

Readable Yesterday, Gone Today: The Silent Crisis of Format Obsolescence in Public Research Data

BAM Dataset
Readable Yesterday, Gone Today: The Silent Crisis of Format Obsolescence in Public Research Data

Photo: Crimson Systems, Public domain, via Wikimedia Commons

There is a particular kind of data loss that leaves no obvious trace. The file is still there. The repository listing still appears in search results. The DOI still resolves. And yet the dataset is, for all practical purposes, gone — encoded in a format that the tools researchers use today cannot reliably open, interpret, or trust.

This is the problem of format obsolescence, and it is quietly hollowing out the open science infrastructure that decades of public investment have been designed to build.

A Problem Hiding in Plain Sight

When researchers and policymakers discuss data preservation, the conversation typically gravitates toward storage — server costs, institutional commitment, the risk of repositories shutting down. These are legitimate concerns. But a dataset that exists on a functioning server in a format no current software can read is not meaningfully preserved. It is archived in the same sense that a cuneiform tablet is archived: technically intact, functionally inaccessible to most of its intended audience.

The scope of the problem is difficult to quantify precisely, which is itself part of the issue. A 2021 audit conducted by researchers at the University of Edinburgh found that a substantial proportion of datasets deposited in major repositories over the preceding two decades used formats that had either been officially discontinued or were no longer supported by their originating software vendors. In agricultural research specifically — a domain that relies heavily on longitudinal datasets tracking soil composition, crop yields, and climate variables over multi-decade periods — the implications are severe. A soil carbon dataset recorded in a mid-1990s version of Quattro Pro, or a pest resistance study stored in a legacy SAS transport file, may be technically present in a federal repository while remaining effectively inaccessible to the graduate student attempting to build on that work today.

Why Institutions Fail to Migrate

The failure to migrate datasets into durable, forward-compatible formats is not primarily a technical problem. The tools and standards to address it have existed for years. Organizations like the Library of Congress, the Digital Preservation Coalition, and the Open Geospatial Consortium have published extensive guidance on preferred archival formats — plain-text CSV files, open-standard XML schemas, HDF5 for large scientific arrays. The knowledge exists. The execution does not follow.

Several structural factors explain the gap. First, data migration is unglamorous work. It generates no publications, attracts no grant funding in its own right, and produces no career advancement for the researchers or data managers who perform it. In an academic environment where incentive structures reward novelty over stewardship, format migration tends to be indefinitely deferred.

Second, many research institutions lack dedicated data curation staff with the technical capacity to identify at-risk formats systematically. A departmental server hosting fifteen years of agricultural field trial data may have no assigned custodian at all — the original principal investigator has retired, the graduate students have dispersed, and the files persist through institutional inertia rather than deliberate management.

Third, the cost of migration is easy to underestimate until it becomes prohibitive. Converting a single well-documented dataset from a legacy format to an archival standard may require only hours of work. Converting thousands of heterogeneous datasets deposited across a decade, many with incomplete or absent metadata, is an undertaking that can exceed the resources available to all but the largest repositories.

The Compounding Cost of Delay

What makes format obsolescence particularly insidious is that it compounds over time in ways that are not linear. A dataset stored in an older Excel binary format (.xls rather than .xlsx) is still recoverable today using a range of conversion tools, though with some risk of fidelity loss. The same dataset left unmigrated for another decade may require specialized forensic software — and may no longer be fully recoverable at any cost.

For agricultural science, where the value of a dataset often increases with its age — long-term soil health records, multi-decade drought response data, historical pesticide application logs — this trajectory is particularly damaging. The datasets most worth preserving are frequently the oldest, and the oldest datasets are disproportionately represented among those stored in obsolete formats.

The downstream consequences extend beyond individual research projects. When systematic reviews and meta-analyses attempt to synthesize evidence across decades of published work, inaccessible underlying data introduces silent gaps. The synthesis appears comprehensive; the omissions are invisible. In a domain like agricultural policy, where federal programs affecting millions of acres and billions of dollars in subsidies are sometimes justified by reference to accumulated scientific evidence, the integrity of that evidence base is not a purely academic concern.

Emerging Responses From Forward-Thinking Repositories

A small number of repositories are beginning to treat format sustainability as a first-order responsibility rather than an afterthought. The approaches vary, but several common strategies are emerging.

Automated format detection at the point of deposit is one of the more promising interventions. Rather than accepting any file a researcher submits, some repositories now run submitted datasets against format registries — such as PRONOM, maintained by the UK National Archives — and flag submissions in formats identified as high-risk for obsolescence. Depositors receive immediate notification and, in some implementations, are offered automated conversion to a recommended archival format before the dataset is accepted into the repository.

Scheduled format audits represent another approach. Rather than treating deposit as a terminal event, some institutions have implemented periodic reviews of their holdings — typically on a five- or ten-year cycle — to identify datasets whose formats have moved from supported to deprecated status since the original deposit. These audits are resource-intensive, but early adopters report that the cost of proactive migration is substantially lower than the cost of emergency recovery when formats become entirely unreadable.

Community-driven format migration, modeled loosely on open-source software maintenance, is a third approach still in early stages. Under this model, researchers who need to access a legacy dataset for their own work contribute the migration effort back to the repository, effectively crowdsourcing the curation burden. The incentive alignment is imperfect — researchers have reason to convert the specific dataset they need, not the thousands they do not — but the model has shown promise in pilot programs at several university data centers.

What Sustained Progress Requires

None of these solutions scales without resources, and resources require recognition that format obsolescence is a genuine threat to the scientific record. That recognition has been slow to develop among the funding agencies and institutional administrators who control the relevant budgets.

The National Science Foundation and the USDA's National Institute of Food and Agriculture both require data management plans as conditions of grant funding. These plans, however, are typically evaluated at the time of application and rarely audited for compliance with format sustainability standards over the life of the award. Strengthening those requirements — and backing them with dedicated funding for curation activities — would represent a meaningful structural intervention.

For the open science mission to be more than aspirational, the files in public repositories must remain genuinely readable by the researchers who need them. Availability without accessibility is an incomplete promise. Addressing format obsolescence is not the most visible challenge in the open data landscape, but it may be among the most consequential — and among those most amenable to systematic, coordinated solutions if the will to pursue them can be sustained.

All Articles

Related Articles

The Toll Gate Remains: Academic Publishers, Preprint Culture, and the Unfinished Business of Open Science

The Toll Gate Remains: Academic Publishers, Preprint Culture, and the Unfinished Business of Open Science

Methodology Under Lock and Key: The Proprietary Protocols Undermining Environmental Science

Methodology Under Lock and Key: The Proprietary Protocols Undermining Environmental Science

From Soil to Satellite: How Public Climate Data Is Leveling the Playing Field for American Farmers

From Soil to Satellite: How Public Climate Data Is Leveling the Playing Field for American Farmers