The Informal Archive: PhD Students Stepping In Where Institutions Have Stepped Out
In the spring of 2022, a fourth-year doctoral student in epidemiology at a large public university in the Southeast found herself in an unusual position. Her advisor had accepted an administrative role at another institution, and the lab — along with its decade-long longitudinal dataset tracking chronic disease outcomes in rural communities — was effectively dissolving. No formal data transfer plan existed. No institutional repository had been designated. The dataset, assembled through years of patient recruitment, clinical measurement, and painstaking quality control, was stored across a combination of university servers, external hard drives, and a cloud storage account tied to a personal email address that would expire with the faculty appointment.
She spent the next four months doing what no one had asked her to do and no one was paying her to do: documenting the dataset's structure, writing codebooks from memory and scattered notes, transferring files to a university-affiliated repository, and emailing former lab members to reconstruct the provenance of variables whose origins had never been formally recorded.
"If I had not done it, it would have been gone," she said. "Not dramatically gone — just quietly unreachable. The files might still exist somewhere, but no one would ever be able to use them."
Her experience is not exceptional. It is, according to researchers who study data preservation and scientific infrastructure, increasingly representative of how a significant portion of American research data survives at all.
The Structural Gap Graduate Students Are Filling
Federal data-sharing mandates, journal data availability policies, and institutional open science initiatives have collectively created the impression that research data is increasingly well-managed and systematically preserved. The reality, as documented in multiple studies of research data practices, is considerably more complicated. Policies mandate sharing. They rarely mandate the sustained curation, documentation, and maintenance that makes shared data actually usable. And they almost never specify what happens to data when the people who created it move on.
Labor economists who study academic research have a term for the kind of work that fills this gap: invisible infrastructure labor. It is the work that makes other work possible, that is rarely credited in publications, and that is disproportionately performed by people at the lower end of academic hierarchies — graduate students, postdocs, and research staff on short-term contracts.
In the domain of research data preservation, this invisible labor has taken on a particular character in recent years. As university libraries have reduced specialized data curation staff, as NIH and NSF data management plan requirements have generated paperwork without always generating practice, and as the pace of lab turnover has accelerated in the post-pandemic academic labor market, graduate students have increasingly found themselves as the last line of defense between functional datasets and permanent inaccessibility.
Networks Without Names
What has emerged in some corners of the research community is something that resists easy categorization: informal networks of early-career researchers who share knowledge, tools, and strategies for data rescue work that their institutions do not formally recognize.
These networks operate through discipline-specific Slack workspaces, mailing lists attached to professional society early-career committees, and word-of-mouth referrals between graduate students at different institutions who have encountered similar situations. They share documentation templates, repository submission guides, and — perhaps most valuably — moral support for work that can feel both urgent and professionally thankless.
Jordan Eklund, a postdoctoral researcher in genomics who has been involved in one such informal network for two years, described the community's function in direct terms: "We are essentially doing triage. We cannot save everything. We try to identify what is irreplaceable — longitudinal cohorts, rare population samples, datasets that took years to build and cannot be reconstructed — and we prioritize those."
The triage metaphor is apt. The volume of data at risk of loss in any given academic year is vastly larger than any informal volunteer network can address. The choices these researchers make about what to prioritize — and what to allow to disappear — are consequential scientific decisions being made without institutional guidance, without formal authority, and without any systematic mechanism for tracking what has been lost.
The Medical Research Stakes
The domain where the consequences of informal data stewardship are most acute is medical and clinical research. Longitudinal health datasets — the kind that track patient populations over years or decades — are among the most scientifically valuable and practically irreplaceable assets in biomedical science. They are also among the most vulnerable to the kind of institutional disruption that triggers data loss.
A clinical researcher who retires without a formal succession plan, a lab that loses its NIH funding without completing a data transfer protocol, a hospital system that changes electronic health record platforms without migrating legacy research extracts — each of these scenarios has produced documented instances of dataset loss that could not subsequently be reconstructed. In medical research, where patient recruitment alone can take years and where the scientific value of a dataset often increases with time, these losses are not recoverable.
The graduate students and postdocs who intervene in these situations are not doing so because they have been trained in data curation or because their institutions have given them authority to act. They are doing so because they understand what the data represents and because, in many cases, they are the only people still present who do.
Credit, Compensation, and Continuity
The sustainability problem is straightforward. Data rescue work performed by graduate students and postdocs is labor-intensive, technically demanding, and professionally unrewarded by the metrics that govern academic career advancement. Publications, grants, and teaching evaluations determine whether a PhD student gets a faculty position. The number of orphaned datasets they have rescued and documented does not appear on a CV in any form that hiring committees recognize.
Several research libraries and data repositories have begun experimenting with formal recognition mechanisms — certificates, co-authorship credits on repository records, letters of support for fellowship applications — that attempt to make this labor visible. These efforts are meaningful but limited. They address recognition without addressing compensation, and they do not change the underlying structural reality that institutions have offloaded a core scientific responsibility onto the most precarious members of their research communities.
A more durable solution would require universities to treat data stewardship as an institutional function with dedicated staffing and sustained funding — not as an emergency measure triggered by crisis, and not as volunteer work performed by people who cannot afford to say no. Several peer institutions in Europe have moved in this direction, embedding professional data curators within research groups rather than centralizing curation in libraries that are often too removed from active research to intervene effectively.
What Is Being Saved, and What Is Not
The informal networks doing this work are, by their own account, succeeding in some cases and failing in many others. The datasets that get rescued tend to be those where a graduate student or postdoc happens to be present, happens to recognize the risk, and happens to have both the technical skills and the time to intervene. The datasets that disappear tend to be those where none of those conditions are met — which, given the scale of research activity in the United States, is most of them.
For a scientific enterprise that depends on cumulative knowledge, on the ability to revisit earlier findings with new methods, and on the long-run value of data collected at significant public expense, this is not an acceptable equilibrium. The informal archive being maintained by early-career researchers is a testament to their commitment to science. It should not have to be.