The reproducibility bargain: why open science cannot survive without institutional subsidy
Mandates for transparency have multiplied, but the labour required to maintain reproducible research remains unfunded and unrewarded.
Lena Wolf
Science Reporter
3 September 2026
6 min read
Photo: Unsplash / GlobalTimesOnline
The promise of open science rested on a seductive premise: that making research data and methods freely available would restore public confidence in scientific findings and accelerate discovery. Transparency, its advocates argued, would serve as both disinfectant and catalyst. Yet a decade after major funding bodies began requiring data-sharing plans and code repositories, the infrastructure supporting this vision shows signs of distress. Repositories struggle to maintain servers, datasets languish without documentation, and researchers increasingly treat compliance as a box-ticking exercise rather than a methodological commitment.
The problem is not ideological resistance but structural mismatch. Preparing a dataset for public release demands meticulous work: anonymising sensitive information, writing metadata that makes variables intelligible to outsiders, creating read-me files that explain analytical choices, and responding to queries from users who encounter errors or ambiguities. This labour is iterative and time-consuming, yet grant budgets rarely allocate funds specifically for data curation. When they do, the sums are modest and the positions temporary.
Academic hiring and promotion committees compound the difficulty. A researcher's curriculum vitae gains little from having maintained an exemplary data repository for five years. What counts is publication in prestigious journals, successful grant applications, and increasingly the ability to demonstrate impact beyond the academy. The careful documentation of code, the patient correspondence with users seeking to replicate findings, the unglamorous work of version control—none of this translates into the metrics that govern career progression.
The consequences are predictable. Early-career researchers, facing precarious employment and intense competition for permanent positions, cannot afford to invest substantial effort in activities that hiring panels will overlook. Senior researchers, already stretched across teaching, administration, and their own projects, delegate data management to doctoral students and postdoctoral fellows who lack both training and incentive to do the work thoroughly. The result is a growing archive of nominally open data that is practically unusable: files without codebooks, scripts that assume directory structures unique to one laboratory, datasets stripped of the contextual information needed to interpret them.
Some repositories have attempted to impose quality standards, requiring peer review of datasets before acceptance or employing professional curators to check submissions. These efforts improve usability but exacerbate the resource problem. Curation is skilled labour, and repositories operating on soft money or volunteer effort cannot sustain it. Several prominent domain-specific archives have announced closures or mergers in recent years, citing unsustainable funding models. The work of building trust in shared data, it turns out, requires not just technical infrastructure but ongoing human attention.
Funding agencies have responded to poor compliance by tightening requirements. Grant conditions now often mandate not merely data deposition but the use of specific repositories, adherence to particular metadata standards, and timelines for release. These rules create a compliance burden that falls disproportionately on smaller research groups, which lack dedicated data managers and must divert scientific personnel to administrative tasks. The irony is stark: policies designed to make research more accessible risk making it more difficult for under-resourced teams to compete for funding.
The tension extends to the question of who bears responsibility for long-term maintenance. A dataset deposited today may be cited and reused for decades, but the researcher who created it will move between institutions, retire, or shift focus to new questions. Should universities host data in perpetuity, even when the original investigator has departed? Should funders establish endowments to cover storage and curation costs long after a grant period ends? The default answer has been to rely on a patchwork of institutional repositories, national archives, and non-profit initiatives, none of which commands stable funding.
Advocates for open science often point to the collaborative ethos of fields like genomics, where data-sharing has become normative and infrastructure investments have been substantial. But genomics benefits from a concentration of funding, a relatively standardised set of data types, and a clear economic rationale for sharing: the value of sequence data multiplies when pooled. Many other disciplines lack these advantages. Observational social science, field ecology, and qualitative research generate heterogeneous data that resist standardisation, and the benefits of sharing are diffuse and difficult to quantify.
There is also a more uncomfortable question about whether full reproducibility is always worth the cost. Repeating an analysis from raw data requires not only access to that data but familiarity with the software environment, statistical techniques, and domain knowledge that shaped the original work. Even with perfect documentation, replication is labour-intensive. If the goal is to verify findings, targeted re-analysis of key results may be more efficient than attempting to reproduce every step. If the goal is to build on existing work, researchers often prefer to use summary statistics or processed datasets rather than starting from scratch.
This is not an argument against transparency but a recognition that transparency exists on a spectrum and that different points on that spectrum entail different costs. A minimal standard might require sharing the final analytical dataset and code, with sufficient documentation to understand what was done. A maximal standard might demand raw data, processing scripts, version histories, and detailed lab notebooks. The former is achievable within existing structures; the latter requires dedicated infrastructure and personnel.
The risk is that by mandating the maximal standard without providing the resources to meet it, funding agencies create a culture of performative compliance. Researchers deposit files to satisfy grant conditions but do not invest in making them genuinely reusable. The archive grows, but its contents remain opaque to all but the most determined users. The appearance of openness substitutes for its substance.
Some institutions have begun to experiment with alternative models. Research data management is being integrated into graduate training, with students learning curation practices alongside statistical methods. Universities are hiring data librarians and research software engineers whose roles include supporting open-science workflows. A few funders now allow applicants to request dedicated data management positions as part of grant budgets, recognising that this work cannot be an afterthought.
These initiatives remain limited in scale and uneven in adoption. They also raise questions about professionalisation. If data curation becomes a distinct career path, does it risk creating a two-tier system in which some researchers produce knowledge while others manage its dissemination? Or does it acknowledge that the skills required for each task are different and that both deserve recognition and reward? The answer may depend on whether data professionals are seen as service providers or as collaborators with intellectual stakes in the research.
The technical challenges, meanwhile, continue to evolve. As computational methods become more complex, reproducing an analysis may require not just code but the entire software environment in which it ran: specific versions of libraries, operating system configurations, even hardware specifications. Containerisation and virtual machines offer partial solutions, but they introduce their own maintenance burdens. A container image that works today may not run in five years if dependencies are deprecated or platforms change.
There is a parallel here with the history of scientific publication. Journals were once sustained by subscription revenue and the donated labour of editors and reviewers. The shift to open-access publishing redistributed costs, with some models charging authors fees and others relying on institutional or philanthropic support. The transition has been uneven and contentious, but it reflects a broader recognition that making knowledge freely available requires someone to pay for the infrastructure. Open data and open methods demand a similar reckoning.
The strongest counter-argument is that science managed without formalised data-sharing for centuries and that the emphasis on reproducibility, while valuable, should not overshadow the production of new knowledge. There is a legitimate concern that administrative requirements are consuming time and resources that could be directed toward experimentation and discovery. The question is whether the transparency gains justify the opportunity costs, and the answer may vary across fields and contexts.
What remains clear is that the current model is unsustainable. Mandates without resources produce compliance theatre. Enthusiasm without compensation exploits the goodwill of researchers who believe in open science but cannot afford to prioritise it over career survival. If funding agencies and institutions are serious about transparency, they must treat the labour of making research reproducible as a fundable, rewardable component of scientific work, not as an unfunded mandate or a moral obligation to be discharged in spare time.
The infrastructure of open science is not merely technical. It is social, built on the willingness of researchers to invest effort in work that benefits others more than themselves. That willingness is not infinite, and it cannot be taken for granted. The question facing the research community is whether it will construct incentive structures that sustain this labour or whether the ideal of open science will remain just that: an ideal, admired in principle but abandoned in practice when the costs become too high.
This article was produced with AI assistance and reviewed against our editorial standards.
Lena Wolf
Science Reporter
Lena Wolf translates complex science into accessible, compelling journalism.