Open Science Has a Blueprint but No Foundation: The Infrastructure Gap Holding Researchers Back
Photo: RubinObs/NOIRLab/SLAC/NSF/DOE/AURA/J. Pinto, CC BY 4.0, via Wikimedia Commons
The Rhetoric Outpaces the Reality
Open science has become one of the dominant organizing principles of contemporary American research policy. Federal agencies have issued data-sharing mandates. Universities have established open access repositories. Journals have adopted transparency requirements for code and data. The language of openness now appears in nearly every institutional strategic plan, grant solicitation, and scholarly communication initiative worth noting.
Yet if you ask a working researcher—a postdoc managing a multi-site dataset, a principal investigator trying to make her lab's computational pipeline reproducible, a graduate student attempting to collaborate in real time with a colleague at another institution—the picture that emerges is considerably less optimistic. The tools they actually need, the integrated, reliable, and interoperable systems that would make open science a daily practice rather than an aspirational standard, largely do not exist. What does exist is a patchwork of partial solutions, incompatible platforms, and workarounds that researchers have assembled through ingenuity and frustration in roughly equal measure.
This is not a minor inconvenience. It is a structural failure with real consequences for the pace and quality of scientific progress.
What Researchers Actually Need
To understand the infrastructure gap, it helps to be specific about what researchers require that current tools do not adequately provide.
Consider data sharing. The ideal is straightforward: a researcher completes a study, deposits her data in a repository, and other scholars can access, interrogate, and build upon it with minimal friction. The reality is that most existing repositories—Zenodo, Figshare, institutional repositories, domain-specific archives—serve their individual purposes reasonably well but do not communicate with one another. A researcher whose project draws on datasets from three different repositories must navigate three different access protocols, three different metadata standards, and three different citation conventions. The integration that would make this seamless does not exist.
Or consider version control, a capability that software developers have taken for granted for decades. The underlying logic of version control—tracking changes, enabling collaboration without overwriting, maintaining a clear record of who contributed what and when—is precisely what scientific research requires. Yet the tools built for software development, principally Git and its associated platforms, are poorly adapted to the realities of scientific data. Large binary files, proprietary data formats, and the non-linear nature of experimental research all strain systems designed for code. The result is that most researchers either abandon version control entirely or adopt fragile workarounds that collapse under the weight of real projects.
Real-time collaboration presents similar challenges. The commercial tools that have become standard in other knowledge-work contexts—shared document editors, project management platforms, communication applications—were not designed with scientific workflows in mind. They do not handle structured data, support reproducible analysis pipelines, or integrate with the reference management and manuscript preparation tools that researchers depend on. The consequence is a constant, low-level friction that accumulates over the course of a project into something genuinely costly.
The Policy-Infrastructure Mismatch
The frustrating irony of the current situation is that the policy environment has never been more favorable to open science. The 2022 memorandum from the White House Office of Science and Technology Policy, which required federal agencies to eliminate embargo periods on publicly funded research, represented a significant shift in the institutional landscape. Funding agencies have followed with increasingly specific requirements for data management plans, preregistration, and open access publication.
But policy mandates cannot substitute for infrastructure. Requiring researchers to share their data is meaningless if the systems for sharing it are inadequate. Mandating reproducibility is counterproductive if the tools for achieving reproducibility are inaccessible to anyone outside a small community of computationally sophisticated specialists. What the policy environment has created, in effect, is a set of obligations that researchers are expected to fulfill using tools that were not designed for the purpose.
This mismatch has a predictable outcome: compliance theater. Researchers deposit data in formats that technically satisfy a mandate but are practically unusable by others. They publish code that runs on their specific computational environment but cannot be executed by anyone else. They write data management plans that describe aspirational practices rather than actual ones. The open science mandate is met on paper while its substantive goals go unachieved.
Why Technologists Have Not Filled the Gap
One might reasonably ask why the commercial technology sector, which has demonstrated considerable capacity for building collaborative platforms in other domains, has not addressed this need. The answer involves a combination of market dynamics and domain complexity that helps explain the persistence of the problem.
The academic research market, while large in aggregate, is fragmented by discipline, institution, and funding structure in ways that make it an unattractive target for conventional software development. The needs of a genomics laboratory differ substantially from those of a social science research team, which differ again from those of a digital humanities project. Building a platform that genuinely serves all of these communities requires deep domain knowledge that commercial developers rarely possess and rarely find economical to acquire.
The institutions best positioned to understand researcher needs—universities and research organizations—have historically lacked both the technical capacity and the organizational incentive to build the required infrastructure. Library systems, research computing offices, and scholarly communication units have done admirable work within significant resource constraints, but they have operated largely in isolation from one another, producing the same fragmentation at the institutional level that characterizes the tool landscape more broadly.
What a Real Solution Would Require
Addressing the infrastructure gap will require a level of coordinated investment that has not yet materialized, but the outlines of a solution are discernible.
First, the research community needs interoperability standards developed collaboratively across institutions and disciplines—not imposed by any single vendor or agency, but negotiated through the kind of scholarly consensus-building that has produced durable standards in other technical domains. Without common standards for metadata, access protocols, and data citation, no individual platform can solve the fragmentation problem.
Second, sustained funding for research infrastructure needs to be recognized as a legitimate scientific investment, not an administrative overhead. The National Science Foundation's investment in cyberinfrastructure has been valuable but remains insufficient relative to the scale of the need. Expanding this commitment, and encouraging parallel investment from private foundations and research universities, is a prerequisite for meaningful progress.
Third, the development of research tools must involve researchers as genuine co-designers, not merely as end users consulted after key decisions have been made. The history of academic technology is littered with platforms built by well-intentioned developers who did not fully understand the workflows they were trying to support.
The Opportunity in the Gap
The infrastructure deficit is a problem, but it is also an opportunity. The research community that builds the integrated, interoperable systems that open science requires will not merely improve the efficiency of existing workflows. It will enable forms of collaboration, synthesis, and discovery that are currently impossible. The next significant advance in research productivity is unlikely to come from a methodological breakthrough or a policy reform. It will come from finally building the foundation that the open science movement has always assumed was already in place.