Breaking Down the Walls: Why Fragmented Data Systems Are Costing American Research Its Competitive Edge
Photo: researchers collaborating around shared data visualization screens in modern university research center, via i.pinimg.com
Imagine two federally funded research teams — one at a major state university in Ohio, another at a private institution in California — spending the better part of three years independently collecting nearly identical datasets on the same biological phenomenon. Neither team is aware of the other's work. Neither has access to a shared repository that would have made such duplication immediately apparent. When both studies are eventually published, the scientific community gains marginally from the redundancy, but the American taxpayer has effectively funded the same work twice.
This is not a hypothetical. Variations of this scenario play out across US research institutions with disconcerting regularity, and the cumulative cost — in dollars, time, and scientific opportunity — is staggering. A 2017 analysis published in PLOS ONE estimated that irreproducibility and related inefficiencies cost the US biomedical research enterprise alone approximately $28 billion annually. Fragmented data infrastructure is not the only driver of that figure, but it is among the most structurally entrenched and, crucially, among the most correctable.
The Siloed Laboratory as a Structural Problem
The tendency toward data siloing in American research is not accidental. It reflects decades of institutional incentives that have rewarded competitive secrecy over collaborative openness. Faculty advancement, grant competitiveness, and journal publishing norms have all historically privileged the researcher who controls unique data over the one who shares it. In this environment, proprietary datasets became professional assets — hoarded rather than circulated, because circulation carried the perceived risk of being scooped.
The infrastructure that grew up around these incentives reflects their logic. Data management systems at individual institutions were built to serve local needs, with little consideration for interoperability with external platforms. Metadata standards vary wildly across disciplines and even within them. File formats, access protocols, and licensing frameworks differ from one repository to the next. The result is a landscape of extraordinary scientific richness that is simultaneously extraordinarily difficult to navigate.
Federal funding agencies have not been innocent bystanders. Despite rhetoric favoring open access and data sharing, grant requirements for data management plans have frequently been treated as compliance exercises rather than genuine accountability mechanisms. The plans are filed; the data is rarely made findable, accessible, interoperable, or reusable in any meaningful sense — the four principles that define what the research community now refers to as FAIR data.
What the International Landscape Reveals
A comparative glance at research infrastructure development abroad should give American policymakers and institutional leaders genuine pause. The European Open Science Cloud, an ambitious initiative backed by the European Commission, represents a coordinated effort to create a federated, cross-border infrastructure for research data across EU member states. While implementation has been uneven and the project faces legitimate criticisms, its ambition and scale reflect a political will to treat scientific data as shared public infrastructure rather than institutional property.
In Australia, the Australian Research Data Commons has made substantial progress in developing discipline-specific data repositories that communicate with one another through standardized metadata frameworks. In China, significant state investment in national research data platforms has created infrastructure that, whatever its limitations in terms of openness and governance, enables a degree of coordination that fragmented American systems struggle to match.
The United States possesses extraordinary scientific assets — world-class universities, deep federal research investment, and an unparalleled concentration of research talent. But assets alone do not determine outcomes. The infrastructure through which those assets are organized and deployed matters enormously, and on that dimension, the US is falling behind in ways that are only beginning to register at the policy level.
Emerging Platforms and Promising Models
There are, however, genuine reasons for optimism. A growing ecosystem of open science platforms has emerged within the United States, demonstrating that interoperable, researcher-friendly data infrastructure is technically achievable. The Open Science Framework, developed by the Center for Open Science in Charlottesville, Virginia, has been adopted by hundreds of thousands of researchers across disciplines and provides a concrete model for how preregistration, data sharing, and collaborative workflow management can be integrated into a single accessible platform.
The NIH's National Center for Advancing Translational Sciences has invested in data commons initiatives that allow biomedical datasets from different funded projects to be queried and analyzed together — a seemingly modest capability that has nonetheless accelerated discoveries that siloed analysis would have delayed by years. The National Science Foundation's investment in the Open Storage Network and related cyberinfrastructure programs similarly reflects a growing recognition that data infrastructure is as fundamental to scientific productivity as laboratory equipment.
What these initiatives share is a commitment to standardization without uniformity — creating the connective tissue that allows diverse repositories to communicate with one another, rather than attempting to consolidate everything into a single monolithic system. This federated model respects the legitimate disciplinary differences in how research data is generated and used while enabling the cross-institutional and cross-disciplinary synthesis that produces the most consequential scientific breakthroughs.
The Policy Imperative
For open science infrastructure to reach its potential in the United States, it requires more than technical solutions. It requires a deliberate realignment of the incentive structures that currently reward data hoarding. Federal funding agencies must move beyond pro forma data management requirements and establish meaningful, enforceable standards for data deposit, documentation, and accessibility. Grant renewal and evaluation processes should explicitly credit researchers for contributions to shared data infrastructure, treating a high-quality, well-documented public dataset as a legitimate scientific output comparable to a peer-reviewed publication.
Universities, for their part, must revisit how they count data sharing contributions in promotion and tenure decisions. The cultural shift required is significant, but it is not without precedent. The transition to open access publishing, while incomplete, has demonstrated that academic norms can change when institutional policies create genuine incentives for doing so.
Philanthropic funders and private foundations also have a role to play. Organizations such as the Gordon and Betty Moore Foundation and the Simons Foundation have already made substantial investments in open science infrastructure; broader engagement from the philanthropic sector could help fill gaps that federal funding cycles are poorly structured to address.
The Stakes of Inaction
The argument for open science infrastructure is sometimes framed as an idealistic appeal to the public good — a vision of science as a commons rather than a competition. That framing is not wrong, but it is incomplete. There is an equally compelling pragmatic case rooted in straightforward competitive logic.
The nations and institutions that build the most capable, most interoperable, most widely adopted research data infrastructure will attract the collaborations, the talent, and ultimately the discoveries that define scientific leadership. The United States has historically led in this domain not by accident but through sustained investment and institutional innovation. Maintaining that leadership in an era of intensifying global competition requires treating data infrastructure as the strategic scientific asset it genuinely is — and acting accordingly before the window for decisive action closes.