The Impact of data and AI on Scientific Production

A Policy Perspective

Barend Mons

Professor Emeritus of Leiden University

Board Member of the GO FAIR foundation

Director of LIFES, the Leiden Initiative for FAIR and Equitable science

Arie Baak

Board member of LIFES Association for FAIR and Equitable science

Koen Jonkers

European Commission, Joint Research Centre (JRC), Brussels, Belgium



Published on July 13th, 2026

Introduction

Artificial Intelligence is transforming how science is conducted, communicated, and validated. It holds genuine promise, enabling pattern recognition at scales no human researcher can match, accelerating workflows, and opening new avenues for discovery. But AI is also fast, not error free, and indiscriminate, while science is time consuming and meticulous by design. The tension between these two realities is already generating serious dysfunction: a flood of AI-generated publications and data, degraded peer review, hallucinated findings, and eroded trust in the scientific record. The skill sets required of future scientists are changing fundamentally, and science policy has not kept pace. Addressing this gap is urgent.

1. AI is Reshaping How Science is Done

Science has moved from a data-sparse to a data-rich environment with remarkable speed. Where researchers once generated modest datasets and communicated findings primarily through human-readable articles, they now operate in an era of exponential data generation. Much of it high-dimensional, distributed, and only tractable by machines. This  trend is however not reflected in actual data reuse. To change that, data itself must now be treated as a first-class scientific output, not a byproduct of articles.

Machines excel at finding patterns in large and complex datasets. But machines have no inherent conceptual model of human-relevant reality. The result is a near-unlimited supply of patterns, many spurious, some misleading. A central challenge for the next generation of scientists will be navigating this pattern jungle: distinguishing meaningful signal from statistical noise. Human oversight is not optional; it is structurally necessary.

Policymakers should also be aware that AI in science extends well beyond Large Language Models, which currently dominate public discourse. The broader landscape includes machine learning for hypothesis generation, automated experimental workflows, and knowledge graph-based reasoning, each carrying distinct risks and opportunities that require tailored governance responses.

2. Data Quality and Infrastructure are the Bottleneck

The quality of AI output in science is directly determined by the quality of data input. FAIR data (Findable, Accessible, Interoperable, and Reusable, and increasingly understood as Fully AI Ready) is the foundational prerequisite for responsible AI-driven science. Without it, AI models are trained on raw, uncurated, potentially biased, or even fabricated content, producing hallucinations that may be life-threatening in clinical contexts and that systematically undermine trust in science. Especially where the reuse of readily findable data is mostly by non-science based actors like the current LLM companies

Equally important is the use of FAIR conceptual models to constrain AI outputs. By providing structured semantic frameworks, where each concept is well defined, how they relate, and what counts as valid inference, these models significantly reduce the risk of hallucination and help ensure that AI-generated findings remain explainable, interpretable and verifiable by humans.

The integrity crisis in scientific publishing compounds these risks. AI systems can now generate plausible-looking articles, data, and reviews almost instantly. One leading AI conference received over 20,000 submissions in 2025 alone. The response (scientists using AI to review AI-generated papers) risks creating a closed noise feedback loop in which flawed systems evaluate one another with no meaningful external check. FAIR data published at the source, where provenance can be verified, is among the most robust tools available to counter this trend.

Underlying all of this is a structural failure in incentives. Researchers are still predominantly rewarded for article count and citation rates, a metric designed for a data-sparse era. Publishing high-quality FAIR data carries little professional recognition and no guaranteed funding coverage. This must change.

3. Governance Gaps are Acute

The technical governance landscape for data in science is in crisis. Decades of dependence on a small number of large cloud providers, combined with legal instruments such as the US Cloud Act, have created a data sovereignty problem that is now acute, particularly in light of recent US policy instability and its effects on internationally shared research infrastructure. Many of the world's core scientific data resources are effectively outside the jurisdictional control of the institutions and communities that generated them.

The FAIR Principles, complemented by the CARE Principles (which address the rights and interests of Indigenous Peoples in data governance) provide a coherent framework for responsible stewardship. Together, they define the conditions under which data should be open or restricted, reusable, and sovereign. The concept of an Internet of FAIR Data and Services (IFDS), in which data remains at its source and is visited by authorised algorithms rather than copied and centralised, offers a practical architectural response. In this model, computational queries travel to distributed data nodes; data does not travel to centralised repositories unless unavoidable.

Equitable participation in AI-driven science also requires that this infrastructure be genuinely accessible. Institutions and communities in the Global South cannot be passive subjects of data reuse. At the other end of the spectrum of connectivity, data intensive industry should not be excluded. Distributed FAIR infrastructure, designed to avoid vendor lock-in and single points of failure, is the mechanism by which meaningful participation becomes possible.

4. What Policymakers Should Do

A coherent policy response requires action on several fronts:

  • Require FAIR data standards in publicly funded research. Funding bodies should make FAIR data publication a condition of grant awards, not a recommendation.

  • Fund data publishing and FAIRification as eligible grant costs. The investment required to generate high-quality, machine-actionable data must be recognised and covered.

  • Make reuse of previously generated data and the supporting infrastructure eligible research costs in grant proposals. The investment required to store and provide high-quality, machine-actionable data must be recognised and covered.

  • Support community-based standards rather than vendor-driven solutions. Open, interoperable infrastructure built on minimal shared standards prevents lock-in and ensures a level playing field.

  • Invest in scientific literacy. Next to scientists, citizens, journalists, and policymakers also require a working understanding of AI's capabilities and limitations. The gap between public perception and technical reality is a governance risk.

  • Mandate provenance and transparency requirements for AI used in research. Any AI system contributing to published scientific findings should be required to disclose training data provenance, model constraints, and limitations.

Conclusion

Science is a global public good. Its integrity, equity, and openness cannot be assumed; they require active policy choices. The window to establish sound governance for AI in scientific production is open, but it will not remain so indefinitely. The decisions made now about data standards, infrastructure, incentives, and oversight will determine whether AI accelerates trustworthy science or systematically undermines it.


Disclaimer: The contents of this publication do not necessarily reflect the position or opinion of the European Commission.

Copyright: © 2026 [author(s)]. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in Frontiers Policy Labs is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.


Please join the conversation and share your perspective.


Previous
Previous

Can science shape its own future?

Next
Next

Modern science: geostrategic, technological, and societal drivers of change