







The peer review system, fundamental to scientific quality control, faces a significant crisis. As journal editors, we often need to send up to 35 invitations just to secure two reviewers, confronting daily the collapse of voluntary participation. This reflects a critical imbalance: while publication pressure intensifies, willingness to evaluate diminishes, creating "literature elephantiasis", i.e., an overwhelming proliferation of papers exceeding human processing capacity. Current compensation models, relying on token recognition and database access, fail to incentivize quality engagement and may encourage ethically problematic practices like excessive self-citation. The unchecked infiltration of artificial intelligence into peer review, with minimal enforcement, further undermines system integrity. We propose transforming peer reviewers into professional referees, modeled on sports officiating. This radical solution involves formal training and certification for reviewers, equipping them to assess scientific merit, methodology, and ethics comprehensively. Like sports referees supported by assistants, scientific referees would collaborate with specialists - including statisticians, methodology experts, and reference checkers - ensuring thorough evaluation while distributing workload effectively. Funding would come from publishers or research funders, recognizing peer review as an essential, compensated component of the research lifecycle. Implementation faces challenges including publisher resistance and funding allocation, which we address through phased transition strategies. This professionalization addresses current inequities where conscientious scientists shoulder disproportionate reviewing burdens while others contribute minimally. Professional reviewers would view evaluation as valued career development rather than unwelcome obligation. Critics citing independence concerns overlook the sports analogy: referees maintain impartiality through professional standards despite league compensation. Quality scientific evaluation requires dedicated expertise, adequate training, and fair remuneration. Science deserves better than a system dependent on goodwill and guilt - it needs professional referees now.
Screening, sorting, and the feedback cycles that imperil peer review
Scholarly journals rely on peer review to identify the science most worthy of publication. Yet finding willing and qualified reviewers to evaluate manuscripts has become an increasingly challenging task, possibly even threatening the long-term viability of peer review as an institution. What can or should be done to salvage it? Here, we develop mathematical models to reveal the intricate interactions among incentives faced by authors, reviewers, and readers in their endeavors to identify the best science. Two facets are particularly salient. First, peer review partially reveals authors’ private sense of their work’s quality through their decisions of where to send their manuscripts. Second, journals’ reliance on traditionally unpaid and largely unrewarded review labor deprives them of a standard market mechanism—wages—to recruit additional reviewers when review labor is in short supply. We highlight a resulting feedback loop that threatens to overwhelm the peer review system: (1) an increase in submissions overtaxes the pool of suitable peer reviewers; (2) the accuracy of review drops because journals must either solicit assistance from less qualified reviewers or ask current reviewers to do more; (3) as review accuracy drops, submissions further increase as more authors try their luck at venues that might otherwise be a stretch. We illustrate how this cycle is propelled by the increasing emphasis on high-impact publications, the proliferation of journals, and competition among these journals for peer reviews. Finally, we suggest interventions that could slow or even reverse this cycle of peer-review meltdown.
A New Paradigm for Scientific Publishing, Peer Review, and Impact Assessment
Scientific publishing and peer review have evolved little in three centuries, while the demands placed on them have grown profoundly. The growing role of artificial intelligence has underscored deep, systemic shortcomings of an aging system that has largely evaded innovation, a system whose origins are appallingly closer to the invention of the printing press than to the internet. We can do better – much better. This article is intended as the beginning of a communal experiment: a living document that critically reviews the modern academic publishing and peer-review system and presents a concrete framework to address what bibliometrics experts¹ have characterized as "the pervasive misapplication of indicators to the evaluation of scientific performance". Building on the Leiden Manifesto, DORA, and a body of scholarship spanning many disciplines and decades, we present a community-governed, non-profit platform organized around three trust-weighted impact factors, for articles, authors, and reviewers, with full algorithmic transparency, an open development log, and structural decoupling of credibility scoring from content moderation and from monetization. We invite the community to discuss, critique, and help shape it.
Reformation of science publishing: the Stockholm Declaration
Science relies on integrity and trustworthiness. But scientists under career pressure are lured to purchase fake publications from ‘paper mills’ that use AI-generated data, text and image fabrication. The number of low-quality or fraudulent publications is rising to hundreds of thousands per year, which—if unchecked—will damage the scientific and economic progress of our societies. The result is editor and reviewer fatigue, irreproducible experiments, misguided experiments, disinformation and escalating costs that devour funding from taxpayers intended for research. It is high time to reevaluate current publishing models and outline a global plan to stop this unhealthy development. A conference was therefore organized by the Royal Swedish Academy of Sciences to draft an action plan with specific recommendations, as follows. (i) Academia should resume control of publishing using non-profit publishing models (e.g. diamond open-access). (ii) Adjust incentive systems to merit quality, not quantity, in a reputation economy where the gaming of publication numbers and citation metrics distorts the perception of academic excellence. (iii) Implement mechanisms to prevent and detect fake publications and fraud which are independent of publishers. (iv) Draft and implement legislations, regulations and policies to increase publishing quality and integrity. This is a call to action for universities, academies, science organizations and funders to unite and join this effort.

Open Evaluation: A Vision for Entirely Transparent Post-Publication Peer Review and Rating for Science
The two major functions of a scientific publishing system are to provide access to and evaluation of scientific papers. While open access (OA) is becoming a reality, open evaluation (OE), the other side of coin, has received less attention. Evaluation steers the attention of the scientific community and thus the very course of science. It also influences the use of scientific findings in public policy. The current system of scientific publishing provides only journal prestige as an indication of the quality of new papers and relies on a non-transparent and noisy pre-publication peer review process, which delays publication by many months on average. Here I propose an OE system, in which papers are evaluated post-publication in an ongoing fashion by means of open peer review and rating. Through signed ratings and reviews, scientists steer the attention of their field and build their reputation. Reviewers are motivated to be objective, because low-quality or self-serving signed evaluations will negatively impact their reputation. A core feature of this proposal is a division of powers between the accumulation of evaluative evidence and the analysis of this evidence by paper evaluation functions (PEFs). PEFs can be freely defined by individuals or groups (e.g. scientific societies) and provide a plurality of perspectives on the scientific literature. Simple PEFs will use averages of ratings, weighting reviewers (e.g. by H-factor) and rating scales (e.g. by relevance to a decision process) in different ways. Complex PEFs will use advanced statistical techniques to infer the quality of a paper. Papers with initially promising ratings will be more deeply evaluated. The continual refinement of PEFs in response to attempts by individuals to influence evaluations in their own favor will make the system ungameable. OA and OE together have the power to revolutionize scientific publishing and usher in a new culture of transparency, constructive criticism, and collaboration.

Faster science, penalties in evaluation, and concerns on quality and impact: Researchers’ use and perceptions of preprints
The preprint ecosystem has expanded rapidly over the past decade, fundamentally altering science communication. Yet, the scholarly community’s attitudes toward this shift remain underexplored. Through a large-scale survey of US and Canadian biomedical scholars, we provide a comprehensive analysis of preprint utilization, perceived impact, and integration into academic credit systems. We find robust engagement across reading, citing, and submitting preprints; however, this activity is driven primarily by a desire for rapid dissemination rather than a foundational commitment to open science. Furthermore, while preprints are valued as networking assets, perceived career penalties during formal academic evaluations stifle broader cultural adoption. Crucially, to navigate the absence of formal peer review, scholars report a heavy reliance on author reputation as a primary heuristic to evaluate a preprint’s credibility and guide their reading and citation decisions. Notably, despite acknowledging preprints’ role in accelerating knowledge sharing, scholars express significant concerns regarding fraud and misinformation, particularly amid declining public trust in science and emerging threats to scientific integrity from artificial intelligence. To resolve these tensions, the preprint ecosystem must evolve beyond prioritizing speed to foster genuine academic dialogue. Simultaneously, evaluation frameworks must adapt to the realities of preprinting, and innovative quality-control mechanisms are urgently needed to balance rapid dissemination with rigorous scientific integrity.

Towards Automating Scientific Review with Google's Paper Assistant Tool
Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis generation to mathematical theorem proving. However, this rapid acceleration is creating a systemic challenge: traditional human peer review cannot scale to match the influx of AI-assisted science. Ultimately, to resolve this tension, we must also deploy AI to accelerate the verification and review process itself. To frame the discussion around this transition, we propose a taxonomy consisting of four progressive levels of AI-human collaboration in scientific evaluation, and discuss various trade-offs involved with each. As a step toward this future, we introduce the Paper Assistant Tool (PAT), an agentic AI framework built for deep scientific review and verification. PAT ingests full scientific manuscripts and produces a comprehensive evaluation, checking theoretical results, validating experiments, suggesting improvements, and identifying potential flaws. By utilizing inference scaling techniques, PAT is able to identify deeper issues than a single model call alone, achieving a 34% improvement over zero-shot recall on mathematical errors in the SPOT benchmark. Pilot deployments of PAT as a pre-submission tool for authors at two major Computer Science conferences -- STOC and ICML -- demonstrate its ability to identify critical errors and suggest substantive improvements to research papers. By catching errors early, PAT eases the cognitive burden placed on referees, while preserving their control over the outcomes of the review process.

You can just review things: A digital ethnography of informal peer review
Across scholarly communities, manuscripts face similar evaluative rituals: editors invite experts to privately assess submissions through formal peer reviews. This closed, loosely structured, and...

You can just review things: A digital ethnography of informal peer review
Across scholarly communities, manuscripts face similar evaluative rituals: editors invite experts to privately assess submissions through formal peer reviews. This closed, loosely structured, and publisher-mediated process is now being supplemented by critiques on open, distributed platforms. We call this practice, a blend of three open peer review variants, informal peer review as it is accessible to outsiders, unmediated by publishers, and conducted across public platforms. Informal peer reviewers range from occasional error detectors to experienced sleuths who identify plagiarism, fraud, errors, conflicts of interest, and conceptual flaws. They may interpret methods, clarify jargon, assess value, and connect to related work. Here, we asked four questions: (1) Who are informal peer reviewers? (2) Where do they work? (3) How do they evaluate research? and (4) What are their impacts? To answer these questions, we conducted a cross-platform digital ethnography with participant observation. We traced discourse across communities over four months and revisited cases after nine and twelve months. From 15 communities, we selected 12 case mentions (10 unique cases) and 8 meta-commentaries from 26 reviewers. Using open and axial coding, we generated 1,080 codes and four themes: reviewers are a motley crew, they self-organize across subpar digital spaces, use deep, uncommon strategies, and they face resistance from authors, publishers, and editors. Informal peer review, we concluded, is a fragile, minimally governed patchwork of people, platforms, and practices, as well as an emerging evidence infrastructure that can be scaled up. We advise advocates and tool-builders to evolve informal review tools, communities, training, and governance by connecting to scholars' values, reducing participation friction, and rewarding attempts to extend the scholarly dialogue.

A billion-dollar donation: estimating the cost of researchers’ time spent on peer review
The amount and value of researchers’ peer review work is critical for academia and journal publishing. However, this labor is under-recognized, its magnitude is unknown, and alternative ways of organizing peer review labor are rarely considered.

AI, peer review and the human activity of science
When researchers cede their scientific judgement to machines, we lose something important.

AI’s Growing Role as Scientific Peer Reviewer | Stanford HAI
Stanford computer scientist James Zou is exploring how AI can accelerate scientific research and peer review. His finding: AI excels at spotting gaps, but judgment calls still need humans.

RegCheck: A tool for structured comparisons between study registrations and papers
Across the social and medical sciences, researchers recognize that specifying planned research activities (i.e., 'registration') prior to the commencement of research has benefits for both the transparency and rigour of science. Despite this, evidence suggests that study registrations frequently go unexamined, minimizing their effectiveness. In a way this is no surprise: manually checking registrations against papers is labour- and time-intensive, requiring careful reading across formats and expertise across domains. The advent of AI unlocks new possibilities in facilitating this activity. We present RegCheck, a modular LLM-assisted tool designed to help researchers, reviewers, and editors from across scientific disciplines compare study registrations with their corresponding papers. Importantly, RegCheck keeps human expertise and judgement in the loop by (i) ensuring that users are the ones who determine which features should be compared, and (ii) presenting the most relevant text associated with each feature to the user, facilitating (rather than replacing) human discrepancy judgements. RegCheck also generates shareable reports with unique RegCheck IDs, enabling them to be easily shared and verified by other users. RegCheck is designed to be adaptable across scientific domains, as well as registration and publication formats. In this paper we provide an overview of the motivation, workflow, and design principles of RegCheck, and we discuss its potential as an extensible infrastructure for reproducible science with an example use case.

Introducing COSIG: The Collection of Open Science Integrity Guides
Investigating the integrity of published scientific papers is key to the scientific process, but the necessary knowledge is in short supply. We present COSIG, an open collection of meta-scientific guides enabling anyone to perform forensic peer review.
How do authors want to use AI for review?
A survey of researchers who compared AI-generated scientific reviews with journal-agnostic human peer review reveals that they overwhelmingly prefer using AI as a self-checking tool before submission rather than as a replacement for human reviewers. It encourages an “author-centric” model in which AI helps researchers improve their manuscripts before they are reviewed by their peers.

Peer Review at the Crossroads
Peer review has long been regarded as a cornerstone of scholarly communication, ensuring high quality and credibility of published research. Although academic journals trace their origins back three centuries, the procedures for evaluating submissions, particularly peer review, have undergone continuous evolvement. Peer review’s formal institutionalization in the mid-20th century represents a significant, yet natural, phase in this ongoing transformation of scholarly communication. By the early 21st century, there emerged an opinion that the conventional model of peer review faces systematic challenges, including inefficiency, bias, and institutional inertia. The study aims to synthesize the evolution, practices, and outcomes of both conventional and innovative peer review models in scholarly publishing. Through a mixed-methods approach combining interpretative literature review and process modeling (Business Process Model and Notation –BPMN), it identifies four frameworks: pre-publication peer review, registered reports, modular publishing, and the Publish-Review-Curate (PRC) model. While the PRC model, which integrates preprints with post-publication review, demonstrates advantages in transparency and accessibility, no single approach emerges as universally ideal. The choice of model depends on disciplinary context, resource availability, and institutional priorities. The analysis underscores the need for adaptable platforms that enable hybrid workflows, balancing rigor with inclusivity. Future research must address empirical gaps in evaluating these innovations, particularly their long-term impact on equity and epistemic norms.
Peer Production in Citizen Science: A Community-Centered Approach on the Example of Personal Science
Citizen science encompasses a wide range of practices where online collaboration for knowledge production plays a significant role. However, the study of forms of online collaboration other than crowdsourcing in citizen science has remained largely unexplored. This thesis aims to fill this gap by investigating peer production as a form of collaboration in online citizen science communities of practice. First, peer production theory was operationalized as a working model and used to analyze collaboration in citizen science case studies. This was followed by a comprehensive participatory design process for a specific use case involving the personal science community of practice. This process resulted in the creation of the “Personal Science Wiki”, an online space for consolidating community knowledge through peer production. Subsequently, a usability and card sorting study identified and resolved issues with the wiki implementation, and provided insights into mental models and content requirements regarding self-research knowledge. The lessons learned from the participatory design process were generalized as process recommendations for designing peer production solutions and knowledge management systems with communities of practice.