







A better way to evaluate scientific papers. Methodological soundness and transparency, not citation counts.
A budding discipline around how science gets read and judged - Aris
Scholarly publishing spent thirty years arguing about access. It has barely begun arguing about what readers do with the paper once they have it. At Aris we call that second argument scholarly interface design. Mike Morrison's community calls it ScienceUX. Either way it is becoming a field, and here is why it is worth your attention.
The new way we’ll do science
Papers should become human-readable views over a graph of data, tools, results, and certificates.

Are Scientific Papers Bad?
Scientific production in the era of Large Language Models
Large Language Models (LLMs) are rapidly reshaping scientific research. We analyze these changes in multiple, large-scale datasets with 2.1M preprints, 28K peer review reports, and 246M online accesses to scientific documents. We find: 1) scientists adopting LLMs to draft manuscripts demonstrate a large increase in paper production, ranging from 23.7-89.3% depending on scientific field and author background, 2) LLM use has reversed the relationship between writing complexity and paper quality, leading to an influx of manuscripts that are linguistically complex but substantively underwhelming, and 3) LLM adopters access and cite more diverse prior work, including books and younger, less-cited documents. These findings highlight a stunning shift in scientific production that will likely require a change in how journals, funding agencies, and tenure committees evaluate scientific works.

Open Evaluation: A Vision for Entirely Transparent Post-Publication Peer Review and Rating for Science
The two major functions of a scientific publishing system are to provide access to and evaluation of scientific papers. While open access (OA) is becoming a reality, open evaluation (OE), the other side of coin, has received less attention. Evaluation steers the attention of the scientific community and thus the very course of science. It also influences the use of scientific findings in public policy. The current system of scientific publishing provides only journal prestige as an indication of the quality of new papers and relies on a non-transparent and noisy pre-publication peer review process, which delays publication by many months on average. Here I propose an OE system, in which papers are evaluated post-publication in an ongoing fashion by means of open peer review and rating. Through signed ratings and reviews, scientists steer the attention of their field and build their reputation. Reviewers are motivated to be objective, because low-quality or self-serving signed evaluations will negatively impact their reputation. A core feature of this proposal is a division of powers between the accumulation of evaluative evidence and the analysis of this evidence by paper evaluation functions (PEFs). PEFs can be freely defined by individuals or groups (e.g. scientific societies) and provide a plurality of perspectives on the scientific literature. Simple PEFs will use averages of ratings, weighting reviewers (e.g. by H-factor) and rating scales (e.g. by relevance to a decision process) in different ways. Complex PEFs will use advanced statistical techniques to infer the quality of a paper. Papers with initially promising ratings will be more deeply evaluated. The continual refinement of PEFs in response to attempts by individuals to influence evaluations in their own favor will make the system ungameable. OA and OE together have the power to revolutionize scientific publishing and usher in a new culture of transparency, constructive criticism, and collaboration.

Connected Papers | Find and explore academic papers
A unique, visual tool to help researchers and applied scientists find and explore papers relevant to their field of work.

Connected Papers | Find and explore academic papers
A unique, visual tool to help researchers and applied scientists find and explore papers relevant to their field of work.

Science should be machine-readable - Marginal REVOLUTION
One of the leading tasks of our time: We develop a machine-automated approach for extracting results from papers, which we assess via a comprehensive review of the entire eLife corpus. Our method facilitates a direct comparison of machine and peer review, and sheds light on key challenges that must be overcome in order to facilitate […]
A New Paradigm for Scientific Publishing, Peer Review, and Impact Assessment
Scientific publishing and peer review have evolved little in three centuries, while the demands placed on them have grown profoundly. The growing role of artificial intelligence has underscored deep, systemic shortcomings of an aging system that has largely evaded innovation, a system whose origins are appallingly closer to the invention of the printing press than to the internet. We can do better – much better. This article is intended as the beginning of a communal experiment: a living document that critically reviews the modern academic publishing and peer-review system and presents a concrete framework to address what bibliometrics experts¹ have characterized as "the pervasive misapplication of indicators to the evaluation of scientific performance". Building on the Leiden Manifesto, DORA, and a body of scholarship spanning many disciplines and decades, we present a community-governed, non-profit platform organized around three trust-weighted impact factors, for articles, authors, and reviewers, with full algorithmic transparency, an open development log, and structural decoupling of credibility scoring from content moderation and from monetization. We invite the community to discuss, critique, and help shape it.
Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation
Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We introduce a masked-citation task to compare human and LLM-generated citation behavior. For each citation context, an LLM generates a replacement citation sentence, producing a counterfactual corpus directly comparable to human citation. We analyze what, whom, and how models cite, using an LLM-as-a-judge to classify citation intent and a 20-million-edge coauthorship network to measure social distance between cited authors. Across six popular LLMs and 1,746 top NLP conference papers (63k+ contexts, 132k+ citations), three patterns emerge: (1) Compared with human citation, LLMs cite significantly less critically; (2) LLMs over-cite popular and older papers, a tendency amplified for contrasting citations where human writing more often draws on recent, niche work; (3) Whereas humans often cite within their close social network, especially for supporting citations, LLMs tend to draw on more socially distant authors. Together, these differences are double-edged: LLM citation reaches beyond a scholar's close collaborators while being less critical and amplifying visibility bias, reshaping the rhetoric and reach of scientific citation.

Can we measure trust in scientific publications? - LSE Impact
Jonathon Alexis Coates outlines how a constellation of static and dynamic indicators could provide a means for assessing the trustworthiness of published research

Do Science <i>Kardashians</i> Get Citation Premium? Self‐Fulfilling Effects of Social Media on Scientific Impact
ABSTRACT We analyze whether the visibility of scientists on social media affects the number of academic citations. We use the global COVID‐19 pandemic as a quasinatural experiment that exogenously increased public attention and the demand for expertise. Using publications on COVID‐related topics by social media stars and their coauthors prior to the outbreak of the pandemic, we find that social media stars' pre‐COVID‐era papers received about – more citations annually per paper after 2019. Quantitatively comparable results are obtained when we use scientists' Kardashian index (K‐index) as a benchmark for stardom, however we find no significant effects when using the intensive margin of scientists' K‐indexes. We provide a brief discussion of policy implications in light of these findings.

In an era where research evaluation methods are evolving, the Research Contribution Claim Network makes trustworthy tracking of non-traditional research output easy!
In this whitepaper, Patrick Hochstenbach (Ghent University Library), Thomas van Himbergen (SURF), Laurents Sesink (SURF) and Herbert Van de Sompel (DANS) introduce the...

Is there a ‘Goldilocks zone’ for paper length?
A hefty paper published in Nature has raised the question of whether data-dense research tomes can still be digestible.

1/ Scholarly publishing spent thirty years arguing about access. It has barely started arguing about what readers do with the paper once they have it. That second argument is turning into a field, forming around how science gets read and judged. Here is the map as I see it: aris.pub/blog/scholarly-interface-desi…
A budding discipline around how science gets read and judged - Aris
aris.pub