







Given the recent growing emphasis on citation justice and related concepts, this paper examines "Citation Analysis," published in Library Trends in 1981, and revisits it through the lens of citation justice. Overarching questions include: How can citation analysis be more just? How can research evaluation go beyond citation analysis to be more just? Sections include a discussion of the concept of citation justice, applications of citation analysis with particular emphasis on evaluative bibliometrics, characterization of assumptions underlying citation analysis, identification of problems posed in dealing with citation data, and an outline of possible approaches to achieving citation justice. Several different entities and actions are discussed with the goal of working toward citation justice. These include author roles, pedagogical approaches, resource compilation, editor and reviewer roles, publisher roles, advocacy, recommendations for research evaluation reform, and higher education institutional roles. Viewing citation analysis through the lens of citation justice reveals significant limitations in citation analysis and suggests ways to correct them–both to ensure that more diverse scholars are part of the scholarly conversation that underlies citation analysis and to encourage approaches to research evaluation that are not dependent solely on citation counts.
Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation
Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We introduce a masked-citation task to compare human and LLM-generated citation behavior. For each citation context, an LLM generates a replacement citation sentence, producing a counterfactual corpus directly comparable to human citation. We analyze what, whom, and how models cite, using an LLM-as-a-judge to classify citation intent and a 20-million-edge coauthorship network to measure social distance between cited authors. Across six popular LLMs and 1,746 top NLP conference papers (63k+ contexts, 132k+ citations), three patterns emerge: (1) Compared with human citation, LLMs cite significantly less critically; (2) LLMs over-cite popular and older papers, a tendency amplified for contrasting citations where human writing more often draws on recent, niche work; (3) Whereas humans often cite within their close social network, especially for supporting citations, LLMs tend to draw on more socially distant authors. Together, these differences are double-edged: LLM citation reaches beyond a scholar's close collaborators while being less critical and amplifying visibility bias, reshaping the rhetoric and reach of scientific citation.


The Research Nexus vision for a more connected scholarly community
Crossref envisions “a rich and reusable open network of relationships connecting research organizations, people, things, and actions; a scholarly record that the global community can build on forever, for the benefit of society”. This Research Nexus expands on the importance of research objects being persistently and uniquely identified. The scholarly community has an established practice of connecting things such as citations to others’ work and it is increasingly critical to identify relationships beyond citations, bringing together published work, unpublished work, institutions, individuals, and identifying the actions that they take e.g., funding, publishing, creating, modifying, citing, and sharing. The Research Nexus brings together metadata and relationships to build a joined-up picture of the scholarly ecosystem and helps everyone identify these relationships and how they change through time. This vision is possible if all parts of the scholarly ecosystem (and beyond) work together, including various scholarly infrastructure organizations.

A New Paradigm for Scientific Publishing, Peer Review, and Impact Assessment
Scientific publishing and peer review have evolved little in three centuries, while the demands placed on them have grown profoundly. The growing role of artificial intelligence has underscored deep, systemic shortcomings of an aging system that has largely evaded innovation, a system whose origins are appallingly closer to the invention of the printing press than to the internet. We can do better – much better. This article is intended as the beginning of a communal experiment: a living document that critically reviews the modern academic publishing and peer-review system and presents a concrete framework to address what bibliometrics experts¹ have characterized as "the pervasive misapplication of indicators to the evaluation of scientific performance". Building on the Leiden Manifesto, DORA, and a body of scholarship spanning many disciplines and decades, we present a community-governed, non-profit platform organized around three trust-weighted impact factors, for articles, authors, and reviewers, with full algorithmic transparency, an open development log, and structural decoupling of credibility scoring from content moderation and from monetization. We invite the community to discuss, critique, and help shape it.
Who Gets Cited? Gender- and Majority-Bias in LLM-Driven Reference Selection
Large language models (LLMs) are rapidly being adopted as research assistants, particularly for literature review and reference recommendation, yet little is known about whether they introduce demographic bias into citation workflows. This study systematically investigates gender bias in LLM-driven reference selection using controlled experiments with pseudonymous author names. We evaluate several LLMs (GPT-4o, GPT-4o-mini, Claude Sonnet, and Claude Haiku) by varying gender composition within candidate reference pools and analyzing selection patterns across fields. Our results reveal two forms of bias: a persistent preference for male-authored references and a majority-group bias that favors whichever gender is more prevalent in the candidate pool. These biases are amplified in larger candidate pools and only modestly attenuated by prompt-based mitigation strategies. Field-level analysis indicates that bias magnitude varies across scientific domains, with social sciences showing the least bias. Our findings indicate that LLMs can reinforce or exacerbate existing gender imbalances in scholarly recognition. Effective mitigation strategies are needed to avoid perpetuating existing gender disparities in scientific citation practices before integrating LLMs into high-stakes academic workflows.

We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic Fields
Natural Language Processing (NLP) is poised to substantially influence the world. However, significant progress comes hand-in-hand with substantial risks. Addressing them requires broad engagement with various fields of study. Yet, little empirical work examines the state of such engagement (past or current). In this paper, we quantify the degree of influence between 23 fields of study and NLP (on each other). We analyzed \textasciitilde77k NLP papers, \textasciitilde3.1m citations from NLP papers to other papers, and \textasciitilde1.8m citations from other papers to NLP papers. We show that, unlike most fields, the cross-field engagement of NLP, measured by our proposed Citation Field Diversity Index (CFDI), has declined from 0.58 in 1980 to 0.31 in 2022 (an all-time low). In addition, we find that NLP has grown more insular—citing increasingly more NLP papers and having fewer papers that act as bridges between fields. NLP citations are dominated by computer science; Less than 8% of NLP citations are to linguistics, and less than 3% are to math and psychology. These findings underscore NLP's urgent need to reflect on its engagement with various fields.
Paperstars
A better way to evaluate scientific papers. Methodological soundness and transparency, not citation counts.

Peer Review at the Crossroads
Peer review has long been regarded as a cornerstone of scholarly communication, ensuring high quality and credibility of published research. Although academic journals trace their origins back three centuries, the procedures for evaluating submissions, particularly peer review, have undergone continuous evolvement. Peer review’s formal institutionalization in the mid-20th century represents a significant, yet natural, phase in this ongoing transformation of scholarly communication. By the early 21st century, there emerged an opinion that the conventional model of peer review faces systematic challenges, including inefficiency, bias, and institutional inertia. The study aims to synthesize the evolution, practices, and outcomes of both conventional and innovative peer review models in scholarly publishing. Through a mixed-methods approach combining interpretative literature review and process modeling (Business Process Model and Notation –BPMN), it identifies four frameworks: pre-publication peer review, registered reports, modular publishing, and the Publish-Review-Curate (PRC) model. While the PRC model, which integrates preprints with post-publication review, demonstrates advantages in transparency and accessibility, no single approach emerges as universally ideal. The choice of model depends on disciplinary context, resource availability, and institutional priorities. The analysis underscores the need for adaptable platforms that enable hybrid workflows, balancing rigor with inclusivity. Future research must address empirical gaps in evaluating these innovations, particularly their long-term impact on equity and epistemic norms.
The Discovery Engine: A Framework for AI-Driven Synthesis and Navigation of Scientific Knowledge Landscapes
Scientific progress relies on the effective accumulation, synthesis, and critical evaluation of knowledge. Traditionally, the well-documented, peer reviewed publication served as the primary standard for filtering and disseminating credible findings within the scientific community. Recently, however, we are witnessing an unprecedented acceleration in research output, a veritable explosion of scientific publications across all disciplines [1]. Yet, this very abundance creates a paradox: the sheer volume threatens to overwhelm the mechanisms designed for its assimilation and synthesis. Researchers, even within highly specialized subfields, face an almost insurmountable challenge in keeping abreast of relevant developments, integrating disparate findings, and identifying the truly novel signals amidst the noise [2]. This information overload contributes to disciplinary fragmentation, hindering the cross-pollination of ideas essential for disruptive innovation [3]. Furthermore, persistent concerns regarding "reproducibility crisis" [2], predatory journals, inflation of research areas[4], growing retractions and the potential influences of bibliometrics on research direction [5] highlight systemic challenges in validating and prioritizing scientific contributions to fundamental knowledge.
Faster science, penalties in evaluation, and concerns on quality and impact: Researchers’ use and perceptions of preprints
The preprint ecosystem has expanded rapidly over the past decade, fundamentally altering science communication. Yet, the scholarly community’s attitudes toward this shift remain underexplored. Through a large-scale survey of US and Canadian biomedical scholars, we provide a comprehensive analysis of preprint utilization, perceived impact, and integration into academic credit systems. We find robust engagement across reading, citing, and submitting preprints; however, this activity is driven primarily by a desire for rapid dissemination rather than a foundational commitment to open science. Furthermore, while preprints are valued as networking assets, perceived career penalties during formal academic evaluations stifle broader cultural adoption. Crucially, to navigate the absence of formal peer review, scholars report a heavy reliance on author reputation as a primary heuristic to evaluate a preprint’s credibility and guide their reading and citation decisions. Notably, despite acknowledging preprints’ role in accelerating knowledge sharing, scholars express significant concerns regarding fraud and misinformation, particularly amid declining public trust in science and emerging threats to scientific integrity from artificial intelligence. To resolve these tensions, the preprint ecosystem must evolve beyond prioritizing speed to foster genuine academic dialogue. Simultaneously, evaluation frameworks must adapt to the realities of preprinting, and innovative quality-control mechanisms are urgently needed to balance rapid dissemination with rigorous scientific integrity.

Fraudulent citations, blamed on AI hallucinations, are becoming more common in research papers
“Fabricated” citations that do not reference real academic papers are spreading in the literature, polluting the public record of science, a new study found

Incorrect Citation Association for Articles in Online-Only Springer Nature Journals
We show that citation metrics of journal articles in many of the online-only Springer Nature journals and associated ones are distorted, going back to articles from 2001. We find that most likely due to an API response error, there are many incorrect references which typically lead to Article Number 1 of a given Volume. Among others, the issue affects journals such as Scientific Reports, Nature Communications, Communications journals, Cell Death & Disease, Light: Science & Applications, as well as many BMC, Discovery and npj journals. Beyond the negative effect of introducing incorrect reference information, this distorts the citation statistics of articles in these journals, with a few articles being massively over-cited compared to their peers, while many lose citations; e.g. both in Scientific Reports and in Nature Communications, 5 of the 10 top cited articles have article numbers of 1. We validate the distorted statistics by assessing data from multiple scientific literature databases: Crossref, OpenCitations, Semantic Scholar, and the journals' websites. The issue primarily arises from the inconsistent transition from page-based referencing of articles to article number-based referencing, as well as the improper handling of the change in the publisher's article metadata API. It seems that the most pressing problem has been present since approximately 2011, which we estimate affects the citation count of millions of authors.

Open Sourcing Scientific Research with Lab Discourse Graphs
Fabricated citations: an audit across 2·5 million biomedical papers
Scientific literature depends on the integrity of its references. Each reference implicitly asserts that a verifiable source exists and supports the claims being made. When references point to non-existent studies, readers, reviewers, and policy makers are unable to evaluate the evidence.

Fabricated citations: an audit across 2·5 million biomedical papers
Scientific literature depends on the integrity of its references. Each reference implicitly asserts that a verifiable source exists and supports the claims being made. When references point to non-existent studies, readers, reviewers, and policy makers are unable to evaluate the evidence.

1/ Scholarly publishing spent thirty years arguing about access. It has barely started arguing about what readers do with the paper once they have it. That second argument is turning into a field, forming around how science gets read and judged. Here is the map as I see it: aris.pub/blog/scholarly-interface-desi…
A budding discipline around how science gets read and judged - Aris
aris.pub