







We’ve recently reached an important milestone for the research nexus: the works in our metadata corpus are now connected with over 2 billion citation links! This is a great opportunity to share a dedicated dataset and discuss why these are important for science.
You are Crossref - Crossref
Crossref runs open infrastructure to link research objects, entities, and actions—creating a lasting and reusable scholarly record that underpins open science. Together with our >25,000 members in 167 countries, we drive metadata exchange and support 2.1 billion monthly API queries, facilitating global research communication, for the benefit of society.

Crossref: The sustainable source of community-owned scholarly metadata
This paper describes the scholarly metadata collected and made available by Crossref, as well as its importance in the scholarly research ecosystem. Containing over 106 million records and expanding at an average rate of 11% a year, Crossref’s metadata has become one of the major sources of scholarly data for publishers, authors, librarians, funders, and researchers. The metadata set consists of 13 content types, including not only traditional types, such as journals and conference papers, but also data sets, reports, preprints, peer reviews, and grants. The metadata is not limited to basic publication metadata, but can also include abstracts and links to full text, funding and license information, citation links, and the information about corrections, updates, retractions, etc. This scale and breadth make Crossref a valuable source for research in scientometrics, including measuring the growth and impact of science and understanding new trends in scholarly communications. The metadata is available through a number of APIs, including REST API and OAI-PMH. In this paper, we describe the kind of metadata that Crossref provides and how it is collected and curated. We also look at Crossref’s role in the research ecosystem and trends in metadata curation over the years, including the evolution of its citation data provision. We summarize the research used in Crossref’s metadata and describe plans that will improve metadata quality and retrieval in the future.

The Research Nexus vision for a more connected scholarly community
Crossref envisions “a rich and reusable open network of relationships connecting research organizations, people, things, and actions; a scholarly record that the global community can build on forever, for the benefit of society”. This Research Nexus expands on the importance of research objects being persistently and uniquely identified. The scholarly community has an established practice of connecting things such as citations to others’ work and it is increasingly critical to identify relationships beyond citations, bringing together published work, unpublished work, institutions, individuals, and identifying the actions that they take e.g., funding, publishing, creating, modifying, citing, and sharing. The Research Nexus brings together metadata and relationships to build a joined-up picture of the scholarly ecosystem and helps everyone identify these relationships and how they change through time. This vision is possible if all parts of the scholarly ecosystem (and beyond) work together, including various scholarly infrastructure organizations.

Manuscript submission systems and metadata completeness in Crossref: Patterns and associations
The importance of open research information, particularly publication metadata, is widely recognised. Crossref is one of the most important infrastructures for registering open metadata as part of DOI record registration. It is widely known, however, that the metadata of many publications is far from complete, with many publishers making certain metadata openly available, but failing to do so for other metadata elements. Publishers’ ability to register this metadata with Crossref depends on their capacity to capture and retain this data in their production workflows. Manuscript submission systems are an important, yet largely overlooked, factor in the extent to which publishers make metadata available through Crossref. In this paper, we present the results of an analysis investigating the relation between the level of metadata that publishers deposit with Crossref and the submission systems that they deploy for their journals. We have looked at the 153 publishers with the largest amounts of publications in Crossref and concentrate on the four most commonly used systems: Editorial Manager, ScholarOne, Open Journal Systems (OJS) and eJournalPress. We show that some submission systems appear better suited to capturing certain metadata elements. However, there are always cases where publishers using the same system differ widely in the level of metadata they register, suggesting that technology is not the only prohibiting factor and other considerations are at play.

Incorrect Citation Association for Articles in Online-Only Springer Nature Journals
We show that citation metrics of journal articles in many of the online-only Springer Nature journals and associated ones are distorted, going back to articles from 2001. We find that most likely due to an API response error, there are many incorrect references which typically lead to Article Number 1 of a given Volume. Among others, the issue affects journals such as Scientific Reports, Nature Communications, Communications journals, Cell Death & Disease, Light: Science & Applications, as well as many BMC, Discovery and npj journals. Beyond the negative effect of introducing incorrect reference information, this distorts the citation statistics of articles in these journals, with a few articles being massively over-cited compared to their peers, while many lose citations; e.g. both in Scientific Reports and in Nature Communications, 5 of the 10 top cited articles have article numbers of 1. We validate the distorted statistics by assessing data from multiple scientific literature databases: Crossref, OpenCitations, Semantic Scholar, and the journals' websites. The issue primarily arises from the inconsistent transition from page-based referencing of articles to article number-based referencing, as well as the improper handling of the change in the publisher's article metadata API. It seems that the most pressing problem has been present since approximately 2011, which we estimate affects the citation count of millions of authors.

Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation
Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We introduce a masked-citation task to compare human and LLM-generated citation behavior. For each citation context, an LLM generates a replacement citation sentence, producing a counterfactual corpus directly comparable to human citation. We analyze what, whom, and how models cite, using an LLM-as-a-judge to classify citation intent and a 20-million-edge coauthorship network to measure social distance between cited authors. Across six popular LLMs and 1,746 top NLP conference papers (63k+ contexts, 132k+ citations), three patterns emerge: (1) Compared with human citation, LLMs cite significantly less critically; (2) LLMs over-cite popular and older papers, a tendency amplified for contrasting citations where human writing more often draws on recent, niche work; (3) Whereas humans often cite within their close social network, especially for supporting citations, LLMs tend to draw on more socially distant authors. Together, these differences are double-edged: LLM citation reaches beyond a scholar's close collaborators while being less critical and amplifying visibility bias, reshaping the rhetoric and reach of scientific citation.

Reproducible, citation-aware automated paper reviews @seanjungblluth.bsky.social - ATmosphereConf 20
Evaluating Multilingual Metadata Quality in Crossref
Introduction: Scholarly research spans multiple languages, making multilingual metadata crucial for organizing and accessing knowledge across linguistic boundaries. These multilingual metadata already exist and are propagated throughout scholarly publishing infrastructure, but the extent to which they are correctly recorded, or how they affect metadata quality more broadly is little understood. Methods: Our study quantifies the prevalence of multilingual records across a sample of publisher metadata and offers an understanding of their completeness, quality, and alignment with metadata standards. Utilizing the Crossref API to generate a random sample of 519,665 journal article records, we categorize each record into four distinct language types: English monolingual, non-English monolingual, multilingual, and uncategorized. We then investigate the prevalence of programmatically-detectable errors and the prevalence of multilingual records within the sample to determine whether multilingualism influences the quality of article metadata. Results: We find that English-only records are still in the vast majority among metadata found in Crossref, but that, while non-English and multilingual records present unique challenges, they are not a source of significant metadata quality issues and, in few instances, are more complete or correct than English monolingual records. Discussion & Conclusion: Our findings contribute to discussions surrounding multilingualism in scholarly communication, serving as a resource for researchers, publishers, and information professionals seeking to enhance the global dissemination of knowledge and foster inclusivity in the academic landscape.

Position paper: persistent identifiers in research infrastructure policy - Crossref
PIDs have become central to national and international open research strategies, but identifiers alone cannot deliver the connected, open record that researchers, institutions, funders, publishers, and policymakers depend on. Effective research infrastructure rests on three interdependent elements: open, persistent identifiers; rich, open, and linked metadata; and the sustainable governance and resilient operation of the organisations involved. Crossref urges policymakers to evaluate all three together.

Introducing Citations on the Anthropic API | Claude
Claude can now cite specific passages from your documents, delivering verifiable responses with built-in source tracking. Update: Now available in Amazon Bedrock. (June 30, 2025) Today, we're launching Citations, a new API feature that lets Claude ground its answers in source documents.

We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic Fields
Natural Language Processing (NLP) is poised to substantially influence the world. However, significant progress comes hand-in-hand with substantial risks. Addressing them requires broad engagement with various fields of study. Yet, little empirical work examines the state of such engagement (past or current). In this paper, we quantify the degree of influence between 23 fields of study and NLP (on each other). We analyzed \textasciitilde77k NLP papers, \textasciitilde3.1m citations from NLP papers to other papers, and \textasciitilde1.8m citations from other papers to NLP papers. We show that, unlike most fields, the cross-field engagement of NLP, measured by our proposed Citation Field Diversity Index (CFDI), has declined from 0.58 in 1980 to 0.31 in 2022 (an all-time low). In addition, we find that NLP has grown more insular—citing increasingly more NLP papers and having fewer papers that act as bridges between fields. NLP citations are dominated by computer science; Less than 8% of NLP citations are to linguistics, and less than 3% are to math and psychology. These findings underscore NLP's urgent need to reflect on its engagement with various fields.
Related work | Data Counterfactuals
Adjacent research areas and starting papers, curated in Semble and synced into the site.

Open Sourcing Scientific Research with Lab Discourse Graphs
Holy shit. ITS ALIVE! Meet #Lanyards v0.1 → 'Linktree for Researchers', built in the #ATmosphere 🌀 - Signup with @bsky.app - Link scholarly profiles #ORCID, #ResearchGate, #GoogleScholar etc. - Add research (by DOI) - Add conference talks - and more! #ATproto rules 🙌 #ATscience #AcademicSky