







Journal of Cheminformatics -

Incorrect Citation Association for Articles in Online-Only Springer Nature Journals
We show that citation metrics of journal articles in many of the online-only Springer Nature journals and associated ones are distorted, going back to articles from 2001. We find that most likely due to an API response error, there are many incorrect references which typically lead to Article Number 1 of a given Volume. Among others, the issue affects journals such as Scientific Reports, Nature Communications, Communications journals, Cell Death & Disease, Light: Science & Applications, as well as many BMC, Discovery and npj journals. Beyond the negative effect of introducing incorrect reference information, this distorts the citation statistics of articles in these journals, with a few articles being massively over-cited compared to their peers, while many lose citations; e.g. both in Scientific Reports and in Nature Communications, 5 of the 10 top cited articles have article numbers of 1. We validate the distorted statistics by assessing data from multiple scientific literature databases: Crossref, OpenCitations, Semantic Scholar, and the journals' websites. The issue primarily arises from the inconsistent transition from page-based referencing of articles to article number-based referencing, as well as the improper handling of the change in the publisher's article metadata API. It seems that the most pressing problem has been present since approximately 2011, which we estimate affects the citation count of millions of authors.

DITA Open Toolkit
The open-source publishing engine for content authored in the Darwin Information Typing Architecture
Stackable Citation Knowledge — Live Demo: Annotate While Reading (CiTeX 2026)
This screencast is the pre-recorded demo segment from the talk "Stackable Citation Knowledge: Building on Nanopublications for Climate and Biodiversity Research", delivered remotely at the CiTeX 2026 Workshop on Citation Extraction and Parsing (DIPF Leibniz Institute, Frankfurt, 28-29 May 2026) by Anne Fouilloux (LifeWatch ERIC) and Jean Iaquinta (Vitenhub AS). The demo walks through the third production path for citation- typed nanopublications: an expert reader annotates a citation found in someone else's paper, rather than the citing author declaring intent (Author path) or an LLM-extraction pipeline inferring it retroactively (Extract path). The annotation is published as a signed CiTO Citation nanopub on the open Science Live network — ORCID-attributed, immutable via Trusty URI, and contestable by counter-nanopubs. Worked example: while reading Wernberg et al. 2016 (Science, 353:169-172, "Climate-driven regime shift of a temperate marine ecosystem"), the reader infers that Wernberg cites reference 11 - Poloczanska et al. 2013 (Nature Climate Change, 3:919-925, "Global imprint of climate change on marine life") - as the global authority for climate-driven marine species redistribution. The annotation is published to Science Live with relation cito:citesAsAuthority and a one-sentence rationale tying it to the authors' planned Mediterranean Iberian extension. This is the alternative to LLM-based citation intent extraction: same downstream signed nanopub, human expert reader instead of a language model in the loop. Complements OpenCitations, WikiCite, and CiTO-extraction tools (GROBID, CEC, GRAPHIA, OFFZIB) by providing a citable target for any typed-citation output. Tools used: - Science Live Zotero plugin (Zotero 7+, plugin v1.0.6) - Science Live platform (platform.sciencelive4all.org) - Citation with CiTO template (built on CiTO and FaBiO ontologies) Related materials: - Talk repository (Slidev source, demo script, references): https://github.com/ScienceLiveHub/citex2026-stackable-citations - Live deck: https://sciencelive4all.org/citex2026-stackable-citations/ - CiTeX 2026 workshop: https://sites.google.com/view/workshop-on-citation-extractio/
Introducing Citations on the Anthropic API | Claude
Claude can now cite specific passages from your documents, delivering verifiable responses with built-in source tracking. Update: Now available in Amazon Bedrock. (June 30, 2025) Today, we're launching Citations, a new API feature that lets Claude ground its answers in source documents.

A Novel Kuhnian Ontology for Epistemic Classification of STM Scholarly Articles
Despite rapid gains in scale, research evaluation still relies on opaque, lagging proxies. To serve the scientific community, we pursue transparency: reproducible, auditable epistemic classification useful for funding and policy. Here we formalize KGX3 as a scenario-based model for mapping Kuhnian stages from research papers, prove determinism of the classification pipeline, and define the epistemic manifold that yields paradigm maps. We report validation across recent corpora, operational complexity at global scale, and governance that preserves interpretability while protecting core IP. The system delivers early, actionable signals of drift, crisis, and shift unavailable to citation metrics or citations-anchored NLP. KGX3 is the latest iteration of a deterministic epistemic engine developed since 2019—originating as Soph.io (2019–2020), advanced as iKuhn (2020–2024), and field-tested through Preprint Watch.
Crossref: The sustainable source of community-owned scholarly metadata
This paper describes the scholarly metadata collected and made available by Crossref, as well as its importance in the scholarly research ecosystem. Containing over 106 million records and expanding at an average rate of 11% a year, Crossref’s metadata has become one of the major sources of scholarly data for publishers, authors, librarians, funders, and researchers. The metadata set consists of 13 content types, including not only traditional types, such as journals and conference papers, but also data sets, reports, preprints, peer reviews, and grants. The metadata is not limited to basic publication metadata, but can also include abstracts and links to full text, funding and license information, citation links, and the information about corrections, updates, retractions, etc. This scale and breadth make Crossref a valuable source for research in scientometrics, including measuring the growth and impact of science and understanding new trends in scholarly communications. The metadata is available through a number of APIs, including REST API and OAI-PMH. In this paper, we describe the kind of metadata that Crossref provides and how it is collected and curated. We also look at Crossref’s role in the research ecosystem and trends in metadata curation over the years, including the evolution of its citation data provision. We summarize the research used in Crossref’s metadata and describe plans that will improve metadata quality and retrieval in the future.

Semantic Scholar | AI-Powered Research Tool
Semantic Scholar uses groundbreaking AI and engineering to understand the semantics of scientific literature to help Scholars discover relevant research.

The Research Nexus vision for a more connected scholarly community
Crossref envisions “a rich and reusable open network of relationships connecting research organizations, people, things, and actions; a scholarly record that the global community can build on forever, for the benefit of society”. This Research Nexus expands on the importance of research objects being persistently and uniquely identified. The scholarly community has an established practice of connecting things such as citations to others’ work and it is increasingly critical to identify relationships beyond citations, bringing together published work, unpublished work, institutions, individuals, and identifying the actions that they take e.g., funding, publishing, creating, modifying, citing, and sharing. The Research Nexus brings together metadata and relationships to build a joined-up picture of the scholarly ecosystem and helps everyone identify these relationships and how they change through time. This vision is possible if all parts of the scholarly ecosystem (and beyond) work together, including various scholarly infrastructure organizations.

Two billion citation links in Crossref help research travel further - Crossref
We’ve recently reached an important milestone for the research nexus: the works in our metadata corpus are now connected with over 2 billion citation links! This is a great opportunity to share a dedicated dataset and discuss why these are important for science.

Scientific Web Claims: A survey of definitions, tasks, datasets and methods
Scientific web claims are seen as scientific claims as observed on the Web, across social media, online news, and other platforms. The growing prevalence of scientific discussions on the Web has intensified the need to process and assess this specific type of claims. Unlike claims from scientific publications, scientific web claims are expressed in lay terms, are often decontextualized, and typically lack proper citations, which poses unique challenges for their identification, verification, and communication. Nevertheless, the correct processing of scientific web claims is crucial to keeping online science discussions accurate and informed, for instance through fact-checking. This survey provides the first systematic overview dedicated specifically to scientific web claims. We review and compare existing definitions, task formulations, datasets, and methodological approaches across three major perspectives: (1) Scientific fact-checking on the Web, (2) Scientific citations on the Web, and (3) Science communication on the Web. Our interdisciplinary analysis integrates insights from natural language processing, information retrieval, artificial intelligence, social sciences, and science communication. We identify major methodological challenges, including the lack of unified definitions, domain-agnostic corpora, and foundational models tailored to science-related online discourse. We also discuss challenges related to the existing interplay between emotions and distortions of science online. By mapping current research efforts and highlighting open problems, this survey lays the groundwork for developing robust datasets, methods, and evaluation frameworks to advance the automated processing of scientific web claims, a necessary capability for strengthening the reliability of science-related online discourse at scale.
A Science Funding System Beyond the Linear Model
Research funders should embrace a new ontology of scientific work to improve the relationship between science and the public.

Unequal Scientific Recognition in the Age of LLMs
Large language models (LLMs) are reshaping how scientific knowledge is accessed and represented. This study evaluates the extent to which popular and frontier LLMs including GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro recognize scientists, benchmarking their outputs against OpenAlex and Wikipedia. Using a dataset focusing on 100,000 physicists from OpenAlex to evaluate LLM recognition, we uncover substantial disparities: LLMs exhibit selective and inconsistent recognition patterns. Recognition correlates strongly with scholarly impact such as citations, and remains uneven across gender and geography. Women researchers, and researchers from Africa, Asia, and Latin America are significantly underrecognized. We further examine the role of training data provenance, identifying Wikipedia as a potential sources that contributes to recognition gaps. Our findings highlight how LLMs can reflect, and potentially amplify existing disparities in science, underscoring the need for more transparent and inclusive knowledge systems.
Bibliograph — Overview
Bibliograph AT Protocol AppView: procedures and queries served by the net.olamaelcu.livtet.biblio lexicon.
Great survey of the "scholarly interface design" frontier that also highlights the emerging role of the atproto.science and modular research ecosystems - @chive.pub @paperstarsorg.bsky.social @cosmik.network @semble.so @margin.at @discoursegraphs.bsky.social @curvenote.com and many more!
leotrs
1/ Scholarly publishing spent thirty years arguing about access. It has barely started arguing about what readers do with the paper once they have it. That second argument is turning into a field, forming around how science gets read and judged. Here is the map as I see it: aris.pub/blog/scholarly-interface-desi…
Elsevier Developer Portal
Elsevier Developer Portal
Elsevier Developer Portal
About the Indicators API Documentation – World Bank Data Help Desk
API
developer.nytimes.com