SemRepo.org
SemRepo is a large-scale RDF knowledge graph of GitHub repositories linked to scientific research. SemRepo captures fine-grained repository-level metadata (e.g., contributors, issues, programming languages) and interlinks this with external scholarly knowledge graphs: repositories to publications in LPWC, repository authors to their profiles in SemOpenAlex, and research artifacts (e.g., datasets, experiments) are linked via MLSea.
Science discussions of retracted articles on Bluesky: public scrutiny or misinformation spreading?
Post-publication peer review (PPPR) has emerged as an important supplement to traditional peer review, with social media playing a growing role in publicising potential problems in published research. However, it remains unclear whether social media discussions of retracted articles primarily reflect good practices, such as exposing flaws and acknowledging retraction status, or bad practices, such as overlooking retractions and continuing to disseminate scientific misinformation. In this study, we collected Bluesky posts referencing scholarly articles from Altmetric and retrieved metadata for the referenced articles using OpenAlex. The final dataset included 284 retracted articles with 79 pre-retraction posts and 857 post-retraction posts, 59 retraction notices with 186 posts, and 609,461 non-retracted articles with 1,344,756 posts. We manually coded Bluesky posts discussing retracted articles to identify instances of good and bad practice. The results show that posts demonstrating good practice (89.9%) substantially outnumbered those demonstrating bad practice (10.1%). Posts reflecting good practice also had more user engagement. In the pre-retraction phase, good practice posts constituted a slight minority (43.0%), whereas in the post-retraction phase they were dominant (94.2%). Most negative posts in the pre-retraction phase (90.0%) had good practice while only 17.3% positive posts in the post-retraction phase showed bad practice. Thus, sentiment analysis can be helpful to filter posts that could flag potential flaws before retraction, but it may struggle to accurately identify the spread of misinformation after retraction. More broadly, this study highlights the potential of Bluesky to support responsible scientific communication, public scrutiny, and research integrity.

Tesseract Academy for the Public Sector - Research, AI & Public Sector Delivery Partner
Research-backed AI, data science, public engagement, survey design, and policy advisory for UK public sector.
Position paper: persistent identifiers in research infrastructure policy - Crossref
PIDs have become central to national and international open research strategies, but identifiers alone cannot deliver the connected, open record that researchers, institutions, funders, publishers, and policymakers depend on. Effective research infrastructure rests on three interdependent elements: open, persistent identifiers; rich, open, and linked metadata; and the sustainable governance and resilient operation of the organisations involved. Crossref urges policymakers to evaluate all three together.

Resilient Data Futures — Discourse Graph
A living, content-addressed, contributable form of the SciOS Resilient Data Futures whitepaper.

In an era where research evaluation methods are evolving, the Research Contribution Claim Network makes trustworthy tracking of non-traditional research output easy!
In this whitepaper, Patrick Hochstenbach (Ghent University Library), Thomas van Himbergen (SURF), Laurents Sesink (SURF) and Herbert Van de Sompel (DANS) introduce the...

AI Aqueducts: Speed. Confidence. And Contaminants.
The AI Pipeline Has No Way to Judge the Science It Uses

The Dataset Friction Framework: measuring user-facing friction as a complement to FAIR
Open research data services have matured to the point where the cost of sustaining them at scale has become a primary design constraint, driving providers to make deliberate choices that may reduce user convenience to keep the service viable. The FAIR (Findable, Accessible, Interoperable, Reuseable) principles describe whether a dataset is well stewarded, and FAIR compliance is often treated as a proxy for usability. FAIR does not capture the cost to a user of finding, accessing, interpreting, and applying a dataset. We introduce the Dataset Friction Framework (DFF) as a complement to FAIR, directly addressing usability. DFF measures user-facing friction across six dimensions, distinguishing engineered friction (deliberate data provider design choices that sustain a service) from accidental friction (defects that require remediation). The framework is validated against 18,556 support tickets from the European Centre for Medium-Range Weather Forecasts (January 2024 to May 2026), which serves 280,000 registered users. Restricting the analysis to tickets raised by external reporters reduces the corpus by 12.3%, but every dimension's internal-staff share falls below this baseline -- confirming that the reported friction signals are genuinely user-facing. We then assess three real datasets across three providers and show that FAIR compliance and DFF friction can disagree in both directions: a 92% FAIR-compliant dataset can still carry substantial friction, and a 42% FAIR score can be an artefact of anti-scraping policy rather than poor stewardship. The two measures are non-redundant and jointly informative: FAIR compliance does not predict DFF friction in either direction. This constitutes the first large-scale empirical application of the framework; cross-institutional validation is identified as the immediate next step.

Who Will Keep Research Data Infrastructure Open and Running?
The scientific community must consider the longevity of open research infrastructure—why it might fail and how to prevent it.

Evaluating Multilingual Metadata Quality in Crossref
Introduction: Scholarly research spans multiple languages, making multilingual metadata crucial for organizing and accessing knowledge across linguistic boundaries. These multilingual metadata already exist and are propagated throughout scholarly publishing infrastructure, but the extent to which they are correctly recorded, or how they affect metadata quality more broadly is little understood. Methods: Our study quantifies the prevalence of multilingual records across a sample of publisher metadata and offers an understanding of their completeness, quality, and alignment with metadata standards. Utilizing the Crossref API to generate a random sample of 519,665 journal article records, we categorize each record into four distinct language types: English monolingual, non-English monolingual, multilingual, and uncategorized. We then investigate the prevalence of programmatically-detectable errors and the prevalence of multilingual records within the sample to determine whether multilingualism influences the quality of article metadata. Results: We find that English-only records are still in the vast majority among metadata found in Crossref, but that, while non-English and multilingual records present unique challenges, they are not a source of significant metadata quality issues and, in few instances, are more complete or correct than English monolingual records. Discussion & Conclusion: Our findings contribute to discussions surrounding multilingualism in scholarly communication, serving as a resource for researchers, publishers, and information professionals seeking to enhance the global dissemination of knowledge and foster inclusivity in the academic landscape.

The Research Nexus vision for a more connected scholarly community
Crossref envisions “a rich and reusable open network of relationships connecting research organizations, people, things, and actions; a scholarly record that the global community can build on forever, for the benefit of society”. This Research Nexus expands on the importance of research objects being persistently and uniquely identified. The scholarly community has an established practice of connecting things such as citations to others’ work and it is increasingly critical to identify relationships beyond citations, bringing together published work, unpublished work, institutions, individuals, and identifying the actions that they take e.g., funding, publishing, creating, modifying, citing, and sharing. The Research Nexus brings together metadata and relationships to build a joined-up picture of the scholarly ecosystem and helps everyone identify these relationships and how they change through time. This vision is possible if all parts of the scholarly ecosystem (and beyond) work together, including various scholarly infrastructure organizations.

The Human Infrastructure of Open Science: Why Mentorship Matters More Than Ever
The transformation of academia into a more open and inclusive ecosystem depends not only on new policies or technologies but also on a fundamental shift in how we nurture early-career researchers through mentorship. This values resilience and integrity as much as productivity. Introduction In the vast, often complex, competitive, and

Creative commons licenses and copyright may not stop academic work being used to train AI - Impact of Social Sciences
Considering the legal standing of creative commons licenses & copyright, Martin Eve suggests legal protections for academic work are unlikely to be forthcoming.

A ten-year drive to credit authors for their work — and why there’s still more to do
Nature - Information about the roles of each author of a paper can help to build trust, integrity and responsible research assessment. Coordinated efforts are needed to consolidate progress.

OSF
Two billion citation links in Crossref help research travel further - Crossref
We’ve recently reached an important milestone for the research nexus: the works in our metadata corpus are now connected with over 2 billion citation links! This is a great opportunity to share a dedicated dataset and discuss why these are important for science.

The Gradual Merging of Repository and CRIS Solutions to Meet Institutional Research Information Management Requirements
Much has been said in recent times about the alleged dichotomy between Institutional Repositories (IRs) and Current Research Information Systems (CRISs). According to this highly ideological argument, IRs would be the platforms to support the non- commercial initiative jointly carried out by HEIs – and specifically their Libraries – in order to freely disseminate their research outputs, whereas CRISs would support the whole institutional research information management (RIM) with special emphasis on projects and funding. RIM being an activity oriented towards reporting for research assessment exercises and thus tightly connected to the institutional funding, the support from the Management at HEIs for CRIS implementation and operation and for the Research Office traditionally in charge of such tasks would be much higher than for the much less relevant IR. Moreover, the awareness of researchers and scholars towards such platforms will usually be much higher for the CRIS – from whose accurate and complete depiction of their research activity their salaries will ultimately depend – and it won’t be unusual to collect complaints on the need to ensure that both systems are simultaneously fed with the appropriate, often duplicated information. According to this conception, it is often hard to get the institutional Research Office and Library to work together for improving the end-user experience by enhancing their system interoperability. While much of this may still be happening at a number of HEIs, the general landscape is swiftly evolving and it's not that accurate anymore to describe the RIM system configuration at institutions in such oversimplified terms. CRIS/IR interoperability is now a fairly widespread feature that will allow both platforms to efficiently exchange information and reinforce each other's features, and especially the borders between what each of these platforms is and does are becoming increasingly blurred. Commercial CRISs are gradually becoming compliant with the OAI-PMH protocol and thus becoming able to offer institutions an integrated repository functionality, while the main open source IR platforms have now developed extended data models that will allow them to deliver features traditionally associated to CRISs such as project and funding management, hence becoming suitable solutions for research institutions where purchasing or developing a highly-sophisticated CRIS is not a top priority. This paper aims to describe the areas where CRIS/IR interoperability is taking place, and will provide a set of use cases for institutional research information system configuration involving IRs, CRISs and a combination of both. These will show how both systems are now increasingly merging for best serving institutions and their researchers.
Towards Modular Open Science
Science is trapped in a paper-shaped box. In this talk at the ATProto community, Rowan Cockett introduces the structural states of scientific components — the floor, the buckets, and the shelf — and makes the case for a second wave of modular science where modularity emerges from packaging rather than being imposed as a precondition. The talk introduces the Open Exchange Architecture (OXA), a new standard for modular and composable scientific content.
No shortcuts to research information citizenship - Digital Science
Being open isn't enough - true "research information citizenship" requires a robust, genuinely open research infrastructure.

Crossref: The sustainable source of community-owned scholarly metadata
Abstract. This paper describes the scholarly metadata collected and made available by Crossref, as well as its importance in the scholarly research ecosystem. Containing over 106 million records and expanding at an average rate of 11% a year, Crossref’s metadata has become one of the major sources of scholarly data for publishers, authors, librarians, funders, and researchers. The metadata set consists of 13 content types, including not only traditional types, such as journals and conference papers, but also data sets, reports, preprints, peer reviews, and grants. The metadata is not limited to basic publication metadata, but can also include abstracts and links to full text, funding and license information, citation links, and the information about corrections, updates, retractions, etc. This scale and breadth make Crossref a valuable source for research in scientometrics, including measuring the growth and impact of science and understanding new trends in scholarly communications. The metadata is available through a number of APIs, including REST API and OAI-PMH. In this paper, we describe the kind of metadata that Crossref provides and how it is collected and curated. We also look at Crossref’s role in the research ecosystem and trends in metadata curation over the years, including the evolution of its citation data provision. We summarize the research used in Crossref’s metadata and describe plans that will improve metadata quality and retrieval in the future.

The research nexus is "a rich and reusable open network of relationships connecting research organisations, people, things, and actions; a scholarly record that the global community can build on forever, for the benefit of society." (crossref.org/documentation/research-nexus/)