







A comparison of metadata retention law in five eye nations
Canada’s Bill C-22 and the security cost of collecting more data
Canada’s Bill C-22 risks forcing secure services to retain more metadata and build access systems. Tailscale explains why the bill should change.
How Ten Publishers Retract Research
Retractions are the primary mechanism for correcting the scholarly record, yet publishers differ markedly in how they use them. We present a bibliometric analysis of 46,087 retractions across 10 major publishers using data from the Retraction Watch database (1997-2026), examining retraction rates, reasons, temporal trends, and geographic distributions, among other dimensions. Normalized retraction rates vary by two orders of magnitude, from Elsevier's 3.97 per 10,000 publications to Hindawi's 320.02. China-affiliated authors account for the largest share of retractions at every publisher. Retraction lags and reason profiles also vary widely across publishers. Among the ten publishers, ACM is an outlier in its retraction profile. ACM's normalized rate is mid-range (5.65), yet 98.3% of its 354 retractions are related to one incident. Seven of the ten most common global retraction reasons (including misconduct, plagiarism, and data concerns) are entirely absent from ACM's record. ACM's first retraction dates to 2020, despite a catalog dating to 1997. ACM self-describes its retraction threshold as "extremely high." We discuss this threshold in relation to the COPE retraction guidelines and the implications of ACM's non-public dark archive of removed works.

Manuscript submission systems and metadata completeness in Crossref: Patterns and associations
The importance of open research information, particularly publication metadata, is widely recognised. Crossref is one of the most important infrastructures for registering open metadata as part of DOI record registration. It is widely known, however, that the metadata of many publications is far from complete, with many publishers making certain metadata openly available, but failing to do so for other metadata elements. Publishers’ ability to register this metadata with Crossref depends on their capacity to capture and retain this data in their production workflows. Manuscript submission systems are an important, yet largely overlooked, factor in the extent to which publishers make metadata available through Crossref. In this paper, we present the results of an analysis investigating the relation between the level of metadata that publishers deposit with Crossref and the submission systems that they deploy for their journals. We have looked at the 153 publishers with the largest amounts of publications in Crossref and concentrate on the four most commonly used systems: Editorial Manager, ScholarOne, Open Journal Systems (OJS) and eJournalPress. We show that some submission systems appear better suited to capturing certain metadata elements. However, there are always cases where publishers using the same system differ widely in the level of metadata they register, suggesting that technology is not the only prohibiting factor and other considerations are at play.
America Doesn’t Need to Invade Canada. It Has Our Data | The Walrus
From cloud servers to AI patents, digital dependence is becoming a new form of power

Data‐ and code‐archiving in the British Ecological Society journals: Present status and recommendations for future improvements
Abstract Data‐ and code‐archiving are important components of open science, as both make research more transparent, reproducible, accountable and credible, allowing future researchers to build on previous work. Despite progress in implementing data‐ and code‐archiving policies in journals publishing ecology and evolution research, issues remain. To be more useful to future researchers, archived data and code must not only be archived but also meet good practice standards. We collected data from 1861 papers published between 2017 and 2024 in the seven British Ecological Society (BES) journals, during a hackathon event. We systematically checked associated data and/or code, metadata, help files and annotations to assess archiving practices. We determined if and where data and code files were archived, whether they could be located, downloaded and opened, and whether they had associated READMEs, digital object identifiers (DOI) and licences. We also recorded the file extensions used to save data/code files, and which programming languages code was written in. 93% of the 1861 papers we examined used data and ~90% used code. While 97% of the 1735 papers that used data also archived it, only 35% of the 1670 papers that used code also archived code. Over 85% of archived data and code could be located, downloaded and opened. Reusability, however, was more limited; around a third of papers did not have a README or similar to explain their data/code files, and the quality of READMEs varied substantially. We recommend that researchers archive their code and that archived code be explicitly mentioned in the Data (or Code) Availability statement. We also encourage researchers to provide more accessible and informative READMEs for data and code. To help achieve these recommendations, we advocate that journals employ Data/Code editors to review data and code quality, research institutions deliver more training in open science practices, and funding bodies set clear expectations on open data and code practices.

Why the U.S. is noticing this Canadian security bill | CBC News
A Liberal government bill that proposes giving police and spies easier access to information during investigations has fallen into the crosshairs of U.S. tech giants and two American congressional committees, threatening to become the latest irritant in the Canada-U.S. relationship.

Position paper: persistent identifiers in research infrastructure policy - Crossref
PIDs have become central to national and international open research strategies, but identifiers alone cannot deliver the connected, open record that researchers, institutions, funders, publishers, and policymakers depend on. Effective research infrastructure rests on three interdependent elements: open, persistent identifiers; rich, open, and linked metadata; and the sustainable governance and resilient operation of the organisations involved. Crossref urges policymakers to evaluate all three together.

Preservation Initiatives - Canadian Association of Research Libraries
DPC RAM Benchmarking Project CARL’s Digital Preservation Working Group (DPWG) is facilitating a national benchmarking exercise using the Digital Preservation Coalition’s Rapid Assessment Model (DPC RAM). This project fulfills one of the main recommendations from […]

The State of Papers, Retractions, and Preprints: Evidence from the CrossRef Database (2004-2024)
A 20-year analysis of CrossRef metadata demonstrates that global scholarly output -- encompassing publications, retractions, and preprints -- exhibits strikingly inertial growth, well-described by exponential, quadratic, and logistic models with nearly indistinguishable goodness-of-fit. Retraction dynamics, in particular, remain stable and minimally affected by the COVID-19 shock, which contributed less than 1% to total notices. Since 2004, publications doubled every 9.8 years, retractions every 11.4 years, and preprints at the fastest rate, every 5.6 years. The findings underscore a system primed for ongoing stress at unchanged structural bottlenecks. Although model forecasts diverge beyond 2024, the evidence suggests that the future trajectory of scholarly communication will be determined by persistent systemic inertia rather than episodic disruptions -- unless intentionally redirected by policy or AI-driven reform.

IIPC WAC 2024 Presentation: A Conceptual Model of Decentralized Storage for Community-Based Archives
Crossref: The sustainable source of community-owned scholarly metadata
This paper describes the scholarly metadata collected and made available by Crossref, as well as its importance in the scholarly research ecosystem. Containing over 106 million records and expanding at an average rate of 11% a year, Crossref’s metadata has become one of the major sources of scholarly data for publishers, authors, librarians, funders, and researchers. The metadata set consists of 13 content types, including not only traditional types, such as journals and conference papers, but also data sets, reports, preprints, peer reviews, and grants. The metadata is not limited to basic publication metadata, but can also include abstracts and links to full text, funding and license information, citation links, and the information about corrections, updates, retractions, etc. This scale and breadth make Crossref a valuable source for research in scientometrics, including measuring the growth and impact of science and understanding new trends in scholarly communications. The metadata is available through a number of APIs, including REST API and OAI-PMH. In this paper, we describe the kind of metadata that Crossref provides and how it is collected and curated. We also look at Crossref’s role in the research ecosystem and trends in metadata curation over the years, including the evolution of its citation data provision. We summarize the research used in Crossref’s metadata and describe plans that will improve metadata quality and retrieval in the future.

Apple Warns Canada's Bill C-22 Could Force Encryption Backdoors
Apple and Meta have opposed a Canadian bill that the companies say could force them to create backdoor access to encrypted user data, should it pass through the country's parliament. Proposed by Canada's ruling Liberal Party, Bill C-22 contains provisions that could be similar to a UK data access provision order sent to Apple last year, depending on how they are implemented.

Financial Post (@financialpost.com)
For more than 100 years, Canada's most trusted source of financial news
The Lawful Access Two-Headed Surveillance Monster: How Bill C-22 Went Off the Rails - Michael Geist
The government’s plans for lawful access have gone off the rails. In recent days, Signal has warned it would pull out of the Canadian market rather than comply with Bill C-22. Windscribe, the Toronto-headquartered VPN provider, has said it would relocate its headquarters out of Canada and NordVPN has warned it would consider following suit. Apple and Meta have both raised public concerns about the bill’s effect on encryption and cybersecurity. The Canadian Chamber of Commerce, the Cybersecurity Advisors Network, civil liberties groups, and a long line of legal and security experts have all called for changes. The chairs of the U.S. House Judiciary and Foreign Affairs Committees have written to Public Safety Minister Gary Anandasangaree warning that the bill threatens U.S. national security and the integrity of cross-border data flows. Even the bill’s own oversight body, the National Security and Intelligence Review Agency, has told the SECU committee it does not have the access it needs for effective oversight. If the government thought it could push through the bill largely unnoticed, it has been proven painfully wrong as there are now trade frictions with the U.S., the prospect of leading companies exiting the Canadian market, and weaker cybersecurity protections for ordinary users. How did Canada’s lawful access plan go awry so quickly?

The Gander Passport: Why a Sovereign Node Beats a Digital Bunker - Jason Butterfield
Why is Gander connected to Bluesky, and why should you care? Explore the architectural shift of the AT Protocol and how Wingspan is building a "Third Way" for Canadian data residency. This is an early adopter’s take on the infrastructure required for true digital land back and the end of vendor lock-in for the human experience.

Evaluating Multilingual Metadata Quality in Crossref
Introduction: Scholarly research spans multiple languages, making multilingual metadata crucial for organizing and accessing knowledge across linguistic boundaries. These multilingual metadata already exist and are propagated throughout scholarly publishing infrastructure, but the extent to which they are correctly recorded, or how they affect metadata quality more broadly is little understood. Methods: Our study quantifies the prevalence of multilingual records across a sample of publisher metadata and offers an understanding of their completeness, quality, and alignment with metadata standards. Utilizing the Crossref API to generate a random sample of 519,665 journal article records, we categorize each record into four distinct language types: English monolingual, non-English monolingual, multilingual, and uncategorized. We then investigate the prevalence of programmatically-detectable errors and the prevalence of multilingual records within the sample to determine whether multilingualism influences the quality of article metadata. Results: We find that English-only records are still in the vast majority among metadata found in Crossref, but that, while non-English and multilingual records present unique challenges, they are not a source of significant metadata quality issues and, in few instances, are more complete or correct than English monolingual records. Discussion & Conclusion: Our findings contribute to discussions surrounding multilingualism in scholarly communication, serving as a resource for researchers, publishers, and information professionals seeking to enhance the global dissemination of knowledge and foster inclusivity in the academic landscape.
