







RO-Crate is a community effort to establish a lightweight approach to packaging research data with their metadata. It is based on schema.org annotations in JSON-LD, and aims to make best-practice in formal metadata description accessible and practical for use in a wider variety of situations, from an individual researcher working with a folder of data, to large data-intensive computational research environments.
About RO-Crate | Research Object Crate (RO-Crate)
What RO-Crate is, how RO-Crate makes research FAIR, and how to get started
crate.social — publish structured records to your ATProto PDS
Crate is a custom-lexicon publishing service for ATProto. Define your record types, import from anywhere, query from everywhere.
Open Exchange Architecture
A specification for representing scientific documents and their components as structured objects.
Open Exchange Architecture
A specification for representing scientific documents and their components as structured objects.

Open Exchange Architecture
A specification for representing scientific documents and their components as structured objects.
Scientific Documents as First-Class Objects on AT Protocol
OXA now defines an AT Protocol lexicon for publishing scientific documents to the Atmosphere. The `pub.oxa.*` namespace lets documents live in any Personal Data Server, making scientific content user-owned, portable, and discoverable alongside social interactions, feeds, and moderation infrastructure.
Scientific Documents as First-Class Objects on AT Protocol
OXA now defines an AT Protocol lexicon for publishing scientific documents to the Atmosphere. The `pub.oxa.*` namespace lets documents live in any Personal Data Server, making scientific content user-owned, portable, and discoverable alongside social interactions, feeds, and moderation infrastructure.
atproto-crates — AT Protocol building blocks for Rust
A Rust workspace of eighteen crates for AT Protocol: content addressing, DID and handle resolution, repositories and Merkle Search Trees, OAuth with DPoP, XRPC services, and event streaming.
DASL — Data-Addressed Structures & Links
A small set of simple, standard primitives to work on content-addressed data.

DASL — Data-Addressed Structures & Links
A small set of simple, standard primitives to work on content-addressed data.

WDC - RDFa, Microdata, and Microformat Data Sets
More and more websites have started to embed structured data describing products, people, organizations, places, and events into their HTML pages using markup standards such as Microdata, JSON-LD, RDFa, and Microformats. The Web Data Commons project extracts this data from several billion web pages. So far the project provides 12 different data set releases extracted from the Common Crawls 2010 to 2023. The project provides the extracted data for download and publishes statistics about the deployment of the different formats.
Crossref: The sustainable source of community-owned scholarly metadata
This paper describes the scholarly metadata collected and made available by Crossref, as well as its importance in the scholarly research ecosystem. Containing over 106 million records and expanding at an average rate of 11% a year, Crossref’s metadata has become one of the major sources of scholarly data for publishers, authors, librarians, funders, and researchers. The metadata set consists of 13 content types, including not only traditional types, such as journals and conference papers, but also data sets, reports, preprints, peer reviews, and grants. The metadata is not limited to basic publication metadata, but can also include abstracts and links to full text, funding and license information, citation links, and the information about corrections, updates, retractions, etc. This scale and breadth make Crossref a valuable source for research in scientometrics, including measuring the growth and impact of science and understanding new trends in scholarly communications. The metadata is available through a number of APIs, including REST API and OAI-PMH. In this paper, we describe the kind of metadata that Crossref provides and how it is collected and curated. We also look at Crossref’s role in the research ecosystem and trends in metadata curation over the years, including the evolution of its citation data provision. We summarize the research used in Crossref’s metadata and describe plans that will improve metadata quality and retrieval in the future.

You are Crossref - Crossref
Crossref runs open infrastructure to link research objects, entities, and actions—creating a lasting and reusable scholarly record that underpins open science. Together with our >25,000 members in 167 countries, we drive metadata exchange and support 2.1 billion monthly API queries, facilitating global research communication, for the benefit of society.

Manuscript submission systems and metadata completeness in Crossref: Patterns and associations
The importance of open research information, particularly publication metadata, is widely recognised. Crossref is one of the most important infrastructures for registering open metadata as part of DOI record registration. It is widely known, however, that the metadata of many publications is far from complete, with many publishers making certain metadata openly available, but failing to do so for other metadata elements. Publishers’ ability to register this metadata with Crossref depends on their capacity to capture and retain this data in their production workflows. Manuscript submission systems are an important, yet largely overlooked, factor in the extent to which publishers make metadata available through Crossref. In this paper, we present the results of an analysis investigating the relation between the level of metadata that publishers deposit with Crossref and the submission systems that they deploy for their journals. We have looked at the 153 publishers with the largest amounts of publications in Crossref and concentrate on the four most commonly used systems: Editorial Manager, ScholarOne, Open Journal Systems (OJS) and eJournalPress. We show that some submission systems appear better suited to capturing certain metadata elements. However, there are always cases where publishers using the same system differ widely in the level of metadata they register, suggesting that technology is not the only prohibiting factor and other considerations are at play.
Introducing the Open Exchange Architecture
Scientific communication still privileges narrative over utility, while the data, code, protocols, and computation that underpin modern research are pushed into hard-to-access supplements. Rowan Cockett introduces the Open Exchange Architecture (OXA), an emerging, community-driven standard designed to unbundle research from its paper-shaped container and expose the rich evidence underneath it.
Linked Open Vocabularies (LOV)
Your entry point to high quality and reusable Vocabularies to describe Linked Data.
