







DPC RAM Benchmarking Project CARL’s Digital Preservation Working Group (DPWG) is facilitating a national benchmarking exercise using the Digital Preservation Coalition’s Rapid Assessment Model (DPC RAM). This project fulfills one of the main recommendations from […]
IIPC WAC 2024 Presentation: A Conceptual Model of Decentralized Storage for Community-Based Archives
Rethinking Digital Storage for Community-Based Archives
Over the last few decades, the standard procedure for storing and backing up data has been ‘put it in the cloud.’ Many of us use cloud…

ArchiveBox - Open-source self-hosted web archiving
Preserve websites, media, bookmarks, feeds, source code, evidence, and research material in durable files you control.
Web Archiving: Playback Tools - Thomas Preece
Web Archiving: Playback Tools - Below I've listed some of the tools I found to playback web archives. The two most popular tools that I found were OpenWayback and PyWb. Of the tests I

Archivist - Storage That Can't Be Stopped
Archivist is a decentralized durable storage network. Data is encrypted, erasure-coded, and dispersed across independent operators. Verified by zero-knowledge proofs. Sovereign storage that survives censorship, provider failure, and time.

DigitaltMuseum
DigitaltMuseum is a common database for Norwegian and Swedish museums and collections. It provides access to more than four million photographs, objects, works of art and buildings.
Universal Version Control
Building tools to help people explore alternatives, keep track of history, and collaborate better, across all kinds of media.

Home Page - Software Heritage
preserves software source code for present and future generations
Data‐ and code‐archiving in the British Ecological Society journals: Present status and recommendations for future improvements
Abstract Data‐ and code‐archiving are important components of open science, as both make research more transparent, reproducible, accountable and credible, allowing future researchers to build on previous work. Despite progress in implementing data‐ and code‐archiving policies in journals publishing ecology and evolution research, issues remain. To be more useful to future researchers, archived data and code must not only be archived but also meet good practice standards. We collected data from 1861 papers published between 2017 and 2024 in the seven British Ecological Society (BES) journals, during a hackathon event. We systematically checked associated data and/or code, metadata, help files and annotations to assess archiving practices. We determined if and where data and code files were archived, whether they could be located, downloaded and opened, and whether they had associated READMEs, digital object identifiers (DOI) and licences. We also recorded the file extensions used to save data/code files, and which programming languages code was written in. 93% of the 1861 papers we examined used data and ~90% used code. While 97% of the 1735 papers that used data also archived it, only 35% of the 1670 papers that used code also archived code. Over 85% of archived data and code could be located, downloaded and opened. Reusability, however, was more limited; around a third of papers did not have a README or similar to explain their data/code files, and the quality of READMEs varied substantially. We recommend that researchers archive their code and that archived code be explicitly mentioned in the Data (or Code) Availability statement. We also encourage researchers to provide more accessible and informative READMEs for data and code. To help achieve these recommendations, we advocate that journals employ Data/Code editors to review data and code quality, research institutions deliver more training in open science practices, and funding bodies set clear expectations on open data and code practices.

The Sounds Resource
Archiving and preserving video game media since 2003!
File Over App: A Philosophy for Digital Longevity
My take on the 'file over app' philosophy and why it’s essential for keeping my data resilient and built to last.

CaviraOSS/OpenMemory
Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.
Solutions for Link Rot on the Modern Web
Web links are fundamental to the web, enabling navigation between pages and citations in research articles. However, web links suffer from "link rot", a phenomenon in which links are likely to become inaccessible over time. This can occur if a link’s site disappears, making it impossible to resolve the hostname or establish a server connection, or the linked page has been deleted, resulting in an HTTP 404 "Not Found" error. Today, a common solution to tackle link rot is to rely on web archives, which capture snapshots of web pages for future reference. However, the modern web has evolved significantly since web archives were first introduced, leading to several limitations in their effectiveness. First, the scale of the web has grown tremendously, making it infeasible to crawl every page whenever it changes. As a result, many broken links either have no archived copies, or the archived copies have stale content. Second, modern web pages rely heavily on increasingly complex and diverse JavaScript. This shift has made it more challenging for preserving fidelity in archived copies, significantly increasing both the computational cost of operating browser-based crawlers and engineering effort required to maintain accurate replay systems. This thesis presents a set of solutions to cope with link rot on the modern web. My work aims to mitigate the various limitations of web archives. First, for broken links without any archived copy, or if the archived copy includes stale content or unavailable functionalities, I built Fable. Fable revives the dead link with the new URL to the same page whenever available. Compared to prior approaches, Fable revives 4.6K broken links—a 50% increase—with much higher accuracy. Second, for pages that require dynamic crawling, I show how web archives can achieve a better tradeoff between efficiency and fidelity. By carefully choosing 8.9% of pages to crawl dynamically and strategically reusing resources from those crawls, an archive can serve 99% of the remaining statically crawled pages without any fidelity loss. Third, for fidelity violations in archived copies that are caused by incorrect edits to crawled scripts, I built FidEx. FidEx reliably detects when an archived page differs from its original version and pinpoints the root cause. After fixing the most common errors pinpointed by FidEx, I reduced the fraction of pages for which FidEx reports a violation of fidelity from 15% to 9%.
Back into blogging and just published something I've been thinking about for a while now – making archival content more resilient and discoverable on atproto. Oral history, interactive transcripts, content addressing, and keeping important stories from being quietly erased. maboa.it/resilient-archives-on-the-at-…
Keeping Archives Alive: Resilience and Discovery on ATProto
maboa.itFor me personally there is a journey here - I was a firm believer in archive everything. This peaked in a period where I was very interested in IPFS, CIDs, CAS, etc. Now, the next step. Post-Archive(-Everything). Through a series of conversations I came to the conclusion that forgetting is probably more powerful than remembering. The links in this collections play a similar note, or are related to my understanding of this.

FOSDEM 2026 - Willow - Protocols for an uncertain future
Willow - Home

I Deleted My Second Brain