







Yesterday a blog post by Cyrille Rossant entitled "Moving away from HDF5" caught my eye. My own tendency at the moment is to use HDF5 more and more, so I was interested in why someone else would want to do the opposite. Here is my conclusion after reading his post, plus some ideas about where scientific data management is or should be heading in my opinion.
The HDF5® Library & File Format - The HDF Group - ensuring long-term access and usability of HDF data and supporting users of HDF technologies
HDF5® Technical Details License: BSD-style Official Media Type: application/vnd.hdfgroup.hdf5 Standard File Extension: .h5, .hdf5 High-performance data management and storage suite Utilize the HDF5 high performance data software library and file format to manage, process, and store your heterogeneous data. HDF5 is built for fast I/O processing and storage. Download HDF5 Documentation What is HDF5®? HETEROGENEOUS DATA HDF® sup ports […]

FAIR Principles - GO FAIR
In 2016, the ‘FAIR Guiding Principles for scientific data management and stewardship’ were published in Scientific Data. The authors intended to provide guidelines to improve the Findability, Accessibility, Interoperability, and Reuse of digital assets. The principles emphasise machine-actionability (i.e., the capacity of… Continue reading →

Data Feminism
Today, data science is a form of power. It has been used to expose injustice, improve health outcomes, and topple governments. But it has also been used to d...

Who Will Keep Research Data Infrastructure Open and Running?
The scientific community must consider the longevity of open research infrastructure—why it might fail and how to prevent it.

Who Will Keep Research Data Infrastructure Open and Running?
The scientific community must consider the longevity of open research infrastructure—why it might fail and how to prevent it.

Why on earth would you publish data as RDF? – Pieter Colpaert
A rebuttal to a blog post of Anne Schuth, I which I argue that RDF will earn its place in real neurosymbolic systems. Its best argument is eventual interoperability across organizations.

2025 Global Data Center Outlook
The data center sector will continue to grow at a phenomenal pace in 2025.

Linked Data is a Political Agenda
A Data Utopia for Science-of-Science
Here I want to briefly sketch out a vision for how to solve a key set of problems facing science-of-science researchers, using the relatively new idea of a ‘data trust.’ In my ideal wor…

The Bazaar of Scientific Knowledge | shishyko!
Is the current form of the scientific paper still optimal in 2025? How do we preserve, and efficiently leverage, the uncut gems of the scientific process?
Living in Data: A Citizen's Guide to a Better Information Future (Paperback)
Jer Thorp’s analysis of the word “data” in 10,325 New York Times stories written between 1984 and 2018 shows a distinct trend: among the words most closely associated with “data,” we find not only its classic companions “information” and “digital,” but also a variety of new neighbors—from “scandal” and “misinformation” to “ethics,” “friends,” and “play.”To live in data in the twenty-first century

How and When to Involve Crowds in Scientific Research | James Evans
The lone academic in a basement lab is an endangered species. Years ago I asked Paul Ginsparg, who founded arXiv, what features his machine-learning filter used to flag speculative submissions. The top three: single-authored, submitted on a weekend, and heavy citation of Newton and Einstein. Science is a contact sport now. Marion Poetz and Henry Sauermann's How and When to Involve Crowds in Scientific Research (https://lnkd.in/g5pwVbfD) is the field manual. What crowds can do: Volume. Zooniverse mobilizes 2.7 million people to classify galaxies and court records. Galaxy Zoo's co-founder hand-classified 50,000 galaxies in one week before concluding that isolation would break him. Reach. NASA spent years failing to predict solar flares, then broadcast the problem. The winner was a semi-retired radio engineer in rural New Hampshire who swapped satellite data for radio data. Experience. Patients and caregivers asked to generate research questions produced 826 that scientists had missed, including the link between aging and wound healing. Bench-to-bedside runs backward. One caution the book underplays: an open call is not an inclusive one. Participation costs in time, money, and access screen people out. Self-selection can narrow the very diversity that makes crowds thrive. As AI floods the labor supply of routine, homogeneous cognition, human crowds become more valuable, not less. Machines process. People notice what nobody thought to ask. Check out my review @ https://lnkd.in/gkzvWrRC and the book @ https://lnkd.in/g5pwVbfD!
Databases in 2025: A Year in Review
The world tried to kill Andy off but he had to stay alive to to talk about what happened with databases in 2025.
Fed up with Big Tech, communities turn to data collectives for control
Data collectives and cooperatives, which let creators control the collection and distribution of their data, are emerging as preferred alternatives to big tech companies.

Column Storage for the AI Era
In the past few years, we’ve seen a cambrian explosion of new columnar formats, challenging the hegemony of Parquet: Lance, Fastlanes, Nimble, Vortex, AnyBlox, F3 (File Format for the Future). The thinking is that the context has changed so much that the design of yore (the previous decade) is not going to cut it moving forward. This seemed a bit intriguing to me, especially since the main contribution of Parquet has been to provide a standard for columnar storage. Parquet is not simply a file format. As an open source project hosted by the ASF, it acts as a consensus building machine for the industry. Creating six new formats is not going to help with interroperability. I spent some time to understand a bit better how things actually changed and how Parquet needs to adapt to meet the demands of this new era. In this post I’ll discuss my findings.
The Limits of Data
Policymakers want to make decisions based on clear data, but important factors are lost when we rely solely on data. A philosopher writes:
