







Data Mesh Architecture: Interoperability, Co-Operation, and Co-Regulation
This paper examines the emergence of data mesh architectures and data intermediaries as key components in emerging data economies. It explores how data mesh arc
CopyFair License - P2P Foundation
= a type of licensing or agreement that aims to re-introduce the principle and practice of reciprocity in markets that use mutualized knowledge (commons), by regulating contributions to these commons for those that commercialize it
What-if: crates.io paid subscription for commercial use
In @nikomatsakis' 2020 blog post, he calls for a shift in focus "from adoption to investment. (...) for Rust to really thrive, we need to see more people paid for their work on Rust teams. I wanna accept that call-to-action and propose one such opportunity for investment. tl;dr: The online service Crates.io could start charging commercial users a nominal fee, collected purely through opt-in payment; no paywalls or licensing changes. How crates.io could turn a profit without hurtin' anybody ...
Parting Clouds: Creating a Competitive Marketplace for Compute | Canadian Anti-Monopoly Project
The Canadian market for cloud computing is monopolized. In a new report, CAMP argues that the most effective response is to unleash competition through procurement, regulation, and competition law enforcement.

Simple Pricing | Machine Learning Infrastructure | Deep Infra
We provide only pay-what-you-use pricing with no long-term contracts or upfront costs for our machine learning models and infrastructure. Learn more!


Pricing
See exactly what you pay. No hidden fees, no traffic limits on fair usage. Flexible plans that scale with your business. Check rates »

Licensing terms for data in the Atmosphere
I’d like to raise a fun and interesting topic, and that’s intellectual property and content licensing. The data that flies around the Atmosphere is all visible and public. We’re at a stage within the Atmosphere where services are beginning to grow, and we’re beginning to see lexicons and data that express creative works that extend in length beyond microblogging. I think that the Standard.site lexicons are a great example. But these creations lead to further questions. Who has permission to...

Designing a multi-sided data platform: findings from the International Data Spaces case
The paper presents the findings from a 3-year single-case study conducted in connection with the International Data Spaces (IDS) initiative. The IDS represents a multi-sided platform (MSP) for secure and trusted data exchange, which is governed by an institutionalized alliance of different stakeholder organizations. The paper delivers insights gained during the early stages of the platform’s lifecycle (i.e. the platform design process). More specifically, it provides answers to three research questions, namely how alliance-driven MSPs come into existence and evolve, how different stakeholder groups use certain governance mechanisms during the platform design process, and how this process is influenced by regulatory instruments. By contrasting the case of an alliance-driven MSP with the more common approach of the keystone-driven MSP, the results of the case study suggest that different evolutionary paths can be pursued during the early stages of an MSP’s lifecycle. Furthermore, the IDS initiative considers trust and data sovereignty more relevant regulatory instruments compared to pricing, for example. Finally, the study advances the body of scientific knowledge with regard to data being a boundary resource on MSPs.

On the Impact of the Utility in Semivalue-based Data Valuation
Semivalue–based data valuation uses cooperative‐game theory intuitions to assign each data point a value reflecting its contribution to a downstream task. Still, those values depend on the...
MotherDuck: Pricing - Get Started
Blazing fast analytics, ducking simple pricing. Plans that fit customer-facing analytics and internal data warehousing use cases.

Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
Permissive licenses like MIT, Apache-2.0, and BSD-3-Clause dominate open-source AI, signaling that artifacts like models, datasets, and code can be freely used, modified, and redistributed. However, these licenses carry mandatory requirements: include the full license text, provide a copyright notice, and preserve upstream attribution, that remain unverified at scale. Failure to meet these conditions can place reuse outside the scope of the license, effectively leaving AI artifacts under default copyright for those uses and exposing downstream users to litigation. We call this phenomenon ``permissive washing'': labeling AI artifacts as free to use, while omitting the legal documentation required to make that label actionable. To assess how widespread permissive washing is in the AI supply chain, we empirically audit 124,278 dataset $\rightarrow$ model $\rightarrow$ application supply chains, spanning 3,338 datasets, 6,664 models, and 28,516 applications across Hugging Face and GitHub. We find that an astonishing 96.5\% of datasets and 95.8\% of models lack the required license text, only 2.3\% of datasets and 3.2\% of models satisfy both license text and copyright requirements, and even when upstream artifacts provide complete licensing evidence, attribution rarely propagates downstream: only 27.59\% of models preserve compliant dataset notices and only 5.75\% of applications preserve compliant model notices (with just 6.38\% preserving any linked upstream notice). Practitioners cannot assume permissive labels confer the rights they claim: license files and notices, not metadata, are the source of legal truth. To support future research, we release our full audit dataset and reproducible pipeline.

"The end result of all of this is that we grew from 2k users to almost 150k, added a ton of heavy new functionality, and still managed to optimize and cut down costs from $.15 per active user per month to just $.03 or so." - @snarfed.org breaks down all the work he's done to optimize Bridgy Fed 1/
Bridging on a budget
blog.anew.socialThis approach to a data commons could extend to so many typically shared things. For businesses to easily and authoritatively publish things like open hours or product catalogs, discographies, every kind of -ography... creative appmakers could do so much with all that
JB Hutch
Testing an idea that I've had for a while, but I think is even better suited for atproto Internet, meet plts.at
Part of the datacounterfactuals.org reading lists. Research on data provenance, dataset documentation, licensing and attribution audits, and technical source-attribution methods for understanding which data sources are available, permitted, or responsible for model behavior.
WASA: WAtermark-based Source Attribution for Large Language Model-Generated Data
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore

Datasheets for Datasets
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

A large-scale audit of dataset licensing and attribution in AI