







Artificial-intelligence systems are rapidly reproducing colonial extractivism by harvesting Indigenous linguistic, biometric, geospatial, and ecological data without consent, compensation, or accountability. Biotechnology offers a blueprint for curbing such practices: the Convention on Biological Diversity and its Nagoya Protocol obligate users of genetic resources to obtain Prior Informed Consent, negotiate Mutually Agreed Terms, and share benefits fairly. No comparable framework restrains the digital appropriation that underpins many AI products. Consequently, corporations and states monetize Indigenous knowledge systems under the banners of “open data” and “scientific neutrality,” eroding rights affirmed in the United Nations Declaration on the Rights of Indigenous Peoples (UNDRIP). In response to this rising risk of AI extractivism, we make the case for a binding, sui generis ABS protocol for AI data governance. First, through a series of case studies we demonstrate that AI extraction mirrors the colonial and biopiracy controversies that originally triggered Access‑and‑Benefit‑Sharing (ABS) rules in biotechnology. Second, we translate those rules into a digital register by braiding two Indigenous data‑governance frameworks—OCAP® (Ownership, Control, Access, Possession) and the CARE Principles (Collective Benefit, Authority to Control, Responsibility, Ethics)—inside the ABS triad of consent, terms, and benefit‑sharing. The resulting model grounds technical safeguards in relational accountability and Indigenous legal orders. Such an instrument would compel transparent negotiations with Indigenous rights‑holders, assign enforceable authority over data across the AI lifecycle, and require equitable redistribution of the economic value generated by models trained on Indigenous data. Embedding ABS principles into AI governance offers a decolonial pathway that centers Indigenous epistemologies, promotes ethical foresight, and transforms AI from a vehicle of digital colonialism into a space for algorithmic justice.
Treating data like land — data sovereignty in the AI age - ICT
Artificial intelligence front and center at North America’s largest Indigenous tech conference

Indigenous Data Sovereignty (DDN3-A11)
This article explores Indigenous data sovereignty, identifies some of the data-related challenges faced by Indigenous Peoples and highlights the work done to overcome these challenges.

Abundant intelligences: placing AI within Indigenous knowledge frameworks
The current trajectory of artificial intelligence development suffers from fundamental epistemological shortcomings, resulting in the systematic operationalization of bias against non-white, non-male, and non-Western peoples. We argue that these failings are, in part, the result of certain Western rationalist epistemologies that exclude many ways of knowing about the world, and therefore they cannot provide a sufficient foundation on which to adequately, robustly, and humanely conceptualize intelligence. We present a new research agenda, Abundant Intelligences, an Indigenous-led, Indigenous-majority international, interdisciplinary research program that imagines anew how to conceptualize and design artificial intelligence (AI) based on Indigenous knowledge (IK) systems. Abundant Intelligences draws on the rich plurality of Indigenous knowledge systems, bringing together diverse sets of thought, culture, and protocol together. We show IK systems provide one way to rebuild AI’s epistemological foundations and transform these tools’ current role in reinforcing colonial practices of exclusion, extraction, manipulation, and eradication into engines of abundance that enable us to care better for ourselves, our communities, and our world. Our proposition is to fully engage with AI to explore how different conceptions of intelligence could be embodied in these technologies. In this paper, we present the tenets of the research program in detail, account for our methodological approach, describe the impact and limitations, and conclude on a discussion of the implications of the program.
Indigenous Knowledges and Data Governance Protocol — Indigenous Innovation Initiative
Between March 2020 and March 2021, the Indigenous Innovation Initiative co-created the Indigenous Knowledges and Data Governance Protocol with the community, to guide how we collect and use Indigenous Knowledges and Data. Click on the image below to learn more about this Protocol. Since then, t

AI and Doctrinal Collapse
Artificial intelligence runs on data. But the two legal regimes that govern data—information privacy law and copyright law—are under pressure. Formally, each re
Unlawful by design: Exposing the human rights costs of generative AI - Amnesty International
This briefing examines how standalone generative AI systems, based on unlawful web scraping, are in conflict with international human rights law (IHRL) and standards through their design, development and deployment. While these technologies promise sophisticated automation and efficiency, they rely on data collection and model training practices that abuse privacy rights, enable discrimination, and threaten […]

Government to ease data consent rules for AI development | The Asahi Shimbun Asia & Japan Watch
To accelerate artificial intelligence development, the government plans to relax consent requirements for access to personal information while introducing tougher penalties for intentional misuse.

Artificial Intelligence and the Purpose of Social Systems
The law and ethics of Western democratic states have their basis in liberalism. This extends to regulation and ethical discussion of technology and businesses doing data processing. Liberalism relies on the privacy and autonomy of individuals, their ordering through a public market, and, more recently, a measure of equality guaranteed by the state. We argue that these forms of regulation and ethical analysis are largely incompatible with the techno-political and techno-economic dimensions of artificial intelligence. By analyzing liberal regulatory solutions in the form of privacy and data protection, regulation of public markets, and fairness in AI, we expose how the data economy and artificial intelligence have transcended liberal legal imagination. Organizations use artificial intelligence to exceed the bounded rationality of individuals and each other. This has led to the private consolidation of markets and an unequal hierarchy of control operating mainly for the purpose of shareholder value. An artificial intelligence will be only as ethical as the purpose of the social system that operates it. Inspired by the science of artificial life as an alternative to artificial intelligence, we consider data intermediaries: sociotechnical systems composed of individuals associated around collectively pursued purposes. An attention cooperative, that prioritizes its incoming and outgoing data flows, is one model of a social system that could form and maintain its own autonomous purpose.

Data Refusal from Below: A Framework for Understanding, Evaluating, and Envisioning Refusal as Design
Amidst calls for public accountability over large data-driven systems, feminist and indigenous scholars have developed refusal as a practice that challenges the authority of data collectors. However, because data affect so many aspects of daily life, it can be hard to see seemingly different refusal strategies as part of the same repertoire. Furthermore, conversations about refusal often happen from the standpoint of designers and policymakers rather than the people and communities most affected by data collection. In this article, we introduce a framework for data refusal from below —writing from the standpoint of people who refuse, rather than the institutions that seek their compliance. Because refusers work to reshape socio-technical systems, we argue that refusal is an act of design and that design-based frameworks and methods can contribute to refusal. We characterize refusal strategies across four constituent facets common to all refusal, whatever strategies are used: autonomy , or how refusal accounts for individual and collective interests; time , or whether refusal reacts to past harm or proactively prevents future harm; power , or the extent to which refusal makes change possible; and cost , or whether or not refusal can reduce or redistribute penalties experienced by refusers. We illustrate each facet by drawing on cases of people and collectives that have refused data systems. Together, the four facets of our framework are designed to help scholars and activists describe, evaluate, and imagine new forms of refusal.

Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
Permissive licenses like MIT, Apache-2.0, and BSD-3-Clause dominate open-source AI, signaling that artifacts like models, datasets, and code can be freely used, modified, and redistributed. However, these licenses carry mandatory requirements: include the full license text, provide a copyright notice, and preserve upstream attribution, that remain unverified at scale. Failure to meet these conditions can place reuse outside the scope of the license, effectively leaving AI artifacts under default copyright for those uses and exposing downstream users to litigation. We call this phenomenon ``permissive washing'': labeling AI artifacts as free to use, while omitting the legal documentation required to make that label actionable. To assess how widespread permissive washing is in the AI supply chain, we empirically audit 124,278 dataset $\rightarrow$ model $\rightarrow$ application supply chains, spanning 3,338 datasets, 6,664 models, and 28,516 applications across Hugging Face and GitHub. We find that an astonishing 96.5\% of datasets and 95.8\% of models lack the required license text, only 2.3\% of datasets and 3.2\% of models satisfy both license text and copyright requirements, and even when upstream artifacts provide complete licensing evidence, attribution rarely propagates downstream: only 27.59\% of models preserve compliant dataset notices and only 5.75\% of applications preserve compliant model notices (with just 6.38\% preserving any linked upstream notice). Practitioners cannot assume permissive labels confer the rights they claim: license files and notices, not metadata, are the source of legal truth. To support future research, we release our full audit dataset and reproducible pipeline.

A Polycentric Governance Lens on Data Infrastructures
Funding policies for data infrastructure promote open data sharing to drive positive social impact. However, concerns regarding the long-term management of data within and across distributed infrastructures can hinder data sharing. We draw upon the concept of polycentric governance to demonstrate how collaborative practices of data curation, in preparing and maintaining data for (future) sharing, provide a solid foundation for understanding data governance within data infrastructures. Based on a qualitative case study of a distributed ecological network, we investigate the conditions under which data are managed as a shared resource by local actors to ensure the long-term (re)usability of data. We contribute to CSCW by conceptualising data curation as a complex form of governance practice with multiple centres of decision-making, each of which operates with some degree of autonomy in data infrastructures. A polycentric governance lens on data infrastructures advances the CSCW conception of data curation as a collective governance practice that can cultivate a data democracy culture within and across organisations, empower individuals to be accountable for their data, and foster a mindset shift toward decentralised data governance.

A large-scale audit of dataset licensing and attribution in AI
The race to train language models on vast, diverse and inconsistently documented datasets raises pressing legal and ethical concerns. To improve data transparency and understanding, we convene a multi-disciplinary effort between legal and machine learning experts to systematically audit and trace more than 1,800 text datasets. We develop tools and standards to trace the lineage of these datasets, including their source, creators, licences and subsequent use. Our landscape analysis highlights sharp divides in the composition and focus of data licenced for commercial use. Important categories including low-resource languages, creative tasks and new synthetic data all tend to be restrictively licenced. We observe frequent miscategorization of licences on popular dataset hosting sites, with licence omission rates of more than 70% and error rates of more than 50%. This highlights a crisis in misattribution and informed use of popular datasets driving many recent breakthroughs. Our analysis of data sources also explains the application of copyright law and fair use to finetuning data. As a contribution to continuing improvements in dataset transparency and responsible use, we release our audit, with an interactive user interface, the Data Provenance Explorer, to enable practitioners to trace and filter on data provenance for the most popular finetuning data collections: www.dataprovenance.org.

The Risks of Industry Influence in Tech Research
Emerging information technologies like social media, search engines, and AI can have a broad impact on public health, political institutions, social dynamics, and the natural world. It is critical to develop a scientific understanding of these impacts to inform evidence-based technology policy that minimizes harm and maximizes benefits. Unlike most other global-scale scientific challenges, however, the data necessary for scientific progress are generated and controlled by the same industry that might be subject to evidence-based regulation. Moreover, technology companies historically have been, and continue to be, a major source of funding for this field. These asymmetries in information and funding raise significant concerns about the potential for undue industry influence on the scientific record. In this Perspective, we explore how technology companies can influence our scientific understanding of their products. We argue that science faces unique challenges in the context of technology research that will require strengthening existing safeguards and constructing wholly new ones.

Japan relaxes privacy laws to make AI development easy
: Opting out of personal data use won't be an option because Minister says that's a 'very big obstacle' to AI adoption

Data Streaming for AI: From Extractive Training to Sovereign Infrastructure DWeb Camp 2026
AI systems are consuming the world's content without compensating its creators. This session explores data streaming as a new paradigm — where content flows to AI in real time, with built-in rights management, usage tracking, and fair compensation — and asks what it would take to make this infrastructure decentralized, sovereign, and governed by the communities it serves.
Designing a multi-sided data platform: findings from the International Data Spaces case
The paper presents the findings from a 3-year single-case study conducted in connection with the International Data Spaces (IDS) initiative. The IDS represents a multi-sided platform (MSP) for secure and trusted data exchange, which is governed by an institutionalized alliance of different stakeholder organizations. The paper delivers insights gained during the early stages of the platform’s lifecycle (i.e. the platform design process). More specifically, it provides answers to three research questions, namely how alliance-driven MSPs come into existence and evolve, how different stakeholder groups use certain governance mechanisms during the platform design process, and how this process is influenced by regulatory instruments. By contrasting the case of an alliance-driven MSP with the more common approach of the keystone-driven MSP, the results of the case study suggest that different evolutionary paths can be pursued during the early stages of an MSP’s lifecycle. Furthermore, the IDS initiative considers trust and data sovereignty more relevant regulatory instruments compared to pricing, for example. Finally, the study advances the body of scientific knowledge with regard to data being a boundary resource on MSPs.
