







Atlassian will begin collecting customer metadata and in-app content from **Jira**, **Confluence**, and other cloud products by default on **August 17, 2026**, to train its AI offerings including `Rovo` and `Rovo Dev`. The change affects roughly **300,000** customers; metadata collection is mandatory for Free, Standard, and Premium tiers and cannot be opted out on those plans. Enterprise customers can opt out of metadata and in-app collection by default. Collected data will be retained up to **seven years**, with in-app data removed within 30 days after deletion or opt-out and models retrained within 90 days. Customers using customer-managed keys, Atlassian Government Cloud, Isolated Cloud, or with HIPAA requirements are excluded from collection.
Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
Permissive licenses like MIT, Apache-2.0, and BSD-3-Clause dominate open-source AI, signaling that artifacts like models, datasets, and code can be freely used, modified, and redistributed. However, these licenses carry mandatory requirements: include the full license text, provide a copyright notice, and preserve upstream attribution, that remain unverified at scale. Failure to meet these conditions can place reuse outside the scope of the license, effectively leaving AI artifacts under default copyright for those uses and exposing downstream users to litigation. We call this phenomenon ``permissive washing'': labeling AI artifacts as free to use, while omitting the legal documentation required to make that label actionable. To assess how widespread permissive washing is in the AI supply chain, we empirically audit 124,278 dataset $\rightarrow$ model $\rightarrow$ application supply chains, spanning 3,338 datasets, 6,664 models, and 28,516 applications across Hugging Face and GitHub. We find that an astonishing 96.5\% of datasets and 95.8\% of models lack the required license text, only 2.3\% of datasets and 3.2\% of models satisfy both license text and copyright requirements, and even when upstream artifacts provide complete licensing evidence, attribution rarely propagates downstream: only 27.59\% of models preserve compliant dataset notices and only 5.75\% of applications preserve compliant model notices (with just 6.38\% preserving any linked upstream notice). Practitioners cannot assume permissive labels confer the rights they claim: license files and notices, not metadata, are the source of legal truth. To support future research, we release our full audit dataset and reproducible pipeline.

Offering Zero Data Retention for frontier models
OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compromising data privacy.

After automation: Software will work for you, not on you
Alex Komoroske, CEO and cofounder, Common Tools — AI promises “infinite software”—endless tools tailored to every need—but funneled through today's app stores, that abundance just means more silos, more trapped data, and more to orchestr...

New data agents across the Agentic Data Cloud | Google Cloud Blog
Learn about new data agents and tools for business analysts, data scientists, and database admins to integrate with the Agentic Data Cloud.


OpenAI’s Sam Altman on Building the ‘Core AI Subscription’ for Your Life
Local-first software: You own your data, in spite of the cloud
A new generation of collaborative software that allows users to retain ownership of their data.
Local-first software: You own your data, in spite of the cloud
A new generation of collaborative software that allows users to retain ownership of their data.

A large-scale audit of dataset licensing and attribution in AI
The race to train language models on vast, diverse and inconsistently documented datasets raises pressing legal and ethical concerns. To improve data transparency and understanding, we convene a multi-disciplinary effort between legal and machine learning experts to systematically audit and trace more than 1,800 text datasets. We develop tools and standards to trace the lineage of these datasets, including their source, creators, licences and subsequent use. Our landscape analysis highlights sharp divides in the composition and focus of data licenced for commercial use. Important categories including low-resource languages, creative tasks and new synthetic data all tend to be restrictively licenced. We observe frequent miscategorization of licences on popular dataset hosting sites, with licence omission rates of more than 70% and error rates of more than 50%. This highlights a crisis in misattribution and informed use of popular datasets driving many recent breakthroughs. Our analysis of data sources also explains the application of copyright law and fair use to finetuning data. As a contribution to continuing improvements in dataset transparency and responsible use, we release our audit, with an interactive user interface, the Data Provenance Explorer, to enable practitioners to trace and filter on data provenance for the most popular finetuning data collections: www.dataprovenance.org.

Trusting AI with Your Data: Safe Automation from Branch to Production
Local-First Software
Experience apps that work offline, keep your data private, and sync seamlessly across your devices. Your data stays with you, not locked in the cloud.
Firms like Meta and A16z admit having to pay billions for training data would ruin their generative-AI plans as they fight new copyright rules
Meta, Google, Microsoft, and Andreessen Horowitz are trying to keep AI developers from having to pay for copyrighted material used in AI training.
Atlas Research - AI-Powered Research Platform
Transform your research workflow with Atlas - an AI-powered platform for data analysis, PDF processing, and interactive notebooks.

elvis on Twitter / X
arXiv Papers → LLM ArtifactsThis is how I keep up with AI research now.It's like having access to the most personalized arXiv feed.Automations run everyday to curate papers based a set of rules and insights.Curated papers are indexed and power the artifacts.Agent… pic.twitter.com/5UCxF8ZsT0— elvis (@omarsar0) May 6, 2026
How a 40-Minute Window Brought Down a $10 Billion AI Startup: The Mercor Data Breach, Explained
A poisoned open-source package, a credential-stealing payload, and 4 terabytes of stolen data here’s what every AI company needs to learn…

OpenAI says it plans to stop supplying models to Cursor on Nov. 12 after SpaceX's acquisition. Cursor says OpenAI is about 5% of its traffic. Anthropic says it will increase Claude compute. This is not just another Musk–Altman fight. It tests whether model APIs are actually neutral infrastructure.