Meta fed its AI on almost everything you’ve posted publicly since 2007
Making Facebook and Instagram private won’t delete that data.

MetaやOpenAIがAIモデル開発に使っていた世界最大級のオンライン海賊版ライブラリ「LibGen」とは?
高性能なAIモデルを開発するには、膨大な量の高品質なデータを用いてトレーニングする必要があります。MetaやOpenAIがAIモデルのトレーニングに使ったとされるオンライン海賊版ライブラリ「Library Genesis(LibGen)」やその倫理的問題について、海外メディアのThe Atlanticが報じました。

The Unbelievable Scale of AI’s Pirated-Books Problem
Meta pirated millions of books to train its AI. Search through them here.
AIトレーニングについてコンテンツ作者に使用許可を求めるなら「国のAI産業が一夜で消滅してしまう」と元Meta幹部のニック・クレッグが語る
イギリスの副首相経験者で、元Metaの国際問題担当社長でもあるニック・クレッグ氏が、AIのトレーニングに使用するコンテンツについて、作成者に使用許可を求めるようになれば「この国のAI産業は一夜で消滅してしまう」と語りました。

It’s remarkably easy to inject new medical misinformation into LLMs
Changing just 0.001% of inputs to misinformation makes the AI less accurate.

Elon Musk’s xAI used child porn to train Grok models, lawsuit says
xAI accused of training Grok on real and AI-generated child pornography.

Fake US thinktank set up and funded by Israel sought to game AI for propaganda
In effort to prime chatbots to make pro-Israel arguments the site published 124 reports, over 560,000 words in nine days, Guardian analysis shows

An AI job boom? Here’s what the tedious, temporary work in data labelling is actually like
Interviews with data workers in China and Australia reveal the precarious, exploitative conditions of this new branch of the gig economy.

We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility
We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

Time Magazine has a separate version of its website with ads only AI can see
Brands are already paying to influence what chatbots say about them

TIME Is Serving AI Bots a Different Website, With Ads Built In
TIME is now serving two different versions of its website. Humans get the magazine. AI crawlers get a stripped down markdown copy with ads baked in that no person will ever see. I fetched one ordinary…

Time has started serving ads to AI agents
Time is selling ads targeting AI agents, betting that markdown pages will make its content (and advertisers) more visible in AI search.

MediaDailyNews: Meta Forges AI Content Deal With Newsmax
Meta will gain access to Newsmax's archived and current library of digital news content, with the ability to train its AI-powered search and discovery tools while supporting user searches on Facebook, Instagram, WhatsApp, and Meta AI.

Book publishers sue Google for copyright infringement over Gemini AI training
Group of major publishers accuses the tech giant of ‘one of the most prolific infringements of copyrighted materials in history’

Anthropic destroyed millions of print books to build its AI models
Company hired Google's book-scanning chief to cut up and digitize "all the books in the world."

AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop
ISBNdb, a company that sources printed books for AI companies to turn into training data, tells clients “the optics problem is real.”
OpenAI may have made a fatal misstep in copyright fight with news orgs
OpenAI may be sanctioned for hiding, deleting ChatGPT logs in NYT copyright fight.

Nvidia can't shake authors' claims it trained AI on pirated books
The case could reshape how artificial intelligence companies are allowed to acquire the massive datasets they need to build their systems.

* I’m neither “pro-AI” nor “anti-AI.” I’ve been blocked for being perceived as both. —Actually, I’m honestly more anti-AI than pro-AI thus far, aside from specialized models and specific use cases, but I’m willing to consider information that’s new to me
AYFKM “Meta will be able to draw upon Newsmax's significant current reporting as well as archived content to support AI queries across Meta's apps and devices.” ir.newsmax.com/news/news-details/2026/Newsma…