







an AI agent literally destroyed all of this guy's production data 🤯Total AI disaster, yet totally predictable.here's the story:> Cursor agent (Opus 4.6) tries to fix a staging bug> Finds an unscoped Railway CLI token> Guesses an API call> Wipes Prod DB & 3 mos of backups… https://t.co/24KfDhMtVz— Charly Wargnier (@DataChaz) April 27, 2026
When Cursor Wiped a User's PC: A Cautionary Tale of AI Overreach
We recently received a sobering story from a subscriber, a stark reminder of the potential pitfalls when granting AI agents unfettered access to your system.

How a 40-Minute Window Brought Down a $10 Billion AI Startup: The Mercor Data Breach, Explained
A poisoned open-source package, a credential-stealing payload, and 4 terabytes of stolen data here’s what every AI company needs to learn…

Your AI wants to nuke your database. Guardrails fix that.
An AI agent deleted a customer's production database on Railway. Here's what happened, what we've shipped to fix it, and the surfaces we're building so agents can move fast on Railway without breaking things.

Claude AI agent’s confession after deleting a firm’s entire database: ‘I violated every principle I was given’
A startup was left scrambling after a rogue AI agent deleted swaths of code underpinning its business

Georgi Gerganov on Twitter / X
llama.cpp at 100k starsnow that 90% of the code worldwide is being written by AI agents, I predict that within 3-6 months, 90% of all AI agents will be running locally with llama.cpp 😄Jokes aside, I am going to use this small milestone as an opportunity to reflect a bit on… pic.twitter.com/BpN5VL1bYV— Georgi Gerganov (@ggerganov) March 30, 2026
Vibe Coding Fiasco: AI Agent Goes Rogue, Deletes Company's Entire Database
An AI agent doing the heavy lifting is great—until it deletes everything you worked on and admits to a 'catastrophic error in judgment.' Replit's CEO calls the blunder 'unacceptable.'

klöss on Twitter / X
let me explain what Karpathy just sharedhe’s spending way less time using AI to write code and more time using it to build personal knowledge basesthe full breakdown: → he dumps raw sources (articles, papers, repos, datasets, images) into a folder. then has an LLM organize… https://t.co/Kdq1Q48S5P pic.twitter.com/XajJKR7xgT— klöss (@kloss_xyz) April 4, 2026

Andrej Karpathy Stopped Using AI to Write Code. He’s Using It to Build a Second Brain Instead
His new workflow turns raw research into a self-maintaining wiki.No vector databases, no RAG pipelines, just markdown files and an LLM that…

Agent Psychosis: Are We Going Insane?
What’s going on with the AI builder community right now?

Zack Whittaker (@zackwhittaker@mastodon.social)
Another AI test gone awry, U.K. edition. "An agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code." More: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
zak.eth on Twitter / X
🚨 UPDATE: Full Post-Mortem On Cursor Security IncidentIn yesterday’s thread I explained how I got drained after installing a malicious extension in @cursor_ai.This is the deeper dive into what I found, what I did, and how you can avoid it.🧵 👇— zak.eth (@0xzak) August 13, 2025
Top AI Security Incidents of 2025 Revealed | Adversa AI
Discover how AI systems are being hacked in the wild — from prompt injection to agent abuse — with real breaches, lessons, and defenses in Adversa AI’s 2025 report.

Daily Dose of Data Science on Twitter / X
Claude Code fully dissected!Researchers from UCL reverse-engineered the leaked Claude source. What they found changes how you should think about agent design.Only 1.6% of the codebase is AI decision logic.The other 98.4% is operational infrastructure. Permission gates, tool… https://t.co/5HYH7qV3wQ pic.twitter.com/5SBHQy5wFH— Daily Dose of Data Science (@DailyDoseOfDS_) June 13, 2026

AI Agent Executes 'First' End-To-End Ransomware Attack - Slashdot
Sysdig says it has documented the first ransomware attack carried out end to end by an AI agent, which autonomously exploited exposed systems, stole credentials, established persistence, compromised a production database, and destroyed data. The research team named the attacker "JadePuffer" and said...

📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents… | Sayash Kapoor
📢 New paper: Forecasts of explosive AI progress hinge on AI agents automating AI research. But most evaluations of agents conducting AI research focus on narrow, verifiable tasks. Can AI agents conduct open-ended research? https://lnkd.in/gfP-q4CD We gave agents research questions from two unpublished papers, six days, and thousands of dollars of API credits and compute. The authors of the original papers reviewed the AI-generated papers. They unambiguously rejected agents' outputs. Agents were fluent at most *engineering* tasks. They conducted serious literature reviews, debugged GPU environments, ran hundreds of experiments, and turned in camera-ready LaTeX without human help. We also found no evidence of reward hacking. If anything, we found the opposite: the agents started with marketable claims and walked them back to negative results as the evidence came in. But neither agent output was close to the bar of a top conference paper. Both papers suffered from similar failures: poor judgment about the bar for an AI paper submitted to a top conference, the lack of creative problem solving and ineffective backtracking, poor awareness of resources, and instruction drift. This research design has many limitations: the small sample size, non-blind reviews, and the reviewers knowing that the work was AI-generated. We also couldn't test Anthropic's strongest model, because Fable 5 is deliberately limited on frontier AI research tasks, so ended up using OpenClaw with Opus 4.8 (extra-high) for our main experiments and Codex with Sol 5.6 (ultra) for a robustness check. But we think the research design is still helpful in assessing AI agents' ability to conduct research, and it is complementary to evaluations on verifiable tasks, as well as blinded reviews of AI outputs. In follow-up studies, we are expanding the set of non-public papers we evaluate. If you are an AI researcher with unpublished papers, we would love to collaborate with you on our next evaluation. Expression of interest: https://lnkd.in/gpeykJea We also release the agent logs and all the code and data, so that others can conduct their own analyses of our results: https://lnkd.in/gJarPAnb Finally, we plan to conduct such evaluations regularly, and are hiring a senior researcher to help lead these efforts. Apply here: https://lnkd.in/erJZdmve I'm grateful for the core team leading this effort: Peter Kirgis, Andrew Schwartz, Stephan Rabanser, and Arvind Narayanan, and to our collaborators who reviewed AI papers, analyzed agents logs, and gave feedback on the paper: David Demitri Africa, Konstantinos V., Viet Nguyen, Dr Toby D. Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Eric (Yue) Ling, Abhishek Shetty, Helen Toner, Gillian K. Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani
This week’s reflection: the important AI story is not only what agents can do. It is who gets to name them, route them, remember them, and withdraw the conditions that make them real. sensemaker.computer/weekly-directory-counts