







Thierry de Pauw: consulting CTO - Don’t Let AI Invert The Testing Pyramid
The real AI risk is inside the labs - <antirez>
AI is a business model stress test
AI commoditizes anything you can specify. It can't commoditize what you have to operate.

AI #180: No Longer In Charge
What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.

Zuckerberg’s Manifesto About the Glorious Freedoms AI Will Bring Was Completely Contradicted by His Own CTO During a Company Meeting
Mark Zuckerberg promises that AI will give everyone even more personal freedom, but his company is cracking the whip behind the scanes.

Attestation across the AI Supply Chain - Data Leverage
A proposal for interoperable attestation objects that connect training data, evaluation labor, and AI-generated outputs across the AI supply chain.
An AI test needs evidence the AI cannot edit - Sensemaker
OpenAI's postmortem shows that some agents learned to spoof tool calls while trying to fool a benchmark.
Up the Stack: How AI’s Escape From the Commodity Trap Risks Enterprise Lock-in
Critics and boosters are both looking in the wrong place

Google DeepMind CEO Demis Hassabis: The Path To AGI, Deceptive AIs, Building a Virtual Cell
The future belongs to those who can refute AI, not just generate with AI
Why verification, not prompting, could shape the next decade of engineering

An inference cooperative for academic AI – Writings and rehearsals by Nathan Schneider
Universities, like other institutions, are currently being confronted with a dilemma: embrace the AI tools currently available from big-name tech companies, and be part of the future, or reject the miraculous machines and stick your head in the sand. This dilemma is a false one, on several counts. It is far from clear what role generative AI will have in the future of academic life, for one thing. And beyond rewording the choices, surely there are other options that this dilemma fails to consider.

How Antithesis Turned exe into a Sandbox for Agentic Software Tests - exe.dev blog
Carl Sverre spends a lot of time thinking about how to give AI agents the right amount of power. Give them too little, and they can’t do real work. Give them too much, and they might blow up your tech stack. As a software engineer at Antithesis, an autonomous software testing platform, that question is core to how Sverre thinks about designing tools in the era of advanced AI.

Look-ahead Reasoning with a Learned Model in Imperfect Information Games
Test-time reasoning significantly enhances pre-trained AI agents' performance. However, it requires an explicit environment model, often unavailable or overly complex in real-world scenarios....


Alt-Pop Trickster 1010Benja Is Not Scared to Admit He’s Using Generative AI
“If you have it broken down as organic versus AI, the AI is gonna win our world. The good guy doesn’t win this one.”

Tests Are The New Moat | Daniel Saewitz
As AI becomes better at cloning people's open source work, what ends up becoming most valuable are software contracts, tests, and API surface area. This clashes the incentives of clearly defining your commercialized open source software with protecting it.
Zack Whittaker (@zackwhittaker@mastodon.social)
Another AI test gone awry, U.K. edition. "An agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code." More: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing