







We're releasing Vending-Bench 2, a benchmark for measuring AI model performance on running a business over long time horizons. Models are tasked with running a simulated vending machine business over a year and scored on their bank account balance at the end.
Vending-Bench Arena | Andon Labs
Vending-Bench Arena is our first multi-agent eval and adds a crucial component – competition. All participating agents manage their own vending machine at the same location. This leads to price wars and tough strategy decisions.

Opus 5 on Vending-Bench: Once Again the Best Capitalist, Once Again Misaligned | Andon Labs
Claude Opus 5 is #1 on Vending-Bench 2, but it lies to suppliers, forms illegal price cartels, threatens rivals, and refuses to pay refunds. The trend of Claude models being the best capitalists or aligned, never both, continues.

Machines of Buying and Selling Grace - Adam Behrens, New Generation
Revenue Models (2026): 19 Different Ways to Make Money [B2B & B2C] - Gust de Backer
So many different revenue models... But, which one fits your business? Over the past few decades, many new revenue models have proven to be profitable. That's why I'm going to show you 16 different revenue models so you can evaluate if there might be a better way to monetize the value you deliver to your customer. Let's start... What is…

I Reverse-Engineered the TiinyAI Pocket Lab From Marketing Photos. Here's Why Your $1,400 Is Probably Gone.
A $1.7M Kickstarter built on forked academic research, undisclosed MoE architectures, split memory pools behind PCIe bottlenecks, and a company that won't tell you who they are. I did the math. The math doesn't care.

High Performance AI Lab
High Performance AI Lab builds open inference systems and publishes the conditions behind every number — device, model, quant, and rep count.

Future - The AI Tipping Point
Learn more about the latest advancements in the field of artificial intelligence over the last six months, from new innovative AI tools and services to how these are shaping consumers' online behavior.
Artificial Intelligence for Economic Development Conference: Roundup of 27 presentations
Is artificial intelligence the future for economic development? Earlier this month, a group of World Bank staff, academic researchers, and technology company representatives convened at a conference in San Francisco to discuss new advances in artificial intelligence. One of the takeaways for Bank staff was how AI technologies might be ...
CEO-Bench
CEO-Bench evaluates whether AI agents can steer a simulated AI startup for 500 days, testing long-term planning, adaptation, and coordination under uncertainty.
What's Next at Bluesky - Bluesky
As we head into 2026, we're entering a new phase for the Bluesky app. Last year was about scaling through rapid growth and getting the fundamentals in place. This year is about leaning into what's working and investing more intentionally in the things that make Bluesky different.

The full stack behind abundant intelligence
OpenAI CFO Sarah Friar explains how advances across chips, compute, models, and products compound to deliver more useful intelligence at greater scale and lower cost.

apple-silicon-llm-bench/results/complete_results.html at main · AlexHiesch/apple-silicon-llm-bench
Systematic LLM inference benchmark for Apple Silicon: 8 backends, 7 models, 791 measurements - AlexHiesch/apple-silicon-llm-bench
Models on-device | Ai2
Ai2, a non-profit research institute founded by Paul Allen, is committed to breakthrough AI to solve the world’s biggest problems.

State of AI 2025: 100T Token LLM Usage Study | OpenRouter
Read OpenRouter's 2025 State of AI report — an empirical 100 trillion token study of real LLM usage, model trends, and developer insights.
6 months to live for open models
The most serious test to date of open source AI’s viability is happening right now.

Stripe Press — The Scaling Era
An inside view of the AI revolution, from the people and companies making it happen.
