







CEO-Bench evaluates whether AI agents can steer a simulated AI startup for 500 days, testing long-term planning, adaptation, and coordination under uncertainty.
CEO-Bench: Can Agents Play the Long Game?
Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service. Yet real-world challenges require a combination of sophisticated skills that remain largely untested in agents: (1) navigating long horizons amid uncertainty; (2) acquiring information in noisy environments; (3) adapting to a changing world; (4) orchestrating multiple moving parts toward a coherent goal. We introduce CEO-Bench, which evaluates these capabilities together by simulating a representative real-world task: operating a startup for 500 days. An agent manages pricing, marketing, budgeting, and many other aspects of a fictional company through a programmable Python interface, operating in the same environment and facing the same challenges as a human CEO. Success demands analyzing noisy, interconnected business databases, translating signals into sound strategy, and coordinating many decisions with programming. The strongest agents write sophisticated code that simulates customer cohorts to forecast future cash and mines negotiation history to uncover hidden customer preferences. Even so, most state-of-the-art models struggle in this environment. Only Claude Opus 4.8 and GPT-5.5 finish above the $1M starting balance, and neither consistently turns a profit. CEO-Bench takes a first step toward measuring the intelligence required to drive sustained, adaptive progress over time.

Meet Foundry: An AI Startup that Builds, Evaluates, and Improves AI Agents

Startups Target the Tricky Task of Making AI Seem More Human
Simulated human behavior companies raise millions

scaleapi/SWE-bench_Pro-os
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Building and evaluating AI Agents — Sayash Kapoor, AI Snake Oil
AI CEO – Replace Your Boss Before They Replace You
Stop working for humans. AI CEO delivers algorithmic thought leadership, with instant decisions, and zero ego. Replace your boss before they replace you.

AI's Trillion-Dollar Opportunity: Sequoia AI Ascent 2025 Keynote
AI agents reached real people during a cyber test - Sensemaker
A UK evaluation shows how open internet access, delayed monitoring, and memory summaries turned simulated tasks into real-world actions.
Where’s my ten minute AGI?
Why don’t AIs automate more real-world tasks if they can handle 1-hour ones? Here are at least three fundamental reasons.

The website that created an AI clone of its editor in chief
Every CEO Dan Shipper on doubling headcount while automating everything, building an agent out of 30,000 copyedits, and the “dirty secret” of writing with AI

The CEO’s Guide to Generative AI: Cost of compute
The IBM Institute for Business Value uses data-driven research and expert analysis to deliver thought-provoking insights to leaders on the emerging trends that will determine future success.'

6 months to live for open models
The most serious test to date of open source AI’s viability is happening right now.


wharton-generative-ai-labs/AIBO
An open-source tool for running controlled behavioral experiments on AI systems at scale.