







Deep Research for scenario planning. Stress-test your strategy against thousands of possible futures with a realtime AI scenario generation engine.


Request for Proposals: The Launch Sequence | IFP
Apply to our rolling effort to find, scope, and build the most important projects to prepare the world for advanced AI

Selective Optimism: a critique of AI 2040
Some context for this post: I’ve been working part-time as a consultant for the AI Futures Project over the last year.


The future belongs to those who can refute AI, not just generate with AI
Why verification, not prompting, could shape the next decade of engineering

Deep Research, information vs. insight, and the nature of science
What AI will accelerate in the scientific process, what it cannot do, and how we can prepare for new manners of scientific investigation.

What is AI Agent Security Plan 2026? Threats and Strategies Explained
Learn what AI agent security is, understand key threats like prompt injection and tool abuse, core AI security principles, and best practices to secure AI agents.


AI 2040: Plan A
A detailed forecast and recommendation for how the US, China and the rest of the world should navigate superintelligence.

AI 2040: Plan A
A detailed forecast and recommendation for how the US, China and the rest of the world should navigate superintelligence.

AI.Gov | President Trump's AI Strategy and Action Plan
Explore President Trump’s AI initiatives focused on innovation, infrastructure, international engagement, and youth education in artificial intelligence.

ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents must explore, infer goals, build internal models of environment dynamics, and plan effective action sequences without explicit instructions. Like its predecessors ARC-AGI-1 and 2, ARC-AGI-3 focuses entirely on evaluating fluid adaptive efficiency on novel tasks, while avoiding language and external knowledge. ARC-AGI-3 environments only leverage Core Knowledge priors and are difficulty-calibrated via extensive testing with human test-takers. Our testing shows humans can solve 100% of the environments, in contrast to frontier AI systems which, as of March 2026, score below 1%. In this paper, we present the benchmark design, its efficiency-based scoring framework grounded in human action baselines, and the methodology used to construct, validate, and calibrate the environments.
Deep Agents
Using an LLM to call tools in a loop is the simplest form of an agent. This architecture, however, can yield agents that are “shallow” and fail to plan and act over longer, more complex tasks. Applications like “Deep Research”, “Manus”, and “Claude Code” have gotten around this limitation by

