







We've been experimenting with running coding agents autonomously for weeks at a time.
My Agentic Coding Setup, July 2026
With the right tools, coding agents can run autonomously, in parallel, and be reachable from anywhere. Here's how I've stitched them together.
Cursor: AI coding agent
Built to make you extraordinarily productive, Cursor is the best AI coding agent.

Towards self-driving codebases · Cursor
We're making a part of our multi-agent research harness available to try today in preview.

Open Coding Agents: Fast, accessible coding agents that adapt to any repo | Ai2
SERA is the first in our family of Open Coding Agents, achieving state-of-the-art performance at low cost.

PrimeIntellect-ai/prime-agent
A self-improving RLM agent for coding workflows and long-running autonomous tasks.

Agent Skills
AI coding agents take the shortest path to done, which usually means skipping the specs, tests, and reviews that make software reliable at scale. Agent Skill...

AI Coding Agent Benchmarks & Leaderboard | Artificial Analysis
We measure real-world performance of coding agents on software engineering tasks, including cost, token usage, and execution time. We compare how performance changes across agents, models, and execution settings.

How to harness AI
"Coding agents" are complicated but intelligible

OpenCode Zen | A curated set of reliable optimized models for coding agents
OpenCode - The open source coding agent.

Cline - AI Coding, Open Source and Uncompromised
Open-source AI coding agent with Plan/Act modes, MCP integration, and terminal-first workflows. Trusted by 5M+ developers worldwide.

Pi: The Minimal Agent Within OpenClaw
A gentle introduction to the Pi coding agent and why I think it’s a glimpse into the future of software.

Prime Agent: A self-improving RLM agent
Prime Agent is our open-source, self-improving coding harness built around two abstractions: the Recursive Language Model (RLM) and the Continual Harness. With Opus 5, it achieves 95.5% on ARC-AGI-3, surpassing the reported human expert baseline.

Rebuilding Cognition's Agentic MapReduce
How do you run large-scale agent tasks across a codebase?
