







An actually actionable guide on how to do them
Eval awareness in Claude Opus 4.6’s BrowseComp performance
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

On evaluating agents – aunhumano
No amount of evals will replace the need to look at the data, once you have a evals good coverage you’ll be able to decrease the time but it’ll be always a must to just look at the agent traces to identify possible issues or things to improve.
evalstate/fast-agent
Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP Support

Writing Test Evals For Our MCP Server - Neon Blog
When we launched our MCP server, we knew it’d be important for it to have tests, just like any other piece of software. Since our MCP server has over 20 tools, it’s important for us to know that LLMs can pick the right tool for the job. So, this was the main aspect we wanted […]


Introducing eve
Introducing eve, the open-source agent framework from Vercel for building, running, and scaling agents in production, with durable execution, sandboxed compute, approvals, channels, tracing, and evals built in.

Operads, Type Level Nats, and Tic-Tac-Toe
This summer I spent some time talking with Edward Kmett about lots of things. (Which really means that he was talking and I was trying to keep up.) One of the topics was operads. The ideas behind o…

How our agents build on-brand pages with design.md
How we built design.md, a single public file any coding agent can load to build on-brand Vercel pages, and the eval loop that decided every rule inside it.

Animating the 6 Basic Emotions: Acting Tips from Pro Animators | Animation Mentor
Get animation tips from our blog series on Animating the 6 Basic Emotions. Polish your demo reel with advice from professional animators!

relay-eval
relay-eval.waow.techSmythOS
pguso/ai-agents-from-scratch

Agentation

Gas Town’s Agent Patterns, Design Bottlenecks, and Vibecoding at Scale
Cognitive Integrity Framework: Formal Foundations for Multiagent Security
Rowboat - Let AI build your agents for you