







535 votes, 59 comments. After spending a year building evaluation frameworks and debugging production LLM systems, I've noticed the same…
If You’re Going To Defend AI And Whine About Its Critics, You Should Probably Be Honest About Its Actual Harms
I think this recent post by AI industry CEO Matt Shumer is worth a read. In it, he basically explains how quickly LLMs (large language models) are evolving to supplant many developers and prog…

Solving a Million-Step LLM Task with Zero Errors
LLMs have achieved remarkable breakthroughs in reasoning, insights, and tool use, but chaining these abilities into extended processes at the scale of those routinely executed by humans,...

My AI Skeptic Friends Are All Nuts
My smartest friends have bananas arguments about LLM coding.

My AI Skeptic Friends Are All Nuts
My smartest friends have bananas arguments about LLM coding.

The LLM Critics Are Right. I Use LLMs Anyway.
I almost agree with all of the LLM critics, yet I still use LLMs a lot. I know this sounds like I am delusional, but I don't think I am alone with it.

The LLM Critics Are Right. I Use LLMs Anyway.
I almost agree with all of the LLM critics, yet I still use LLMs a lot. I know this sounds like I am delusional, but I don't think I am alone with it.

Best LLM for Coding 2026 | AI Coding Model Rankings & Benchmarks
Which AI model writes the best code? We rank every major LLM — open and closed source — across SWE-bench, HumanEval, LiveCodeBench, and Terminal-Bench coding benchmarks. Compare the best LLMs for coding, software engineering, and programming.

The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities (Version 1.0)
There's Something Fundamentally Wrong With LLMs
LLMs aren't trained on the "vast majority of speech," experts warn, a major blind spot that could have sweeping consequences.

The two worlds of programming: why developers who make the same observations about LLMs come to opposite conclusions
Writing at the end of the world, from Hveragerði, Iceland
KillBench: Discovering Hidden Biases of LLMs
1M+ experiments exposing bias in critical AI decision-making

Peter Gostev on Twitter / X
I've got a fun new benchmark for you where most LLMs are doing pretty badly - "Bullshit Benchmark".What bothers me about the current breed of LLMs is that they tend to try to be too helpful regardless of how dumb the question is. So I've built 55 'bullshit' questions that don't… pic.twitter.com/4o4quN5EFR— Peter Gostev (@petergostev) February 24, 2026

Have we been measuring AI political bias wrong? A better approach is possible.
Why ideological preferences and epistemic failure in LLMs are not the same thing — and why the difference matters

Good design hasn’t changed with AI — John Pham, SF Compute
State of AI 2025: 100T Token LLM Usage Study | OpenRouter
Read OpenRouter's 2025 State of AI report — an empirical 100 trillion token study of real LLM usage, model trends, and developer insights.