







Solving a Million-Step LLM Task with Zero Errors
LLMs have achieved remarkable breakthroughs in reasoning, insights, and tool use, but chaining these abilities into extended processes at the scale of those routinely executed by humans,...

The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities (Version 1.0)
LLMs and performative productivity
It's worth asking whether LLMs are actually making us more productive at all—and if so, what we might be sacrificing in return.

ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Recent large language models (LLMs) advancements sparked a growing research interest in tool assisted LLMs solving real-world challenges…

Optimizing Agentic Workflows using Meta-tools
Agentic AI enables LLM to dynamically reason, plan, and interact with tools to solve complex tasks. However, agentic workflows often require many iterative reasoning steps and tool invocations,...

There's Something Fundamentally Wrong With LLMs
LLMs aren't trained on the "vast majority of speech," experts warn, a major blind spot that could have sweeping consequences.

Illusions of Understanding from Outsourcing Thinking to LLMs
Some illusions of understanding are an inevitable part of the research process, while others can be avoided or overcome by careful critical thinking and observation. We are facing an increased risk of avoidable illusions as more research activities are delegated to large language models (LMM). LLMs can be useful but they cannot think, and their use can undermine our thinking and understanding. Thinking for ourselves is hard and error prone but worthwhile - and there are no shortcuts to understanding.
The scientific case for being nice to your chatbot
New research confirms that LLMs often perform better when you encourage them. But why?

Dan Shipper 📧 on Twitter / X
this is true and is a big reason why you don’t need to be a highly technical researcher to use LLMs in surprising and novel ways https://t.co/TuxNzXzToU— Dan Shipper 📧 (@danshipper) July 27, 2025
Cognitive exponents and LLM leverage
I know a few people for whom LLMs have been a near-immediate multiplier of attention and effort. I know a lot for whom LLMs clearly make them worse at thinking and doing things. So: why?
The LLM Critics Are Right. I Use LLMs Anyway.
I almost agree with all of the LLM critics, yet I still use LLMs a lot. I know this sounds like I am delusional, but I don't think I am alone with it.

The LLM Critics Are Right. I Use LLMs Anyway.
I almost agree with all of the LLM critics, yet I still use LLMs a lot. I know this sounds like I am delusional, but I don't think I am alone with it.

Fine-Tuning LLMs is a Huge Waste of Time
People think they can use Fine-Tune for Knowledge Injection. People are Wrong


I feel like using LLMs to flag intent/semantics mismatches (eg do var names seem to match what you actually do) or indirect violations of API contracts or whatever as a supplement to static analysis to produce *better* code would be at least as high impact, and it's like 1% of the discourse