







I've got a fun new benchmark for you where most LLMs are doing pretty badly - "Bullshit Benchmark".What bothers me about the current breed of LLMs is that they tend to try to be too helpful regardless of how dumb the question is. So I've built 55 'bullshit' questions that don't… pic.twitter.com/4o4quN5EFR— Peter Gostev (@petergostev) February 24, 2026
One year as an AI Engineer: The 5 biggest misconceptions about LLM reliability I've encountered
535 votes, 59 comments. After spending a year building evaluation frameworks and debugging production LLM systems, I've noticed the same…
Dan Shipper 📧 on Twitter / X
this is true and is a big reason why you don’t need to be a highly technical researcher to use LLMs in surprising and novel ways https://t.co/TuxNzXzToU— Dan Shipper 📧 (@danshipper) July 27, 2025
LLMs believe false statements even after explicit warnings that they're false
Fine-tuning tests show "bias... toward confidently representing the claims as true."

The LLM Critics Are Right. I Use LLMs Anyway.
I almost agree with all of the LLM critics, yet I still use LLMs a lot. I know this sounds like I am delusional, but I don't think I am alone with it.

The LLM Critics Are Right. I Use LLMs Anyway.
I almost agree with all of the LLM critics, yet I still use LLMs a lot. I know this sounds like I am delusional, but I don't think I am alone with it.

When benchmarks go bad - what I learned from measuring performance wrong - Holly Cummins
The world of performance analysis is littered with flawed claims, cognitive biases, dangerous intuitions, and beguiling fallacies. Sadly…

There's Something Fundamentally Wrong With LLMs
LLMs aren't trained on the "vast majority of speech," experts warn, a major blind spot that could have sweeping consequences.

This Web Tool Sabotages AI Chatbots By Making Them Really, Really Slow
Artist Sam Lavigne created ‘Slow LLM’ to make people question their dependence on tools like Claude and ChatGPT. Or at least, make them super annoying to use.

My AI Skeptic Friends Are All Nuts
My smartest friends have bananas arguments about LLM coding.

My AI Skeptic Friends Are All Nuts
My smartest friends have bananas arguments about LLM coding.

LLM Leaderboard 2026 — Compare Top AI Models
Compare the latest LLM benchmarks for GPT, Claude, Gemini and more. Updated rankings across reasoning, coding, math, and multilingual tasks with pricing and speed data.
The scientific case for being nice to your chatbot
New research confirms that LLMs often perform better when you encourage them. But why?

Andrew Ho on Twitter / X
Not really observations that others before me haven’t made, but:- Despite the seemingly magical nature of LLMs, reflection over a >3 month timescale suggests my total productivity hasn’t increased by over 100%, or perhaps even by over 50%, and a lot of time is actually wasted…— Andrew Ho (@andrewho03) September 3, 2026
I read this result as: LLMs do more bullshit citations, name-dropping without engaging.
infoDOCKET
Citing Less Critically: #LLMs Reshape the Rhetoric and Reach of #Scientific #Citation (New Research Article (preprint); via @arxiv.bsky.social) arxiv.org/abs/2609.01432 #scholcomm #citations #libraries #AI #GenAI
I feel like using LLMs to flag intent/semantics mismatches (eg do var names seem to match what you actually do) or indirect violations of API contracts or whatever as a supplement to static analysis to produce *better* code would be at least as high impact, and it's like 1% of the discourse