







The world of performance analysis is littered with flawed claims, cognitive biases, dangerous intuitions, and beguiling fallacies. Sadly…
Why it’s getting harder to measure AI performance
The most famous chart in AI might be obsolete soon.

When Using AI, Users Fall for the Dunning-Kruger Trap in Reverse - Neuroscience News
A new study reveals that when interacting with AI tools like ChatGPT, everyone—regardless of skill level—overestimates their performance.

The Optimization Trap: Why Too Much Efficiency Makes Us Fragile with Olivier Hamant
People Reject Algorithms in Uncertain Decision Domains Because They Have Diminishing Sensitivity to Forecasting Error
Will people use self-driving cars, virtual doctors, and other algorithmic decision-makers if they outperform humans? The answer depends on the uncertainty inherent in the decision domain. We propose that people have diminishing sensitivity to forecasting error and that this preference results in people favoring riskier (and often worse-performing) decision-making methods, such as human judgment, in inherently uncertain domains. In nine studies ( N = 4,820), we found that (a) people have diminishing sensitivity to each marginal unit of error that a forecast produces, (b) people are less likely to use the best possible algorithm in decision domains that are more unpredictable, (c) people choose between decision-making methods on the basis of the perceived likelihood of those methods producing a near-perfect answer, and (d) people prefer methods that exhibit higher variance in performance (all else being equal). To the extent that investing, medical decision-making, and other domains are inherently uncertain, people may be unwilling to use even the best possible algorithm in those domains.

Peter Gostev on Twitter / X
I've got a fun new benchmark for you where most LLMs are doing pretty badly - "Bullshit Benchmark".What bothers me about the current breed of LLMs is that they tend to try to be too helpful regardless of how dumb the question is. So I've built 55 'bullshit' questions that don't… pic.twitter.com/4o4quN5EFR— Peter Gostev (@petergostev) February 24, 2026
Underspecified Human Decision Experiments Considered Harmful
Decision-making with information displays is a key focus of research in areas like human-AI collaboration and data visualization. However, what constitutes a decision problem, and what is required for an experiment to conclude that decisions are flawed, remain imprecise. We present a widely applicable definition of a decision problem synthesized from statistical decision theory and information economics. We claim that to attribute loss in human performance to bias, an experiment must provide the information that a rational agent would need to identify the normative decision. We evaluate whether recent empirical research on AI-assisted decisions achieves this standard. We find that only 10 (26%) of 39 studies that claim to identify biased behavior presented participants with sufficient information to make this claim in at least one treatment condition. We motivate the value of studying well-defined decision problems by describing a characterization of performance losses they allow to be conceived.

Standards around generative AI
Accuracy, fairness and speed are the guiding values for AP’s news report, and we believe the mindful use of artificial intelligence can serve these values and over time improve how we work.
Cognitive Bias Lab | Learn to Make Better Decisions
Explore cognitive biases with interactive tests, simulations, and real-world examples. Free platform to sharpen decision-making and critical thinking — no sign-up needed.

Algorithmic Bias · Open Encyclopedia of Cognitive Science
Algorithmic bias refers to prejudicial, discriminatory, unjust, inaccurate, or otherwise disparate performance or outcomes from algorithmic systems based on racial, gender, or other attributes of an individual or a group. The concept of algorithmic bias emerged at the intersection of computer science, artificial intelligence (AI) research, critical data studies, human–computer interaction, law, philosophy, and similar disciplines. Although problems and discrepancies at the model level denote the most commonly studied form of bias, the term algorithmic bias is also used as a shorthand to describe a multitude of problems and challenges at various steps of the AI pipeline from ideation, problem framing, training data curation and processing, model training and validation, and deployment as well as emergent issues that arise from interaction with the real world. Potential sources of bias, appropriate metrics to define, measure, and mitigate bias, and the utility and merit of technical approaches to bias mitigation are fiercely debated in the current AI landscape.

Alex Cheema on Twitter / X
This is why we need open benchmarks for local AI.Otherwise it turns into tribalism and name calling.We will be publishing the largest database of open benchmarks for local AI, tested on 1,000+ real hardware setups. Every device, every interconnect, different… https://t.co/ZsU3PCdSsZ— Alex Cheema (@alexocheema) March 9, 2026
Unfortunately, You Need to Know What the Jevons Paradox is
Unfortunately, You Need to Know What the Jevons Paradox is
Back-to-basics: on poor conceptualizations in AI work - (Un)rigorous AI
TL;DR — Poor conceptual foundations can severely undermine the credibility and reliability of knowledge claims. (And, no, your metric is not your construct.)
When the Scoreboard Becomes the Game, It’s Time to Recalibrate Research Metrics - The Scholarly Kitchen
Today's guest post discusses research metrics and their relationship to research integrity, inclusivity, and long-term impact.


AI 2027

What will be left for us to work on?

OpenAI and Hugging Face partner to address security incident during model evaluation
Taking more seriously the claim that recent ML models do not "reason", it still is quite odd the particular ways that superhuman game-playin…

General-purpose large language models outperform specialized clinical AI tools on medical benchmarks

StoryScope: Investigating idiosyncrasies in AI fiction