







"A philosophy of games to help us win back control over what we value. The philosopher C. Thi Nguyen-one of the leading experts on the philosophy of games and the philosophy of data-takes us deep into the heart of games, and into the depths of bureaucracy, to see how scoring systems shape our desires. Games are the most important art form of our era. They embody the spirit of free play. They show us the subtle beauty of action everywhere in life in video games, sports, and boardgames-but also cooking, gardening, fly-fishing, and running. They remind us that it isn't always about outcomes, but about how glorious it feels to be doing the thing. And the scoring systems help get us there, by giving us new goals to try on. Scoring systems are also at the center of our corporations and bureaucracies-in the form of metrics and rankings. They tell us exactly how to measure our success. They encourage us to outsource our values to an external authority. And they push on us to value simple, countable things. Metrics don't capture what really matters; they only capture what's easy to measure. The price of that clarity is our independence. The Score asks us is this the game you really want to be playing?"-- Provided by publisher
Introducing DAVIES: A framework for Identifying Talent Across the Globe — American Soccer Analysis
In the world of sports, the search for an all-encompassing player evaluation metric is never-ending. Baseball was the first to develop its metric with Wins Above Replacement. Basketball followed suit with Player Efficiency Rating, and Hockey WAR has come into the fold within the past year. The US So

The Score (book)
The Score: How to Stop Playing Somebody Else's Game is a book by C. Thi Nguyen.
When the Scoreboard Becomes the Game, It’s Time to Recalibrate Research Metrics - The Scholarly Kitchen
Today's guest post discusses research metrics and their relationship to research integrity, inclusivity, and long-term impact.

Index
OpenSkill: Multiplayer Rating System. No Friction. In the multifaceted world of online gaming, an accurate multiplayer rating system plays a crucial role. A multiplayer rating system measures and c...

Game Balance Isn't Real
Input goals and output goals
The other week someone introduced me to the idea of “input goals” and “output goals”, by Oz Chen. Oz writes about “personal development and content strategy”, so…

When benchmarks go bad - what I learned from measuring performance wrong - Holly Cummins
The world of performance analysis is littered with flawed claims, cognitive biases, dangerous intuitions, and beguiling fallacies. Sadly…

Integrative experiments identify how punishment affects welfare in public goods games
Despite decades of research, the conditions under which punishment promotes cooperation remain unclear. Through an integrative experiment varying 14 design parameters of public goods games across 360 experimental conditions (147,618 decisions from 7100 participants), we reveal substantial heterogeneity in punishment effectiveness: Its impact on welfare ranges from 43% improvement to 44% reduction depending on the game parameters. To characterize these patterns, we developed models that outperformed human forecasters in predicting punishment effectiveness in new experiments. Communication emerges as the most important factor, followed by contribution framing (opt out versus opt in), contribution type (variable versus all-or-nothing), game length, and outcome visibility, though these factors often interact. The results reframe the debate from whether punishment works to when it does, demonstrating how integrative experiments enable discovery of generalizable patterns in social phenomena. , Editor’s summary People face conflicts between maximizing personal gain versus supporting collective interests. If we cooperatively recycle or donate to charities, it benefits society, but it also costs us time and resources that could be selfishly preserved for ourselves. We impose penalties to deter those undesirable or selfish behaviors, but under what conditions do punishments or penalties effectively modify behavior to benefit group welfare? Alsobay et al . systematically and simultaneously varied 14 factors together instead of in isolation. Punishment was unequivocally most effective when paired with consistent communication, particularly over time. Another effective factor was “opting out” or withdrawing some, but not all, endowments already in the public fund. These methodological advances revealed when, rather than whether, punishment works. —Ekeoma Uzogara , INTRODUCTION Human societies face many situations where individual and collective interests conflict, often referred to as social dilemmas. Costly peer punishment has been studied for more than 25 years in public goods games (stylized behavioral experiments in which individuals decide how much to contribute to a shared pool that benefits everyone) as a mechanism to promote cooperation. Prior research has identified many contextual factors that moderate punishment’s effectiveness, including game length, communication, group size, punishment cost, and so on. However, the specific conditions under which punishment improves group welfare remain unclear. RATIONALE We argue that this lack of clarity derives from the dominant experimental paradigm, in which any given study manipulates only one or a few theoretically informed factors. Because such studies differ in many ways (different experimental procedures, populations), their results are often difficult to compare or integrate. Consequently, one can list many factors that have some effect, but cannot say how much each matters relative to the others, or how they work together, and as a result, cannot predict when punishment will help or harm welfare in new settings. To address this fundamental knowledge gap, we use an integrative experimental design and systematically vary 14 parameters across 360 conditions (147,618 decisions from 7100 participants) to elucidate when punishment improves versus undermines welfare in public goods games, which factors matter most, and how they interact. RESULTS The effect of punishment on welfare ranged from 43% improvement to 44% reduction depending on the specific combination of game parameters. To characterize this heterogeneity, we trained a model that outperformed all 553 human forecasters (laypeople and experts) in predicting whether punishment would help or harm welfare in new experiments. Communication emerged as roughly three times more important than any other factor, followed by contribution framing (opt in versus opt out), contribution type (variable versus all-or-nothing), game length, and peer outcome visibility (whether participants can see others’ earnings). These factors often interact. For example, longer games enhance punishment’s effectiveness only when communication is available, and contribution framing effects depend on both contribution type and outcome visibility. CONCLUSION Many phenomena in social science are shaped by many factors whose interactions are consequential, yet the dominant experimental paradigm often limits its inquiry to “does a given effect exist?” and examines hypothesized factors in isolation. As a result, research programs can accumulate many partial explanations without a clear picture of how they combine to determine outcomes across settings. Knowing that factors matter individually is fundamentally different from knowing how much each matters and how they interact. The integrative approach implemented here offers one way forward. It varies many factors simultaneously within a shared design space, evaluates models by their predictive accuracy on new experiments, and probes those models to constrain and develop theory. Our hope is that integrative experiment designs, combined with models that integrate prediction and explanation, represent a path toward more cumulative social science. Integrative experiment reveals when punishment helps versus harms. We systematically varied 14 design parameters across 360 experimental conditions. The effect of punishment on cooperation efficiency ranged from −44% to +43% depending on the specific game parameters. Communication emerged as three times more important than any other factor, followed by contribution framing, contribution type, and game length.

17. A Value for n-Person Games
17. A Value for n-Person Games was published in Contributions to the Theory of Games, Volume II on page 307.
Specification gaming: the flip side of AI ingenuity
Specification gaming is a behaviour that satisfies the literal specification of an objective without achieving the intended outcome. We have all had experiences with specification gaming, even if not by this name. Readers may have heard the myth of King Midas and the golden touch, in which the king asks that anything he touches be turned to gold - but soon finds that even food and drink turn to metal in his hands. In the real world, when rewarded for doing well on a homework assignment, a student might copy another student to get the right answers, rather than learning the material - and thus exploit a loophole in the task specification.
The hidden ‘rules of the game’ that dictate how we navigate the world | Psyche Videos
How free are we really, if human behaviour embodies the complex, intertwined webs of society and history?

On the Impact of the Utility in Semivalue-based Data Valuation
Semivalue–based data valuation uses cooperative‐game theory intuitions to assign each data point a value reflecting its contribution to a downstream task. Still, those values depend on the...
AI, Pluralism, and (Social) Compensation
One strategy in response to pluralistic values in a user population is to personalize an AI system: if the AI can adapt to the specific values of each individual, then we can potentially avoid many of the challenges of pluralism. Unfortunately, this approach creates a significant ethical issue: if there is an external measure of success for the human-AI team, then the adaptive AI system may develop strategies (sometimes deceptive) to compensate for its human teammate. This phenomenon can be viewed as a form of social compensation, where the AI makes decisions based not on predefined goals but on its human partner's deficiencies in relation to the team's performance objectives. We provide a practical ethical analysis of the conditions in which such compensation may nonetheless be justifiable.

Playing to Win: How Strategy Really Works
How Strategy Really Works. This approach grew out of the strategy practice at Monitor Company and subsequently became the standard process at P&G.

'It's rigged': The dirty secret behind what IQ scores really measure | BBC Science Focus Magazine
Besides being next to useless for actually measuring intelligence, here's why we should kill IQ scores for good