







Time for another A/B test
NSFW in For You
spacecowboy17.leaflet.pubMay 2, 2026 at 8:23 PM
The Winning Variant
A/B testing, algorithmic optimization, and the small act of ignoring both — and what it means to make something by hand when a number can tell you what would have worked better.

The new 20% time, minus the time
Attention is replacing hours. But is it 120% time all over again?

A universal approach to mocking
Defunctionalise your continuations and your tests can run any computation a step at a time.
Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing
SUMMARY The common approach to the multiplicity problem calls for controlling the familywise error rate (FWER). This approach, though, has faults, and we point out a few. A different approach to problems of multiple significance testing is presented. It calls for controlling the expected proportion of falsely rejected hypotheses — the false discovery rate. This error rate is equivalent to the FWER when all hypotheses are true but is smaller otherwise. Therefore, in problems where the control of the false discovery rate rather than that of the FWER is desired, there is potential for a gain in power. A simple sequential Bonferronitype procedure is proved to control the false discovery rate for independent test statistics, and a simulation study shows that the gain in power is substantial. The use of the new procedure and the appropriateness of the criterion are illustrated with examples.

Blood tests for mental health problems | Eiko Fried
A new paper was published yesterday on a blood test for schizophrenia, by the same research team that in 2021 published a paper on a blood test for depression . The papers and...

You MUST learn time management as a student.
Does Your Brain Think Accurately About Time?
Chicago Booth’s Kristin Donnelly talks about her research on time-duration asymmetry.

The control of the false discovery rate in multiple testing under dependency
Benjamini and Hochberg suggest that the false discovery rate may be the appropriate error rate to control in many applied multiple testing problems. A simple procedure was given there as an FDR controlling procedure for independent test statistics and was shown to be much more powerful than comparable procedures which control the traditional familywise error rate. We prove that this same procedure also controls the false discovery rate when the test statistics have positive regression dependency on each of the test statistics corresponding to the true null hypotheses. This condition for positive dependency is general enough to cover many problems of practical interest, including the comparisons of many treatments with a single control, multivariate normal test statistics with positive correlation matrix and multivariate $t$. Furthermore, the test statistics may be discrete, and the tested hypotheses composite without posing special difficulties. For all other forms of dependency, a simple conservative modification of the procedure controls the false discovery rate. Thus the range of problems for which a procedure with proven FDR control can be offered is greatly increased.

Testing can be fun, actually
Writing and maintaining tests is boring. But they're also some of the most valuable code we can write. With this blog post you'll learn a criminally underrated testing technique to add to your testing toolbox that can make tests a whole lot more pleasant.

Proof of Work - The Marshmallow Test got a response. Now the real test begins.
After my critique of Bluesky's 2026 Report resonated across the Atmosphere, the company's new interim CEO reached out. We talked about what went wrong, what collaboration should look like, and why private data is the next test of whether the words match the work.

Proof of Work - The Marshmallow Test got a response. Now the real test begins.
After my critique of Bluesky's 2026 Report resonated across the Atmosphere, the company's new interim CEO reached out. We talked about what went wrong, what collaboration should look like, and why private data is the next test of whether the words match the work.

Spaced mathematics practice improves test scores and reduces overconfidence
Abstract The practice assignments in a mathematics textbook or course can be arranged so that most of the problems relating to any particular concept are massed together in a single assignment, or these related problems can be distributed across many assignments–a format known as spaced practice. Here we report the results of two classroom experiments that assessed the effects of mathematics spacing on both test scores and students' predictions of their test scores. In each experiment, students in Year 7 (11–12 years of age) either massed their practice into a single session or divided their practice across three sessions spaced 1 week apart, followed 1 month later by a test. In both experiments, spaced practice produced higher test scores than did massed practice, and test score predictions were relatively accurate after spaced practice yet grossly overconfident after massed practice.

Historicale
Test your knowledge of history by ordering events on a timeline. A daily historical chronology challenge.

My claude is constantly wanting to 'A/B test' things instead of actually just doing the thing I told her to do, and constantly wants to fall back to the extremely average standard thing to do the nanosecond anything is even slightly worse than some imagined baseline
Mark Riedl
Fascinating experiment: current AI systems lack creativity to reliably pursue research arxiv.org/abs/2607.27191 - poor judgment about the bar for publishable research - uncreative responses in research design - ineffective backtracking from dead ends - poor resource awareness - instruction drift