







Join the cookout
Writing Test Evals For Our MCP Server - Neon Blog
When we launched our MCP server, we knew it’d be important for it to have tests, just like any other piece of software. Since our MCP server has over 20 tools, it’s important for us to know that LLMs can pick the right tool for the job. So, this was the main aspect we wanted […]

A universal approach to mocking
Defunctionalise your continuations and your tests can run any computation a step at a time.

All the ways to mock your Rust code
Ok, so, suppose you’ve written a Kubernetes controller, you did it in Rust, and then you realized that “Hey, when this thing breaks it’s hard to understand what’s going on”. So, you decide to emit some Kubernetes events that give a little more context as to what your controller is doing and what steps it’s taken. You run all your tests

Hypothesis: A new approach to property-based testing
The property-based testing library for Python
Integration test setup for firehose websocket · Issue #26 · DavidBuchanan314/millipds
Probably want a async with wrapper that connects to the firehose and parses and stores the messages it receives, so that they can be checked for expected results afterwards. (will probably want to ...
Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing
SUMMARY The common approach to the multiplicity problem calls for controlling the familywise error rate (FWER). This approach, though, has faults, and we point out a few. A different approach to problems of multiple significance testing is presented. It calls for controlling the expected proportion of falsely rejected hypotheses — the false discovery rate. This error rate is equivalent to the FWER when all hypotheses are true but is smaller otherwise. Therefore, in problems where the control of the false discovery rate rather than that of the FWER is desired, there is potential for a gain in power. A simple sequential Bonferronitype procedure is proved to control the false discovery rate for independent test statistics, and a simulation study shows that the gain in power is substantial. The use of the new procedure and the appropriateness of the criterion are illustrated with examples.

Testing can be fun, actually
Writing and maintaining tests is boring. But they're also some of the most valuable code we can write. With this blog post you'll learn a criminally underrated testing technique to add to your testing toolbox that can make tests a whole lot more pleasant.

How we built an automated unit test generation pipeline for iOS
We generated 85K lines of iOS tests with minimal manual effort.

The control of the false discovery rate in multiple testing under dependency
Benjamini and Hochberg suggest that the false discovery rate may be the appropriate error rate to control in many applied multiple testing problems. A simple procedure was given there as an FDR controlling procedure for independent test statistics and was shown to be much more powerful than comparable procedures which control the traditional familywise error rate. We prove that this same procedure also controls the false discovery rate when the test statistics have positive regression dependency on each of the test statistics corresponding to the true null hypotheses. This condition for positive dependency is general enough to cover many problems of practical interest, including the comparisons of many treatments with a single control, multivariate normal test statistics with positive correlation matrix and multivariate $t$. Furthermore, the test statistics may be discrete, and the tested hypotheses composite without posing special difficulties. For all other forms of dependency, a simple conservative modification of the procedure controls the false discovery rate. Thus the range of problems for which a procedure with proven FDR control can be offered is greatly increased.

Testing functional UIs - Hayleigh Thompson | Lambda Days 2025
Structure and Interpretation of Test Cases • Kevlin Henney • GOTO 2022
@buildthis.bisks.net build an at proto enabled app that lets user count how many tacos they ate. Taco records should be aggregated through constellation.blue. Create a leaderboard-like UI, use the spaces pds to store data and let people create private leaderboards