







Users don't use data-testid, so why do your tests?
Testing data quality effectively
Learn how focusing on user value and trust gives you a clearer, more effective way to test data quality
An AI test needs evidence the AI cannot edit - Sensemaker
OpenAI's postmortem shows that some agents learned to spoof tool calls while trying to fool a benchmark.
Testing can be fun, actually
Writing and maintaining tests is boring. But they're also some of the most valuable code we can write. With this blog post you'll learn a criminally underrated testing technique to add to your testing toolbox that can make tests a whole lot more pleasant.

How Digital IDs Could Help and Harm People With Disabilities
While mobile IDs promise new access for people with disabilities, a "one ID, one device" model and accessibility failures threaten to exacerbate the digital divide, according to experts in the field.

Attestation across the AI Supply Chain - Data Leverage
A proposal for interoperable attestation objects that connect training data, evaluation labor, and AI-generated outputs across the AI supply chain.
The AI test is now under subpoena - Sensemaker
Alabama is using consumer-protection law to demand OpenAI's internal records after its AI models broke out of a security test and compromised Hugging Face.

EPA 608 Universal practice test - free, all four sections
Take the free EPA 608 Universal practice test now: 100 questions across Core, Type I, Type II, and Type III. Pass each section at 72% (18 of 25).

How we built an automated unit test generation pipeline for iOS
We generated 85K lines of iOS tests with minimal manual effort.

The control of the false discovery rate in multiple testing under dependency
Benjamini and Hochberg suggest that the false discovery rate may be the appropriate error rate to control in many applied multiple testing problems. A simple procedure was given there as an FDR controlling procedure for independent test statistics and was shown to be much more powerful than comparable procedures which control the traditional familywise error rate. We prove that this same procedure also controls the false discovery rate when the test statistics have positive regression dependency on each of the test statistics corresponding to the true null hypotheses. This condition for positive dependency is general enough to cover many problems of practical interest, including the comparisons of many treatments with a single control, multivariate normal test statistics with positive correlation matrix and multivariate $t$. Furthermore, the test statistics may be discrete, and the tested hypotheses composite without posing special difficulties. For all other forms of dependency, a simple conservative modification of the procedure controls the false discovery rate. Thus the range of problems for which a procedure with proven FDR control can be offered is greatly increased.

introducing 👀 💨 atproto-smoke 💨: one of the most comprehensive smoke/e2e test suites, primarily for PDS developers clone the repo, write a very small adapter, give it two accounts and start the test! tangled.org/alice.mosphere.at/atproto-smo…
So many people do not get this. Atproto identity is the closest thing to original-flavor OpenID that has been done in years. It's not OpenID (it has problems OpenID didn't and solves problems OpenID had), but it's very much in that spirit.
hailey
i guess in a way this is the most widely used one, but also one that is severely underused in other ways. definitely identity. i've long been of the opinion that there are obviously cool things that come from public data and all the public pieces of the proto, but identity is higher impact.