







Your LLM Doesn't Write Correct Code. It Writes Plausible Code.
One of the simplest tests you can run on a database:

LLMs believe false statements even after explicit warnings that they're false
Fine-tuning tests show "bias... toward confidently representing the claims as true."

LLM cliché highlighter
Paste text below — or load it from a URL — and it highlights sentences that match known LLM clichés, including the tells catalogued in Wikipedia’s Signs of AI writing guide. Chain patterns like “no X, no Y” get a badge counting their items; tap or hover a highlight to see which cliché it hit.
CLM - New claim for sharing
nodeTypeId: nodeKSiOfvkjvS-H6MHfDUdg nodeInstanceId: 019d688f-89ff-7119-a0df-d07aa0f3340f publishedToGroups:
test highlights - underreacted
Hypothesis: A new approach to property-based testing
The property-based testing library for Python
Proof of Work - The Marshmallow Test got a response. Now the real test begins.
After my critique of Bluesky's 2026 Report resonated across the Atmosphere, the company's new interim CEO reached out. We talked about what went wrong, what collaboration should look like, and why private data is the next test of whether the words match the work.

Proof of Work - The Marshmallow Test got a response. Now the real test begins.
After my critique of Bluesky's 2026 Report resonated across the Atmosphere, the company's new interim CEO reached out. We talked about what went wrong, what collaboration should look like, and why private data is the next test of whether the words match the work.

Using LLM in the shebang line of a script
This comment on Hacker News inspired me to investigate patterns for using my LLM CLI tool in a shebang line:

Writing Test Evals For Our MCP Server - Neon Blog
When we launched our MCP server, we knew it’d be important for it to have tests, just like any other piece of software. Since our MCP server has over 20 tools, it’s important for us to know that LLMs can pick the right tool for the job. So, this was the main aspect we wanted […]

An AI test needs evidence the AI cannot edit - Sensemaker
OpenAI's postmortem shows that some agents learned to spoof tool calls while trying to fool a benchmark.
The /llms.txt file – llms-txt
A proposal to standardise on using an /llms.txt file to provide information to help LLMs use a website at inference time.

Take caution in using LLMs as human surrogates | PNAS
Recent studies suggest large language models (LLMs) can generate human-like responses, aligning with human behavior in economic experiments, survey...

This morning at #ic2s2 I have the chance to present ongoing work on applying a SEM from survey methods to LLM text annotations: 👉 Evaluating LLM Text Annotations Without Ground Truth 📍10:45, Mansfield (210), Understanding LLMs w/ @maximiliankreutner.bsky.social Alex Cernat & @mstrohm.bsky.social
When I worked in software dev, I was taught to check three non-English languages when testing text strings: German to test the largest possible version of a string, Japanese or Chinese to test the shortest possible version, and Thai to check the TALLEST possible version.
Announcing a new version of our 2024 paper on linguistic hypothesis generation from LMs! @najoung.bsky.social and I have systematized our hypothesis generation framework, added stringent criteria for model selection, 10x-ed our learning trials, and included an epigraph from Jeff Elman 🙏!