







For fans of Talkie-1930 and all the methodological questions raised by historical models, here's a new entrant into the field, TypewriterLM, trained up to 1913. Corpus, instruction-tuning datasets, and event dataset are released. arxiv.org/abs/2606.02991
Pretraining Language Models on Historical Text
arxiv.orgJun 10, 2026 at 5:27 PM
Introducing talkie: a 13B vintage language model from 1930
This is a 24/7 live feed of Claude Sonnet 4.6 prompting talkie-1930-13b-it in order to explore its knowledge, capabilities, and inclinations. talkie’s outputs reflect the culture and values of the texts it was trained on, not the views of its authors.
Large Language Models As The Tales That Are Sung
Gene Wolfe, Albert Lord, machine culture.

Vintage chatbot lives in the past like an elderly relative
: Talkie's training data stops at the end of 1930, and its creators hope it'll help us better understand how AI thinks

Models overview
Claude is a family of state-of-the-art large language models developed by Anthropic. This guide introduces the available models and compares their performance.
Do LLMs write like humans? Variation in grammatical and rhetorical styles
As large language models (LLMs) have grown in power and become more widely available, research has focused on their ability to complete various tasks and the biases they exhibit when doing so. In this study, we instead examine their writing style in detail. We show that instruction-tuned models, which are trained to answer questions and solve problems, have a distinct noun-heavy, informationally dense writing style, even when prompted to match the style of informal speech and writing. These findings suggest that instruction-tuned models generate text that does not align with genre conventions familiar to human audiences, and demonstrate the value of linguistic variables in evaluating the output of LLMs., Large language models (LLMs) are capable of writing grammatical text that follows instructions, answers questions, and solves problems. As they have advanced, it has become difficult to distinguish their output from human-written text. While past research has found some differences in features such as word choice and punctuation and developed classifiers to detect LLM output, none has studied the rhetorical styles of LLMs. Using several variants of Llama 3 and GPT-4o, we construct two parallel corpora of human- and LLM-written texts from common prompts. Using Douglas Biber’s set of lexical, grammatical, and rhetorical features, we identify systematic differences between LLMs and humans and between different LLMs. These differences persist when moving from smaller models to larger ones and are larger for instruction-tuned models than base models. This observation of differences demonstrates that despite their advanced abilities, LLMs struggle to match human stylistic variation. Attention to more advanced linguistic features can hence detect patterns in their behavior not previously recognized.

Early Tools for Thought, Mark Bernstein @ Tools For Thought Rocks
In Which A Book Lexicon is Born
Hi everyone, We recently held a sync to discuss creating a shared book lexicon for the Atmosphere Protocol. I’ve shared the full recording (thanks Ronen!) and notes from the meeting here: Fathom with some early notes here: https://app.notion.com/p/Book-Lexicon-Session-1-211b2b43aabb8007b098d063972ca58c (also do we have a ATProto version of Notion yet?) The book logging social scene is highly fragmented. Dozens of Goodreads alternatives are competing but not interoperating, leading to a lack of...

The LLMentalist Effect: how chat-based Large Language Models rep…
The new era of tech seems to be built on superstitious behaviour


10 things I want to work on after the conference - Alex's Blog
a few people found this tool I made a while ago useful in @alex.bsky.team's lexicon workshop yesterday at #atmosphereconf, so resharing here for anyone else who wants to write their lexicons in typescript prototypey.org
Mapping the Mind of a Large Language Model
We have identified how millions of concepts are represented inside Claude Sonnet, one of our deployed large language models. This is the first ever detailed look inside a modern, production-grade large language model.

Announcing a new version of our 2024 paper on linguistic hypothesis generation from LMs! @najoung.bsky.social and I have systematized our hypothesis generation framework, added stringent criteria for model selection, 10x-ed our learning trials, and included an epigraph from Jeff Elman 🙏!
1. We—@eduede.bsky.social, @mjcrockett.bsky.social, Kevin Gross, and I—have a new preprint on the arXiv today, based on ideas that emerged during an @sfiscience.bsky.social workshop in November 2024: The unintended consequences of large language models as a labor-augmenting technology in science.
The unintended consequences of large language models as a labor-augmenting technology in science
arxiv.orga few people found this tool I made a while ago useful in @alex.bsky.team's lexicon workshop yesterday at #atmosphereconf, so resharing here for anyone else who wants to write their lexicons in typescript prototypey.org