







"Cramms" is not a word
Harper | Privacy-First Offline Grammar Checker
Blazing-fast, open-source grammar & spell checking that never sends your words to the cloud.

How linguistics learned to stop worrying and love the language models
Language models (LMs) can produce fluent, grammatical text. Nonetheless, some maintain that language models don’t really learn language and also, even if they did, that would not be informative for the study of human learning and processing. On the other side, there have been claims that the success of LMs obviates the need for studying linguistic theory and structure. We argue that both extremes are wrong. LMs can contribute to fundamental questions about linguistic structure, language processing, and learning. They force us to rethink arguments and ways of thinking that have been foundational in linguistics. While they do not replace linguistic structure and theory, they serve as model systems and working proofs of concept for gradient, usage-based approaches to language. We offer an optimistic take on the relationship between language models and linguistics.

👋 Jan on Twitter / X
You don't need a 32B model to fix grammar. You just need one that's actually trained to do it.GRMR-V3 (1B–4.3B) is a set of models that correct grammar reliably without touching meaning. https://t.co/7kkk39Mk7v— 👋 Jan (@jandotai) June 5, 2025
No apps no masters
Posted on Friday 9 Aug 2024. 1,415 words, 11 links. By Matt Webb.

Home - 80/20 Japanese
Finally make sense of Japanese grammar Master Japanese grammar through visual explanations that show you why the language works the way it does, so you can build your own sentences with confidence, not just memorize phrases. “ Your course is AMAZING! It’s a lifesaver to having me communicate properly with my coworkers and getting around […]

The Stanford NLP Group
Performing groundbreaking Natural Language Processing research since 1999.

atproto made simple: publishing lexicons - underreacted
atproto made simple: publishing lexicons - underreacted
Large Language Muddle | The Editors
The AI upheaval is unique in its ability to metabolize any number of dread-inducing transformations. The university is becoming more corporate, more politically oppressive, and all but hostile to the humanities? Yes — and every student gets their own personal chatbot. The second coming of the Trump Administration has exposed the civic sclerosis of the US body politic? Time to turn the Social Security Administration over to Grok. Climate apocalypse now feels less like a distant terror than a fact of life? In three years, roughly a tenth of US energy demand will come from data centers alone.

Computation and its Connotations
A Review of Language Machines by Leif Weatherby

Exploring CRDT Lexicons
I’ve (alongside many others I’m sure) been working on CRDT work with atproto, and I think we’ll need to come up with some standard for how to represent CRDT data as lexicons. @chris.pardy.family has written a leaflet exploring what such a record could look like and is looking for feedback and thoughts: CRDT's on ATProto I think it’s a good start. It doesn’t tie us to any particular CRDT, though it doesn’t guarantee interop between different CRDT types, which seems fine to me. I’m thinking may...

Do LLMs write like humans? Variation in grammatical and rhetorical styles
As large language models (LLMs) have grown in power and become more widely available, research has focused on their ability to complete various tasks and the biases they exhibit when doing so. In this study, we instead examine their writing style in detail. We show that instruction-tuned models, which are trained to answer questions and solve problems, have a distinct noun-heavy, informationally dense writing style, even when prompted to match the style of informal speech and writing. These findings suggest that instruction-tuned models generate text that does not align with genre conventions familiar to human audiences, and demonstrate the value of linguistic variables in evaluating the output of LLMs., Large language models (LLMs) are capable of writing grammatical text that follows instructions, answers questions, and solves problems. As they have advanced, it has become difficult to distinguish their output from human-written text. While past research has found some differences in features such as word choice and punctuation and developed classifiers to detect LLM output, none has studied the rhetorical styles of LLMs. Using several variants of Llama 3 and GPT-4o, we construct two parallel corpora of human- and LLM-written texts from common prompts. Using Douglas Biber’s set of lexical, grammatical, and rhetorical features, we identify systematic differences between LLMs and humans and between different LLMs. These differences persist when moving from smaller models to larger ones and are larger for instruction-tuned models than base models. This observation of differences demonstrates that despite their advanced abilities, LLMs struggle to match human stylistic variation. Attention to more advanced linguistic features can hence detect patterns in their behavior not previously recognized.

i was annoyed that this isn't explained how i want in the docs so i wrote my own mini-guide to publishing lexicons:
atproto made simple: publishing lexicons
underreacted.leaflet.pubYep. This is what is missing in the ATProto vision. The lexicons should've been a like schema.org instead of domain name-based. Everyone who wants to improve it can discuss it, similar to, well, schema.org. The frontend must be separate from the content. It's beneficial in many ways.
jack
we have standard.site but wen standard.pics? rn I’m juggling between @grain.social, @flashes.blue & @sprk.so for image sharing 😭