







‘Quite simply the best historical dictionary of English slang there is, ever has been […] or is ever likely to be’ — Journal of English Language and Linguistics
Online Etymology Dictionary
The online etymology dictionary (etymonline) is the internet's go-to source for quick and reliable accounts of the origin and history of English words, phrases, and idioms.

It's giving incel: The evolution of internet slang : Code Switch
How have recommendation algorithms affected language? Linguist Adam Aleksic — aka the Etymology Nerd — says most “Gen-Z slang” is either appropriated from Black people or incels. This week, we trace how -maxxing went from the eugenicist looksmaxxing subculture to trending TikToks to the Pentagon tweeting about “lethality maxxing.” And we ask what’s actually at stake when we use words without knowing where they come from.

An AT Proto lexicon currently under consideration and development
An AT Proto lexicon currently under consideration and development · GitHub

The Grid: A Lecture on Cybotron and Techno-Vernacular Expressionism by DeForrest Brown, Jr.

What Language is This? Ask Your Tokenizer
Language Identification (LID) is an important component of many multilingual natural language processing pipelines, where it facilitates corpus curation, training data analysis, and cross-lingual evaluation of large language models. Despite near-perfect performance on high-resource languages, existing systems remain brittle in low-resource and closely related language settings. We introduce UniLID, a simple and efficient LID method based on the UnigramLM tokenization algorithm, leveraging its probabilistic framing, parameter estimation technique and inference strategy. In short, to predict a string's language label, we simply ask: under which language's unigram distribution is this string most likely? Our formulation is data- and compute-efficient, supports incremental addition of new languages without retraining existing models, and can naturally be integrated into existing language model tokenization pipelines. Empirical evaluations against widely used baselines, including fastText, GlotLID and CLD3, show that UniLID achieves competitive performance on standard benchmarks, substantially improves sample efficiency in low-resource settings -- reaching ~70% accuracy with as few as five labeled samples per language -- and delivers large gains on fine-grained dialect identification.

The Wikipedia Revolution: How a Bunch of Nobodies Creat…
"Imagine a world in which every single person on the pl…

Road to a book lexicon (Part 2) - Nick The Sick
The one where I actually define a lexicon

2025's Word of the Year (Half the Answer #60, with Savi Namboodiripad and Daniel Midgley)
Caitlin and Trent chat with linguistics scholars Savi Namboodiripad and Daniel Midgley on what word most exemplifies 2025. Was it a new coinage? A pre-existing word that gained a new meaning or usage this year? A word that seemed to be on everybody's lips in 2025? Half the
The Depths of Wikipedians—Asterisk
A conversation about yogurt wars, German hymns, tropical cyclones, and the people who make Wikipedia function.

EsDictionary - The EsDeeKid Translator
Th relationship between domains and lexicons seems like a short-term win, but it remains unclear for me if it’s good in the long-term.
atproto.brussels 🇧🇪
Loosely related question I have about lexicons being tied to a domain. Domains are not free. So, say your lexicon becomes the industry standard and thousands use it across many apps. It seems fragile to me that the domain belongs to an individual and not a public entity. Is there a solution to that?