







Stanford CS336 Language Modeling from Scratch I 2025
Universal Decompositional Semantics on Universal Dependencies
We present a framework for augmenting data sets from the Universal Dependencies project with Universal Decompositional Semantics. Where the Universal Dependencies project aims to provide a syntactic a
The Stanford NLP Group
Performing groundbreaking Natural Language Processing research since 1999.

The last six months in LLMs in five minutes
I put together these annotated slides from my five minute lightning talk at PyCon US 2026, using the latest iteration of my annotated presentation tool. # I presented this lightning …

Becoming unLLMable
slides + transcript of my presentation at the Sana AI Summit in Stockholm

Transmogrify Record - Lexicon Garden
Transform records between ATProtocol lexicon schemas
Differences in Text Generated by Diffusion and Autoregressive Language Models
Diffusion language models (DLMs) are promising alternatives to autoregressive language models (ARMs), yet the intrinsic differences in their generated text remain underexplored. We first find empirically that off-the-shelf DLMs exhibit lower $n$-gram entropy, higher semantic coherence, and higher semantic diversity. To understand the cause, we conduct controlled experiments that decouple the effects of training objectives and decoding algorithms. Results suggest that the DLM training objective contributes to the increases in semantic coherence and semantic diversity, but has a minor influence on entropy. These differences are primarily driven by the bidirectional context; other components in the training objective, such as input masking, label masking, and the weighting function, have a much weaker influence. Further, our experiments demonstrate that the reduction in entropy stems from DLMs' decoding algorithms, particularly confidence-based remasking strategies. We provide a theoretical understanding for this entropy reduction phenomenon. Together, our work uncovers key mechanisms underlying the differences between DLMs and ARMs in text generation, and informs future design of training objectives and decoding algorithms in DLMs.

Packrat parsing: | Proceedings of the seventh ACM SIGPLAN international conference on Functional programming
For decades we have been using Chomsky's generative system of grammars, particularly context-free grammars (CFGs) and regular expressions (REs), to express the syntax of programming languages and protocols. The power of generative grammars to express ...

Generating event descriptions under syntactic and semantic constraints
With the goal of supporting scalable lexical semantic annotation, analysis, and theorizing, we conduct a comprehensive evaluation of different methods for generating event descriptions under both synt
Structured Outputs with Will Kurt and Cameron Pfiffer - Weaviate Podcast #119!
Named Entity Recognition with NLTK and SpaCy
NER is used in many fields in Natural Language Processing (NLP)

LLM in a Flash: Efficient Large Language Model Inference with Limited Memory
Large language models (LLMs) are central to modern natural language processing, delivering exceptional performance in various tasks…

Thank you @steefjan.bsky.social for this nice write-up of my keynote at QCon London this week: infoq.com/news/2026/03/qcon-local-first… My slides are now up as well: speakerdeck.com/ept/mitigating-geopolitical-r…
QCon London 2026: Kleppmann on Mitigating Europe's Cloud Dependency with Local-First Software
www.infoq.com