







Actual Computer's first public software release: toks, the best tokenizer on earth. The exact same ids as Hugging Face tokenizers, 13x to 151x faster on one CPU core.
Tokenization Tax 2026: LLM Tokenizer Comparison | ellamind
A measurement study of 20 LLM tokenizers across 12 languages and 21 kinds of text. On German, Claude 4.7+ needs 2.01x the tokens of the OpenAI reference, and Cohere Command A+ the fewest.
The Bitter Lesson is coming for Tokenization
Highlights the desire to replace tokenization with a general method that better leverages compute and data. We'll see tokenization's fragility and review the Byte Latent Transformer arch.

Tock OS A Rust Based Open Platform for Transparent and Secure Root of Trust Devices
Fastest 1000000 tokens
Introducing the Tolaria Alliance! 🦸♂️
A small set of tools that power my Tolaria coding workflow, and also fund my work!

Road to Tonk Substrate — Tonk
We're building a living, malleable software environment that allows you to shape, own, and share your digital world without asking anyone's permission.

BytePlus Free Trial: 4K AI Image/Video Models + LLM Tokens | 200 Free Images
BytePlus free trial, Seedream 4.0, Seedance 1.0, DeepSeek V3.1, GPT-OSS-120B, free 4K AI images, AI video generation, BytePlus LLM tokens

ZTA: Zero Token Architecture - Kelsey Hightower | PlatformCon 2026
Toga
Toga is a Python native, OS native, cross-platform GUI toolkit. Toga consists of a library of base components with a shared interface to simplify platform-agnostic GUI development.
Mia on Twitter / X
Fast and Furious ⚡️Update your DeepSeek v4.1 Flash on 3x DGX Sparks- Prose decode tok/s improved by 17%- Code decode tok/s improved by 39%- Mixed decode tok/s improved by 29%Prose single stream is now 50+ tok/s and 85 tok/s at 4 concurrent streams.Changes to 4x DGX… pic.twitter.com/0cMC1kgvDR— Mia (@MiaAI_lab) September 24, 2026
Megakernel: Matching Apple Silicon Efficiency at 2x the Throughput on a RTX 3090
The first megakernel for hybrid DeltaNet/Attention LLMs. 413 tok/s at 1.87 tok/J on a 2020 RTX 3090, matching M5 Max efficiency at 1.8x throughput.

Tolaria
A second brain for the AI era. Free forever, local-first, Markdown-based, Git-ready, and AI-friendly.
The State of On-Device LLMs
Xuan-Son Nguyen, an engineer at Hugging Face, specializes in on-device large language models (LLMs) and runtime optimization, working extensively with llam...
The State of On-Device LLMs
Xuan-Son Nguyen, an engineer at Hugging Face, specializes in on-device large language models (LLMs) and runtime optimization, working extensively with llam...