







exact tokenizers ids, as fast as the hardware allows
Actual Computer › Introducing toks
Actual Computer's first public software release: toks, the best tokenizer on earth. The exact same ids as Hugging Face tokenizers, 13x to 151x faster on one CPU core.

The Bitter Lesson is coming for Tokenization
Highlights the desire to replace tokenization with a general method that better leverages compute and data. We'll see tokenization's fragility and review the Byte Latent Transformer arch.

Tokenization Tax 2026: LLM Tokenizer Comparison | ellamind
A measurement study of 20 LLM tokenizers across 12 languages and 21 kinds of text. On German, Claude 4.7+ needs 2.01x the tokens of the OpenAI reference, and Cohere Command A+ the fewest.
Hyperfast AI: Rethinking Design for 1000 tokens/s
I recently spoke at AI Tinkerers Raleigh about hyperfast inference systems and how they’re fundamentally changing AI application design. If you haven’t heard of Cerebras (or however they pronounce it), you’re in for a treat—this is one of the most exciting areas of research in AI right now.

ai/nanoid
A tiny (118 bytes), secure, URL-friendly, unique string ID generator for JavaScript
Fastest 1000000 tokens
Mem0 Research Paper: Token-Efficient Memory Algorithm
Benchmarked across LoCoMo, LongMemEval, and BEAM, achieves competitive accuracy while using under 7,000 tokens per retrieval call. For comparison, full-context approaches on these benchmarks routinely consume 25,000+ tokens per query.

Fast Embeddings on GPUs
Fast and accurate search is vital to all of Perplexity, from Search and Computer to our API Platform. Behind the scenes, the heavy lifting is done by embedding


Port amazing nope-id optimizations by ai · Pull Request #602 · ai/nanoid
A tiny (118 bytes), secure, URL-friendly, unique string ID generator for JavaScript - Port amazing nope-id optimizations by ai · Pull Request #602 · ai/nanoid
QuickDID - AT Protocol Identity Resolution Service
High-performance handle-to-DID resolution service for the AT Protocol ecosystem. Resolve Bluesky and AT Protocol handles instantly.
Prompt caching: 10x cheaper LLM tokens, but how? | ngrok blog
A far more detailed explanation of prompt caching than anyone asked for: how tokens, embeddings, and attention make cached LLM tokens 10x cheaper and faster.

Engineering High-Performance Parsers with Data-Oriented Design
Notes from building Yuku: the AST is flat arrays of u32 indices instead of a pointer tree, and memory layout, allocation, strings, unicode, and serialization all follow from that one decision.
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
Scaling language models unlocks impressive capabilities, but the accompanying computational and memory demands make both training and deployment expensive. Existing efficiency efforts typically target either parameter sharing or adaptive computation, leaving open the question of how to attain both simultaneously. We introduce Mixture-of-Recursions (MoR), a unified framework that combines the two axes of efficiency inside a single Recursive Transformer. MoR reuses a shared stack of layers across recursion steps to achieve parameter efficiency, while lightweight routers enable adaptive token-level thinking by dynamically assigning different recursion depths to individual tokens. This allows MoR to focus quadratic attention computation only among tokens still active at a given recursion depth, further improving memory access efficiency by selectively caching only their key-value pairs. Beyond these core mechanisms, we also propose a KV sharing variant that reuses KV pairs from the first recursion, specifically designed to further decrease memory footprint. Across model scales ranging from 135M to 1.7B parameters, MoR forms a new Pareto frontier: at equal training FLOPs and smaller model sizes, it significantly lowers validation perplexity and improves few-shot accuracy, while delivering higher throughput compared with vanilla and existing recursive baselines. These gains demonstrate that MoR is an effective path towards large-model quality without incurring large-model cost.

ZTA: Zero Token Architecture - Kelsey Hightower | PlatformCon 2026
something that has come up fairly recently with LLMs - for coding, specifically - is that it’s become a lot easier to burn stupefying amounts of tokens on stuff very fast, with agents running 24/7 or managing more agents (see: Yegge’s Gas Town) even with low inference costs that adds up in a hurry
Jesse Felder
‘While some cling to the promise of an AI “revolution,” the cost of adoption is proving a stubborn bottleneck. These developments also suggest that the economics of replacing human labor with AI may be more complicated than some early forecasts originally implied.’ fortune.com/2026/05/22/microsoft-ai-cost-…