







How LLMs Actually Work
A from-the-ground-up walkthrough of how modern LLMs work, from tokens to transformer blocks to the next-token loop
Understanding Multimodal LLMs
An introduction to the main techniques and latest models

Can LLMs Be Computers? | Percepta
We build a computer inside a transformer — executing arbitrary C programs for millions of steps with exponentially faster inference via 2D attention heads.

Can LLMs Be Computers? | Percepta
We build a computer inside a transformer — executing arbitrary C programs for millions of steps with exponentially faster inference via 2D attention heads.

The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities (Version 1.0)
LLM Compressor
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
The State of On-Device LLMs
Xuan-Son Nguyen, an engineer at Hugging Face, specializes in on-device large language models (LLMs) and runtime optimization, working extensively with llam...
The State of On-Device LLMs
Xuan-Son Nguyen, an engineer at Hugging Face, specializes in on-device large language models (LLMs) and runtime optimization, working extensively with llam...
The Big LLM Architecture Comparison
From DeepSeek-V3 to Kimi K2: A Look At Modern LLM Architecture Design

Your Laptop Isn’t Ready for LLMs. That’s About to Change
The quest to run large AI models locally on an individual's machine are driving the biggest change in laptop architecture in decades.

Francois Chaubard on Twitter / X
I believe that LSTM based architectures (constant hidden state, constant flops / step) will crush transformers as we know them in 2-3 years. we will laugh about KV Caches and how dumb we were.. to achieve this, the LSTM will need to have a large external memory bank with…— Francois Chaubard (@FrancoisChauba1) July 30, 2026
Towards Feasible, Private, Distributed LLM Inference
Exploring how the Secure Transformer Inference Protocol (STIP) protects inputs, outputs, and model weights with lightweight permutations enabling efficient, privacy-safe LLM inference at scale.
Introducing SubQ: The First Fully Subquadratic LLM
Subquadratic is a frontier AI research and infrastructure company building a new class of LLMs.

Introducing SubQ 1.1 Small
Subquadratic is a frontier AI research and infrastructure company building a new class of LLMs.

Sakana AI on Twitter / X
We’re excited to introduce Doc-to-LoRA and Text-to-LoRA, two related research exploring how to make LLM customization faster and more accessible.https://t.co/wGKDNhBcJXBy training a Hypernetwork to generate LoRA adapters on the fly, these methods allow models to instantly… pic.twitter.com/gId3J6hgEr— Sakana AI (@SakanaAILabs) February 27, 2026