







Compact LLM Battle Arena: Frugal AI Face-Off!
**An Edge-First Generalized LLM LoRA Fine-Tuning Framework for Heterogeneous GPUs**
A Blog post by QVAC on Hugging Face
Mesh LLM: distributed AI computing on iroh
How Mesh LLM pools existing GPU resources across machines into a single OpenAI-compatible API, built on iroh.
"AI" centralization
So it seems like NVIDIA has agreed to buy Hugging Face. NVIDIA is the company building most of the chips used in “AI” data centers and Hugging face runs probably the biggest repository for data sets and open weight “AI” models in the world (Hugging Face also offers inference but not at relevant rates and […]

ChatGPT Images 2.0 Tops Arena With Big Jump Over Nano Banana 2
OpenAI’s models had created plenty of buzz on Arena when they were released under their codenames, and the official model has done pretty...

Arena AI: The Official AI Ranking & LLM Leaderboard
Chat, compare, vote for the world's best AI models. Join the community shaping the public leaderboard for LLMs, image, and code models through real-world evaluation.

Arena AI: The Official AI Ranking & LLM Leaderboard
Chat, compare, vote for the world's best AI models. Join the community shaping the public leaderboard for LLMs, image, and code models through real-world evaluation.

The State of On-Device LLMs
Xuan-Son Nguyen, an engineer at Hugging Face, specializes in on-device large language models (LLMs) and runtime optimization, working extensively with llam...
The State of On-Device LLMs
Xuan-Son Nguyen, an engineer at Hugging Face, specializes in on-device large language models (LLMs) and runtime optimization, working extensively with llam...
Exploring LLMs with MLX and the Neural Accelerators in the M5 GPU
Mac with Apple silicon is increasingly popular among AI developers and researchers interested in using their Mac to experiment with the…

Nvidia has been in talks to acquire Hugging Face for more than $13 billion
Nvidia has held talks to acquire Hugging Face for more than $13 billion as the chip giant expands its AI dealmaking.
Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye
Differential acceleration of cyber, math, and AI

clem 🤗 on Twitter / X
Introducing Kernels on the Hugging Face Hub ✨What if shipping a GPU kernel was as easy as pushing a model?- Pre-compiled for your exact GPU, PyTorch & OS- Multiple kernel versions coexist in one process- torch.compile compatible- 1.7x–2.5x speedups over PyTorch baselines pic.twitter.com/U0qDdxCWkd— clem 🤗 (@ClementDelangue) April 14, 2026
Arena Leaderboard - a Hugging Face Space by lmarena-ai
This app displays the LMArena leaderboard in a full‑screen view, letting you see the latest rankings of language models at a glance. Just open the page and the leaderboard loads automatically—no in...
Modern GPU Programming For MLSys — Modern GPU Programming For MLSys
Machine learning systems sit at the heart of modern AI workloads. In these systems, performance often comes down to the quality of a small number of GPU kernels. Attention kernels, LLM prefill and decode kernels, low-precision block-scaled GEMMs, fused MoE layers, and other large fused kernels all directly shape end-to-end speed in both training and serving.
Modern GPU Programming For MLSys — Modern GPU Programming For MLSys
Machine learning systems sit at the heart of modern AI workloads. In these systems, performance often comes down to the quality of a small number of GPU kernels. Attention kernels, LLM prefill and decode kernels, low-precision block-scaled GEMMs, fused MoE layers, and other large fused kernels all directly shape end-to-end speed in both training and serving.
OpenAI Chat GPT OSS 20b Open Source LLM Full Local Ai Review