







A Blog post by QVAC on Hugging Face
GPU Poor LLM Arena - a Hugging Face Space by k-mktr
Compact LLM Battle Arena: Frugal AI Face-Off!
Ellora: Enhancing LLMs with LoRA - Standardized Recipes for Capability Enhancement
A Blog post by Asankhaya Sharma on Hugging Face

clem 🤗 on Twitter / X
Introducing Kernels on the Hugging Face Hub ✨What if shipping a GPU kernel was as easy as pushing a model?- Pre-compiled for your exact GPU, PyTorch & OS- Multiple kernel versions coexist in one process- torch.compile compatible- 1.7x–2.5x speedups over PyTorch baselines pic.twitter.com/U0qDdxCWkd— clem 🤗 (@ClementDelangue) April 14, 2026
H100 PCIe vs SXM vs NVL: Which H100 GPU Is Fastest and Most Cost-Effective for Fine-Tuning LLMs?
Benchmarking all three H100 variants for full, LoRA, and QLoRA fine-tuning

Nvidia has been in talks to acquire Hugging Face for more than $13 billion
Nvidia has held talks to acquire Hugging Face for more than $13 billion as the chip giant expands its AI dealmaking.
The State of On-Device LLMs
Xuan-Son Nguyen, an engineer at Hugging Face, specializes in on-device large language models (LLMs) and runtime optimization, working extensively with llam...
The State of On-Device LLMs
Xuan-Son Nguyen, an engineer at Hugging Face, specializes in on-device large language models (LLMs) and runtime optimization, working extensively with llam...
"AI" centralization
So it seems like NVIDIA has agreed to buy Hugging Face. NVIDIA is the company building most of the chips used in “AI” data centers and Hugging face runs probably the biggest repository for data sets and open weight “AI” models in the world (Hugging Face also offers inference but not at relevant rates and […]

Nvidia agrees to buy Hugging Face for $12.9 billion, The Information reports
Nvidia has agreed to buy Hugging Face, a repository of open-source AI models, for $12.9 billion, The Information reported on Wednesday, citing a person with knowledge of the deal.

Vaibhav (VB) Srivastav on Twitter / X
Hello local LLM enjoyers, starting today you should be able to find if you can run a GGUF directly from the Hugging Face page! 🔥Super psyched to see this out in the wild - this has been a big ask from the community!P.S. It'll pick all the hardwares you've defined under your… pic.twitter.com/7PTFCUHftl— Vaibhav (VB) Srivastav (@reach_vb) April 1, 2025
Modern GPU Programming For MLSys — Modern GPU Programming For MLSys
Machine learning systems sit at the heart of modern AI workloads. In these systems, performance often comes down to the quality of a small number of GPU kernels. Attention kernels, LLM prefill and decode kernels, low-precision block-scaled GEMMs, fused MoE layers, and other large fused kernels all directly shape end-to-end speed in both training and serving.
Modern GPU Programming For MLSys — Modern GPU Programming For MLSys
Machine learning systems sit at the heart of modern AI workloads. In these systems, performance often comes down to the quality of a small number of GPU kernels. Attention kernels, LLM prefill and decode kernels, low-precision block-scaled GEMMs, fused MoE layers, and other large fused kernels all directly shape end-to-end speed in both training and serving.
Osaurus — All your AI. One app.
Chat with GPT-5, Claude, Llama, and more — or download local MLX models from Hugging Face. Supercharge Cursor with powerful tools. Open source and built for Mac.
RightNow AI - YC-Backed GPU Research Lab
YC-backed GPU research lab building the RightNow CUDA editor, RunInfra inference infra, Forge kernels, and publishing AutoMegaKernel and related papers on arXiv.

mem-agent: Persistent, Human Readable Memory Agent Trained with Online RL
A Blog post by Dria on Hugging Face