







Highlights Support for mxfp4 quantization (Metal, CPU) More performance improvements, bug fixes, features in CUDA backend mx.distributed supports NCCL back-end for CUDA What's Changed [CUDA]...
Cheng on Twitter / X
We have been expecting this since ollama's first pull request to MLX. It is just the beginning, CUDA & CPU backends are still improving and hopefully we will have one framework unifying inference & training for all platforms. https://t.co/EaBmEaNJhZ— Cheng (@zcbenz) March 31, 2026
NVlabs/cuda-oxide
cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign language bindings, just Rust.
Ash Hart on Twitter / X
MCDMA | Metal CUDA Direct Memory Access 🚀If you have a Spark and an Apple Silicon Mac, MCDMA gives you a direct RDMA path between CUDA memory and Metal-side unified memory over USB-C.Registered memory, rkeys, one-sided READ/WRITE, two-sided SEND/RECV with credit flow… pic.twitter.com/gFt8tx9xII— Ash Hart (@ashxhart) August 18, 2026

LM Studio on Twitter / X
MXFP4 support for openai/gpt-oss landed in MLX!⌘ + Shift + R to update your runtime.🚀 https://t.co/hnM3e0L2qw pic.twitter.com/JfGgD57duJ— LM Studio (@lmstudio) August 29, 2025

RightNow AI - YC-Backed GPU Research Lab
YC-backed GPU research lab building the RightNow CUDA editor, RunInfra inference infra, Forge kernels, and publishing AutoMegaKernel and related papers on arXiv.

Ivan Fioravanti ᯅ on Twitter / X
"We are releasing Open Source implementations for CoreAILanguageModel and MLXLanguageModel for running a myriad of local models on the Apple Neural Engine or your Mac's GPU" 👀 From #WWDC26: What’s new in the Foundation Models framework video: https://t.co/1NtKWYhNRs pic.twitter.com/HVtBsr3tjL— Ivan Fioravanti ᯅ (@ivanfioravanti) June 9, 2026
Zhijian Liu on Twitter / X
🔥 DFlash x MLX is happening!Shoutout to @aryagm01 for the early work on this. We're building on the momentum. Native MLX support, more models (Qwen3.5), up to 4x faster. Lossless!👉 https://t.co/9CtLKDptNI pic.twitter.com/pMdaCH4oKi— Zhijian Liu (@zhijianliu_) April 15, 2026
Cua on Twitter / X
1/ Today, as part of our broader research into Apple Silicon virtualization, we're releasing a process-scoped Metal capability layer for macOS VMs. On one M1 Ultra, prompt / generation:TinyLlama: 11.08× / 16.36×Gemma 4 12B: 7.20× / 14.54×Muse Glimmer 30B: 7.55× / 8.87× pic.twitter.com/6UwxgoDowT— Cua (@trycua) August 11, 2026

Ollama is now powered by MLX on Apple Silicon in preview· Ollama Blog
Today, we're previewing the fastest way to run Ollama on Apple silicon, powered by MLX, Apple's machine learning framework.

mlx-examples/stable_diffusion at main · ml-explore/mlx-examples
Examples in the MLX framework. Contribute to ml-explore/mlx-examples development by creating an account on GitHub.
What’s MXFP4? The 4-Bit Secret Powering OpenAI’s GPT‑OSS Models on Modest Hardware
A Blog post by Rakshit Aralimatti on Hugging Face
AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing?
$3 Million USD dataset open sourced, 1 Mil+ Context Length, Multiturn, Sub Agents 95%+ KVCache HitRate, GB300 NVL72, MI355, B200

Cua: Scale computer fleets for computer-use agents
Scale Linux, Windows, macOS, and Android computer fleets for computer-use agents with one open-source MCP/CLI driver.

LLMs Can Now Write GPU Kernels That Beat torch.compile - Break AI Scaling Limits in 7 Days
We're now seeing multi-agent systems that take your PyTorch code and produce CUDA or Triton kernels with 2x to 14x speedups over torch.compile(mode='max-autotune-no-cudagraphs'). Not on toy benchmarks. On real models like Llama-3.1-8B, Whisper, and Stable Diffusion. Learn proven techniques to shift the scaling law intercept and achieve 10-50% performance gains.
