







An extremely fast WASM Apple ProRes video decoder
Optimizing On-Device Inference for Apple Silicon
A custom local engine that improves prefill and decode throughput

Combining NVIDIA DGX Spark + Apple Mac Studio for 4x Faster LLM Inference with EXO 1.0
Disaggregating Prefill and Decode: Faster First Tokens, Faster Streams

Combining NVIDIA DGX Spark + Apple Mac Studio for 4x Faster LLM Inference with EXO 1.0
Disaggregating Prefill and Decode: Faster First Tokens, Faster Streams

Why MLX ā Prince Canuma, Neywa Labs
NixOS Linux 26.05 Yarara 1440p 60fps record + stream testing #foss #gnu #linux #nixos #kde #plasma
clandestine.eth š¦š on Twitter / X
Heterogeneous acceleration on Apple Silicon achieved.ANE + GPU running in parallel.Mirror SD with DFlash, ported to MLX ā targeting ANE + GPU simultaneously.The M-series was designed for this. We just hadn't unlocked it yet. pic.twitter.com/raSH0CMN4Vā clandestine.eth š¦š (@0xClandestine) April 15, 2026

Quantizer support in Mediabunny v1.52.0 | Mediabunny
Mediabunny v1.52.0 adds quantizer-based video encoding for AVC, HEVC, VP9, and AV1, enabling constant-quality video encoding.

ICML WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
Real-time Automatic Speech Recognition (ASR) is a fundamental building block for many commercial applications of ML, including live captioning, dictation, meeting transcriptions, and medical scribes. Accuracy and latency are the most important factors when companies select a system to deploy. We present WhisperKit, an optimized on-device inference system for real-time ASR that significantly outperforms leading cloud-based systems. We benchmark against server-side systems that deploy a diverse set of models, including a frontier model (OpenAI gpt-4o-transcribe), a proprietary model (Deepgram nova-3), and an open-source model (Fireworks large-v3-turbo).Our results show that WhisperKit matches the lowest latency at 0.46s while achieving the highest accuracy 2.2\% WER. The optimizations behind the WhisperKit system are described in detail in this paper.

Megakernel: Matching Apple Silicon Efficiency at 2x the Throughput on a RTX 3090
The first megakernel for hybrid DeltaNet/Attention LLMs. 413 tok/s at 1.87 tok/J on a 2020 RTX 3090, matching M5 Max efficiency at 1.8x throughput.

Apple Vision Pro
Featuring the new powerful M5 chip and comfortable Dual Knit Band, Apple Vision Pro seamlessly blends digital content with your physical space.

ARahim3/mlx-tune
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning ā natively on MLX. Unsloth-compatible API.
iPhone Air has an A19 Pro chip, which is has native matmul full NVIDIA inference speed in laptops, maybe even phones too? news.ycombinator.com/item?id=45186015