







Autonomous AI agents that profile, diagnose, and optimize GPU inference across your entire stack — from kernels to models to production pipelines.
Open-Source Agentic Inference Benchmark | InferenceX
Compare AgentX, InferenceX's long-context, multi-turn coding scenario, with fixed-sequence AI inference across chips and frameworks. Public NVIDIA and AMD runs update when configurations change.

Open Source AI Inference Benchmark | InferenceX
Compare AI inference performance across GPUs and frameworks. Real benchmarks on NVIDIA GB200, B200, AMD MI355X, and more. Free, open-source, continuously updated.
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
High Performance AI Lab
High Performance AI Lab builds open inference systems and publishes the conditions behind every number — device, model, quant, and rep count.

Together AI | The AI Native Cloud
Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research.

TheStage AI – Faster, Cheaper AI Inference
Accelerate models on NVIDIA & edge. Full guides for setup, optimization & deploy. ANNA, QLIP, Elastic Models, CLI & API. Built for AI teams & devs.

The Universal Execution Layer for AI
Optimize any AI model on any engine, across all hardware. Dria’s topology-aware compiler and peer-to-peer runtime merge CPUs, GPUs, NPUs & chiplets into one fabric—maximising utilisation, cutting inference cost and ending vendor lock-in.

SambaNova | The Fastest AI Inference Platform
Discover SambaNova - the complete AI platform delivering the fastest AI inference, fine-tuning, and scalable solutions for agentic AI easily integrated into existing data center infrastructures.

Advancing AI Infrastructure for Agentic AI with NVIDIA DOCA In-Silicon Security | NVIDIA Technical Blog
The AI era is driving a new class of infrastructure: AI factories that transform data into intelligence for autonomous AI agents operating at unprecedented scale. Powered by accelerated computing…

Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye
Differential acceleration of cyber, math, and AI

Fastino trains AI models on cheap gaming GPUs and just raised $17.5M led by Khosla | TechCrunch
Tech giants like to boast about trillion-parameter AI models that require massive and expensive GPU clusters. But Fastino is taking a different approach.

ZETIC | On-Device AI for Everything - for any model, on any device, in any framework
Built by ex-Qualcomm AI Engineer. Automate on-device AI deployment with full NPU optimization. Benchmark on 100+ physical devices and ship in hours with just 3 lines of code.

Introducing Pipette: A benchmarking suite for on-device intelligence — Blog
Meet Pipette, an open-source platform for reproducible on-device AI benchmarks across models, quantization, runtimes and hardware.
Import AI 464: Fables writes GPU kernels; AI automation; and analog computation
Is this the beginning of a new world?
