







AI inference at the edge. ggml has 22 repositories available. Follow their code on GitHub.
Supported AI models in GitHub Copilot - GitHub Docs
Learn about the supported AI models in GitHub Copilot.


Open-Source Agentic Inference Benchmark | InferenceX
Compare AgentX, InferenceX's long-context, multi-turn coding scenario, with fixed-sequence AI inference across chips and frameworks. Public NVIDIA and AMD runs update when configurations change.
Open Source AI Inference Benchmark | InferenceX
Compare AI inference performance across GPUs and frameworks. Real benchmarks on NVIDIA GB200, B200, AMD MI355X, and more. Free, open-source, continuously updated.
TheStage AI – Faster, Cheaper AI Inference
Accelerate models on NVIDIA & edge. Full guides for setup, optimization & deploy. ANNA, QLIP, Elastic Models, CLI & API. Built for AI teams & devs.

Overview - GroqDocs
Fast LLM inference, OpenAI-compatible. Simple to integrate, easy to scale. Start building in minutes.

Strata Demo video
Strata Demo video
Glimmer · Reproducible AI science
Glimmer turns a research project into a navigable knowledge graph you can explore, run, verify, and extend — reproducibly.
Run DeepSeek-R1 Dynamic 1.58-bit
DeepSeek R-1 is the most powerful open-source reasoning model that performs on par with OpenAI's o1 model. Run the 1.58-bit Dynamic GGUF version by Unsloth.

SambaNova | The Fastest AI Inference Platform
Discover SambaNova - the complete AI platform delivering the fastest AI inference, fine-tuning, and scalable solutions for agentic AI easily integrated into existing data center infrastructures.

Cloudflare Workers AI | Open-source AI inference
Workers AI facilitates the scalable development & deployment of AI applications at the edge.

GGML and llama.cpp join HF to ensure the long-term progress of Local AI
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
