







A new way to group VMs, with shared resources and predictable pricing. Plus new Personal and Work plans.
An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference Age
LLM inference has become a global-scale, heterogeneous workload spanning agents, retrieval, tool-use, code execution and multi-modal reasoning. These workloads naturally enable context reuse from overlapping inputs, creating a major opportunity to store and reuse the contexts' KV Caches instead of recomputing them. However, model-side advances that shrink the KV Cache and system-side advances that reduce compute, storage, and transfer costs are evolve independently within legacy cloud boundaries. We argue that future inference infrastructure should allow decoupling of compute and KV Cache storage across cloud and datacenters. The network becomes an active distribution channel; bandwidth, latency and pricing directly determines how the KV Cache should be managed. We propose a vision for an Internet for the KV Cache, with KV Cache management working as a content-distribution system. In this view, KV Cache storage and recompute decisions are driven by model, infrastructure, and application metrics, to enable adaptive, content-driven decisions for minimizing latency and cost.

Welcome to Commons Computer
Commons Computer is an experiment in sharing compute resources and admin & management support across multiple people and projects.
Why exe.dev VMs are persistent - exe.dev blog
On the design decision to make VMs persistent, with persistent disks.
Virtual Cards | Old Open Collective Docs
Virtual cards are an additional benefit offered to collectives through some hosts. Hosts create virtual cards and assign them to a Collective. Anyone with access to that card can then use it to make payments on behalf of the Collective. This is particularly useful for covering recurring costs like hosting a website.
Heaps do lie: debugging a memory leak in vLLM. | Mistral AI
The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.
Vercel Sandbox
Vercel Sandbox allows you to run arbitrary code in isolated, ephemeral Linux VMs.
Deploy Hyper-Converged Ceph Cluster
Proxmox VE unifies your compute and storage systems, that is, you can use the same physical nodes within a cluster for both computing (processing VMs and containers) and replicated storage. The traditional silos of compute and storage resources can be wrapped up into a single hyper-converged appliance. Separate storage networks (SANs) and connections via network attached storage (NAS) disappear. With the integration of Ceph, an open source software-defined storage platform, Proxmox VE has the ability to run and manage Ceph storage directly on the hypervisor nodes.
Cloudblast - Pricing
Secure and scalable virtual machine hosting, featuring built-in DDoS protection to ensure continuous, safe operation for your critical applications.

AI Coding & Cloud | Meetup
This group started January 2016.We are about extraordinary advances in software creation, maintenance, and hosting.We are a gathering of good natured people for encouragement and support of what you do.This is your group! How you can help:1. Help to publicize events at work, with friends, and especi

Docker Model Runner Integrates vLLM for High-Throughput Inferencing
How Docker Model Runner integrates vLLM as an inference backend, letting developers run safetensors models with high-throughput serving, PagedAttention, streami
Confidential Inference via Trusted Virtual Machines
Announcing a new collaborative research paper on Confidential Inference, a set of tools to improve the security of our model weights and of our users' data

Commons Computer
A modern team knowledge base for your internal documentation, product specs, support answers, meeting notes, onboarding, & more…