Skip to content

superfluid documentation

One serving daemon for local LLM inference — across runtimes and machines.

Coding agents and any OpenAI, Anthropic or Ollama client reach one superfluid server, which drives baseRT, llama.cpp and MLX workers on this machine and on other machines that join it

superfluid is a serving daemon for local LLM inference. It puts one scheduler, one durable session log and one set of APIs (OpenAI, Anthropic, Ollama) in front of several inference runtimes: llama.cpp for GGUF files, MLX for MLX directories, and the baseRT engine for .base bundles.

superfluid serve unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M

Getting started

Serving

Integrations

Models

Features

Deployment

Design

Performance

Contributing