Skip to content

superfluid documentation

superfluid is a serving daemon for local LLM inference. It puts one scheduler, one durable session log and one set of APIs (OpenAI, Anthropic, Ollama) in front of several inference runtimes: llama.cpp for GGUF files, MLX for MLX directories, and the baseRT engine for .base bundles.

superfluid serve unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M

Getting started

Serving

Integrations

Models

Features

Deployment

Design

Performance

Contributing