Skip to content

Command-line reference

superfluid is one binary with a few subcommands. Environment variables and the on-disk layout are in Configuration.

command what it does
superfluid serve <model> run the server
superfluid launch <agent> run a coding agent against a local model, starting a server if none answers
superfluid stop stop the server launch started
superfluid runtimes [<model>] list runtimes, or say which one serves a model
superfluid runtime <subcommand> install, update, repair, remove or pin a runtime
superfluid session <subcommand> inspect or export a session from a running server
superfluid media <subcommand> store media in, or sweep, a running server's media pool
superfluid top <host:port>... live monitor over one or more servers
superfluid --help                 # command overview (also: superfluid, superfluid help)
superfluid serve --help           # every serve flag, grouped, with its default
superfluid help runtime           # one command's help
superfluid --version              # also -V, superfluid version

Flags take their value as --flag value or --flag=value. The last occurrence of a scalar flag wins; repeatable flags accumulate.

Exit codes

code meaning
0 normal exit (closed socket, quit monitor, finished drain)
1 run-time failure: a port or socket that cannot be bound, a sessions directory another server holds, a damaged session log, a worker that did not start, a model that did not load, a failed install or pull
2 usage or configuration refusal, before anything loads. An unknown flag names the closest valid one; a bad value names the flag and what it expected
3 (superfluid-workerd, superfluid-noded) the model did not load

superfluid serve

superfluid serve <model> [--model <model> ...] [flags]

A model is a path or a Hugging Face-style id org/model[:tag]. Everything else has a default: the session log goes to ~/.superfluid/sessions, the session socket to ~/.superfluid/superfluid.sock, and the HTTP API listens on 127.0.0.1:8453. A runtime that was never installed is installed on first use.

superfluid serve unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M        # llama.cpp
superfluid serve mlx-community/Qwen3-4B-4bit           # MLX (Apple silicon)
superfluid serve ./Qwen3-4B-Q4_K_M.gguf                 # runtime picked by format
superfluid serve ./model.gguf --port 9000 --api-key "$KEY" --max-batch 4

Model

flag default description
--model <model> required A model path or id to serve; repeatable. A bare first argument is the same. The first model is the default for requests that name none.
--runtime <id>[@<install>] inferred Runtime for every model: basert, llamacpp, mlx (aliases native, llama.cpp, llama-cpp); auto clears it. Repeatable.
--runtime <model>=<id>[@<install>] Runtime for one model, named by path, file stem or id.
--offline off Fetch nothing: look for an id only where its runtime keeps models, and install no runtime.
--pull-<name> [value] Pass --<name> to the runtime's pull tool (for example --pull-file model.gguf). See Supported models.
--model-dir <dir> none Every model under <dir> becomes known and loads on first request. See Serving multiple models.
--idle-timeout <sec> 0 (never) Unload a model unused this long, if it can be loaded again; never the default model.
--basert-lib <path> search libbaseRT file or directory for the basert runtime (sets BASERT_LIB).

With no --runtime, a model id goes to llamacpp when its repo name contains gguf, to mlx when it is mlx-community/... or its repo name contains mlx, and to basert otherwise. A path goes to the runtime that reads its format.

Server

flag default description
--http <addr:port> 127.0.0.1:8453 Address of the OpenAI, Anthropic and Ollama HTTP API.
--no-http Serve the session API on the socket only.
--host <addr> / --port <N> 127.0.0.1 / 8453 The HTTP address as two flags; --http wins.
--socket <path> $SUPERFLUID_HOME/superfluid.sock Native session API socket. Falls back to $TMPDIR/superfluid-<uid>.sock when the default path is too long for a unix socket.
--sessions <dir> $SUPERFLUID_HOME/sessions Session store: log, parks, media, files, batches. One server per directory.
--api-key <key> none Require Authorization: Bearer <key> or X-Api-Key on every route but /health.
--key-policy <file.json> none Per-key class, rate, concurrency and admin rights; turns authentication on. See Security.
--rate-limit <rpm> 0 (off) Requests per minute per client IP.
--drain-timeout <sec> 60 On SIGUSR1, how long in-flight requests get before exit.
--nonstream-keepalive <sec> 0 (off) Write a space every <sec> into a pending non-streaming response.
--web <127.0.0.1:port> off Loopback browser transport for the session API, with a minted token.
--web-origin <url> none An allowed Origin for --web; repeatable.

Generation

flag default description
--max-context <N> auto Context window per lane, floor 512. auto sizes it for the device: the memory left after the weights and a reserve, 60% of it for KV, divided over the lanes, at least 4096; one conversation may grow past its lane's share into what the others leave, up to the model's trained window. llamacpp sizes on Metal and on CUDA with a dedicated GPU, mlx on Metal; elsewhere, for models on two runtimes, and with --model-dir the window is 8192 and startup says why.
--max-batch <N> 8 Concurrent decode lanes. Alias --max-batch-size.
--max-tokens <N> fill the context Generation cap for requests that set none.
--kv-bits <N> 0 (auto) KV cache type: 4 (q4_0), 8 (q8_0), 16 (f16) or 84 (K q8_0 / V q4_0). basert only.
--dialect <name> auto Chat codec: auto, chatml, template, atem or raw. See Chat templates.
--tool-call-parser <name> auto Override the tool-call format: json, atem, gemma, harmony or glm; must agree with the template.
--temperature <F> model default Default temperature, 0-2.
--top-p <F> model default Default nucleus truncation, 0-1.
--top-k <N> model default Default top-k truncation.
--min-p <F> model default Default min-p truncation, 0-1.
--repeat-penalty <F> model default, else 1.0 Default repetition penalty, 0-2; 1.0 disables.

Sampling flags apply only to requests that name no value of their own. See Sampling.

Sessions and caching

flag default description
--park / --no-park off Seal a finishing session's KV to disk and resume from it later.
--park-lossy off Park in the lossy Q8 tier (about half the size). Implies --park and --kv-bits 16.
--park-budget-gb <N> 20 Size bound per park directory, oldest evicted first; 0 disables the sweep.
--sessions-cache off Start a fresh session log on every start (the old one kept as wal.log.prev). Not with --park.
--tool-lease-ms <N> 0 (none) Deadline for every tool call; a later result is recorded as expired.
--prefix-blob-budget-mb <N> engine default Recurrent-state snapshot cache on hybrid models (basert).

See Sessions.

Files API

flag default description
--files-max-bytes <N> unbounded Refuse a /v1/files upload past this many stored bytes.
--files-expiry <sec> never Delete a file this long after upload.
--files-sweep <sec> 300 How often the expiry sweep runs.

Scheduling and QoS

flag default description
--tick-decode-budget <N> 0 (--max-batch × grant) Decode tokens per tick across all lanes.
--tick-target-ms <N> 2000 Tick duration while every lane is busy (250 max on llamacpp, mlx).
--prefill-budget <N> 4096 Prompt tokens per tick, 64-4096; 0 is adaptive.
--starvation-ticks <N> 8 Ticks before a starved lane is served first; also the queue aging period.
--class-lanes <class>=N[,...] uncapped Lane cap per QoS class, e.g. background=2,agent=6.
--http-default-qos <class> agent Class for HTTP requests that name none.
--no-http-qos-header honoured Ignore the client's x-superfluid-qos header.
--http-allow-batch-invariant off Allow x-superfluid-batch-invariant requests.
--worker-process / --no-worker-process on Run each model in its own worker process. Off only for a single model on a linked runtime.
--os-pressure / --no-os-pressure on React to OS memory pressure.
--pressure-high <pct> 85 KV pool occupancy that triggers cache eviction.
--pressure-low <pct> 70 Occupancy eviction brings it down to.
--pin-budget-pct <pct> 50 Share of the pool pinned prefixes may hold; 0 records pins without enforcing.

See Scheduling.

Speculative decoding

flag default description
--speculate <directive> off auto, mtp-head, prompt-lookup, dflash:<path>, dspark:<path>, eagle3:<path>, draft-model:<path>, a drafter .base path, or off. On llamacpp and mlx, only prompt-lookup (and auto).
--spec-draft-tokens <N> per strategy Draft depth, 1-15.
--spec-adaptive / --no-spec-adaptive on Adapt depth per lane count.
--spec-min-yield <F> 0.75 Accepted tokens per round below which a request stops drafting; 0 disables.
--spec-yield-rounds <N> 24 Rounds observed before the yield floor is judged.
--spec-throughput-gate / --no-spec-throughput-gate on Turn speculation off at lane counts where it loses.
--spec-gate-probe-tokens <N> 16 Length of a plain-decode probe.
--spec-min-speedup <F> 1.08 Speed-up the gate requires.
--spec-gate-reprobe <N> 32 Requests after an abandonment before re-measuring; 0 makes it final.
--spec-gate-reprobe-max <N> 1024 Cap on the doubling re-probe window.
--spec-max-temperature <F> none Cap a temperature that comes from the model's default while speculating.
--dspark-confidence <F> engine default DSpark confidence threshold.
--spec-bitexact / --no-spec-bitexact off Make speculative and plain decoding bit-exact (slower).

See Speculative decoding.

Code completion (FIM)

flag default description
--fim-model <path> none A dedicated fill-in-the-middle model in its own worker.
--fim-max-context <N> min(window, 8192) The FIM model's window.
--fim-max-batch <N> 2 The FIM model's lanes.
--completion-deadline-ms <N> 0 (none) A completion that cannot start in time expires. Alias --request-timeout.
--completion-rate <F> 10 Completions per second per client; 0 is unmetered.
--completion-burst <N> 20 Burst a client may send back to back.

Logging and telemetry

flag default description
--log-filter <directive> info EnvFilter syntax, e.g. superfluid_daemon::scheduler=debug,info.
-v, --verbose Same as --log-filter debug.
--log-dir <dir> stderr JSON log files superfluid.<n>.jsonl (64 MiB each, 8 kept).
--log-file <path> none Redirect stderr to this file (append).
--tui / --no-tui on in a terminal Live monitor in the terminal.
--otlp-metrics <url> off Push metrics to <url>/v1/metrics (http:// only).
--otlp-interval-ms <N> 15000 Metrics push interval.
--otlp-endpoint <url> off Push operational spans to <url>/v1/traces (http:// only).
--otlp-header <k=v> none Header on every export; repeatable.
--otlp-service-name <s> superfluid service.name resource attribute.
--otlp-filter <directive> info,superfluid_daemon::scheduler=debug Which spans export.
--otlp-queue <N> 8192 Span queue; full drops spans.
--otlp-batch-ms <N> 1000 Span batch interval.
--otlp-timeout-ms <N> 5000 Export request timeout.

See Observability.

Fleet head

flag default description
--fleet <host:port>[,...] off Serve as a fleet head over these superfluid-noded nodes.
--fleet-auth <path> none Shared credential sent to every node.
--fleet-policy <policy> load-aware Placement: load-aware, least-loaded or round-robin.
--fleet-pool-high <pct> 90 KV occupancy at which load-aware routes around a node.
--fleet-conns-per-node <N> 1 Connections, and so concurrent generations, per node.

See Distributed serving.

Retired flags

Flags from earlier releases are accepted with a notice naming what replaced them: --paged-kv, --prefix-cache, --continuous-batching [N], --prefix-cache-file, --prefix-cache-save-interval, --files-dir, --media-dir, --metallib, --prefill-chunk, --gpu-wait-timeout-ms, --decode-replay, --paged-weights, --no-paged-weights-retry, --no-baked-decode.

superfluid launch

superfluid launch                                   # list the agents and whether each is installed
superfluid launch <agent> [--model <m>] [--http <addr:port>] [--api-key <key>] [--print] [-- <agent args>...]

Runs claude, codex, pi, opencode, hermes, cline, openclaw or the ollama CLI pointed at the server on --http, starting superfluid serve <model> in the background when nothing answers there. --print shows what would run; arguments after -- go to the agent. See superfluid launch.

superfluid stop

Sends the server launch started SIGTERM and waits for it to drain. A server started with superfluid serve is not touched.

superfluid runtimes

superfluid runtimes [<model>] [--json]

Without a model: one row per runtime and install, with state (ready, not installed, broken, missing), version, formats, device and worker location, then each problem with its fix. With a model: which runtime reads it and what that runtime serves and refuses; exits 1 when none here serves it.

superfluid runtime

superfluid runtime install <id> [--version V] [--backend B] [--from <url|path>] [--sha256 H] [--untested] [--dry-run]
superfluid runtime update <id>
superfluid runtime repair <id> [<install>]
superfluid runtime remove <id> [<install>]
superfluid runtime use <id> <install> | --auto

See Runtimes.

superfluid session

superfluid session inspect <id> [--socket <path>] [--json]
superfluid session export <id> [--socket <path>] --format jsonl|otlp-jsonl [--out <dir>] \
    [--include-content] [--model <name>] [--endpoint <otlp-http-url> [--allow-content-egress]]

inspect prints a session's summary and events. export writes its events as JSONL or OTLP trace lines, or pushes them to a collector. Content is omitted unless --include-content, and leaves the host only with --allow-content-egress. --socket defaults to $SUPERFLUID_HOME/superfluid.sock. See Sessions.

superfluid media

superfluid media put <file> [--socket <path>]    # store a file, print its hash
superfluid media gc [--socket <path>]            # sweep blobs no session references

superfluid top

superfluid top [--api-key <key>] <host:port> [<host:port> ...]

Live per-node view over each server's /metrics. A keyed node needs --api-key, else it shows as 401. q or Esc quits.

superfluid-workerd and superfluid-noded

superfluid-workerd is the worker the daemon spawns per model; you do not run it to serve. superfluid-workerd check --engine <id> prints the runtime report superfluid runtimes reads. Installed runtimes ship their own workers (superfluid-worker-llamacpp, -mlx, -basert) with the same interface; see Writing a runtime adapter.

superfluid-noded is the fleet node agent; its flags are in Distributed serving.