Distributed serving (fleet mode)¶
Fleet mode runs one superfluid as a head that owns the HTTP API and the chat template, and one or more node agents (superfluid-noded) that each own a model and sample its tokens. Each request is placed as a whole session on one node; only sampled tokens cross the network.
Warning
Fleet mode is early. The head keeps no durable session log, serves two routes, and the link between head and nodes is unencrypted TCP.
Run a node¶
superfluid-noded --engine native --model /models/qwen.base \
--listen 100.64.0.2:8454 --max-context 32768 --max-batch 8 \
--identity box-a --auth-file /etc/superfluid/fleet.key
| flag | default | meaning |
|---|---|---|
--listen <ip:port> |
0.0.0.0:8454 |
where the head connects |
--engine <id> |
mock |
native (basert), llamacpp or mlx; must be linked into this build |
--model <path> |
required (except mock) |
the model |
--model-name <id> |
the model's id | the name the head matches on; set it the same on every node |
--max-context <N> |
4096 | the node's window, floor 512 |
--max-batch <N> |
8 | the node's lanes |
--identity <name> |
superfluid-node |
the name in the head's logs |
--auth-file <path> |
none | shared credential the head must present |
--insecure-no-auth |
off | accept any peer on a non-loopback address |
--handshake-timeout-ms <N> |
10000 | time a connecting head has to say hello |
--idle-timeout-ms <N> |
120000 | drop an idle connection |
A node bound off loopback refuses to start without --auth-file unless --insecure-no-auth is given. A node has no sampling flags and no durable state: its session store is a scratch directory under the system temp dir. It logs to stderr (RUST_LOG sets the filter), including each head it refuses for a key that does not match.
Run the head¶
superfluid serve /models/qwen.base --http 0.0.0.0:8453 \
--fleet 100.64.0.2:8454,100.64.0.3:8454 \
--fleet-auth /etc/superfluid/fleet.key
At start the head connects to every node and reports each one, or why it cannot reach it:
superfluid: fleet node 'box-a' at 100.64.0.2:8454: qwen, context window 32768 tokens, 8 lanes
superfluid: fleet node 100.64.0.3:8454: closed the connection at the handshake; check that --fleet-auth holds the key in the node's --auth-file; retrying it every 10 s
superfluid: context window 32768 tokens (the largest node's)
| flag | default | meaning |
|---|---|---|
--fleet <host:port>[,...] |
the nodes | |
--fleet-auth <path> |
none | credential sent to every node |
--fleet-policy |
load-aware |
load-aware, least-loaded or round-robin |
--fleet-pool-high <pct> |
90 | load-aware avoids nodes at or above this KV occupancy |
--fleet-conns-per-node <N> |
1 | connections per node, and so concurrent generations there |
The head loads no engine but still reads the model for its tokenizer and chat template; its model id must match what the nodes advertise (the head warns about a node that serves another). Its context window is the nodes', so it needs no --max-context. It logs where each session is placed and every move to another node. --dialect, --tool-call-parser and --temperature/--top-p/--top-k/--min-p apply on the head. --api-key, --key-policy and --rate-limit protect its listener; a policy's class and max_class do not apply, since the head has no scheduler.
Refused on the head (set them on each node's own superfluid serve): --class-lanes, --http-default-qos, --no-http-qos-header, --http-allow-batch-invariant, --spec-max-temperature, --fim-model, --repeat-penalty, --otlp-endpoint.
Placement policies¶
| policy | how it picks |
|---|---|
load-aware (default) |
prefers nodes whose window fits the prompt, skips nodes above --fleet-pool-high, then the lowest estimated completion time from each node's reported rates. A node measures its rates only while it is generating, so an idle node keeps the speed it showed when busy; one that has not generated yet counts at the mean of the others |
least-loaded |
fewest sessions in progress plus requests in flight on that node, whatever its speed |
round-robin |
rotate through capable nodes; use it to prove a fleet is distributing |
A prompt no node can hold is refused with context_length_exceeded.
What the head serves¶
POST /v1/chat/completions, streaming or not, with tools (OpenAI API).GET /v1/models: the one model the fleet serves.
Refused with 400: non-zero presence_penalty/frequency_penalty, a repeat_penalty other than 1.0, ignore_eos, stream_options.continuous_usage_stats, and image or audio parts. Not read: logprobs, logit_bias, response_format, tool_choice. There is no /health, /metrics, Anthropic, Ollama, embeddings, files or batches surface on the head. A generation has a ten-minute wall-clock ceiling.
Failure handling¶
| event | what happens |
|---|---|
| node fails mid-stream | delivered tokens become context; the session is re-placed on the next capable node, which generates the rest. No token is delivered twice |
| node unreachable | skipped; the head tries it again every 10 s and uses it once it answers |
| node restarts | comes back empty; the head reconnects and new placements use it |
| connection idle | the head sends a keepalive every 15 s, under the node's --idle-timeout-ms; a connection a node closed anyway is replaced before it is used |
| head restarts | in-flight requests fail; placement state is lost |
| head misses a lease renewal | the node fences the session and emits nothing more |
Limits¶
- One model per fleet.
- No KV transfer between nodes: a failover re-prefills.
--fleet-conns-per-nodeis the only concurrency lever on a node, whatever its lane count.- No QoS class crosses the wire.