Skip to content

Distributed serving (fleet mode)

Fleet mode runs one superfluid as a head that owns the HTTP API and the chat template, and one or more node agents (superfluid-noded) that each own a model and sample its tokens. Each request is placed as a whole session on one node; only sampled tokens cross the network.

Warning

Fleet mode is early. The head keeps no durable session log, serves two routes, and the link between head and nodes is unencrypted TCP.

fleet mode: the head serves the HTTP API and places each session on a node over Link F; nodes run the worker and sample

Run a node

superfluid-noded --engine native --model /models/qwen.base \
    --listen 100.64.0.2:8454 --max-context 32768 --max-batch 8 \
    --identity box-a --auth-file /etc/superfluid/fleet.key
flag default meaning
--listen <ip:port> 0.0.0.0:8454 where the head connects
--engine <id> mock native (basert), llamacpp or mlx; must be linked into this build
--model <path> required (except mock) the model
--model-name <id> the model's id the name the head matches on; set it the same on every node
--max-context <N> 4096 the node's window, floor 512
--max-batch <N> 8 the node's lanes
--identity <name> superfluid-node the name in the head's logs
--auth-file <path> none shared credential the head must present
--insecure-no-auth off accept any peer on a non-loopback address
--handshake-timeout-ms <N> 10000 time a connecting head has to say hello
--idle-timeout-ms <N> 120000 drop an idle connection

A node bound off loopback refuses to start without --auth-file unless --insecure-no-auth is given. A node has no sampling flags and no durable state: its session store is a scratch directory under the system temp dir. It logs to stderr (RUST_LOG sets the filter), including each head it refuses for a key that does not match.

Run the head

superfluid serve /models/qwen.base --http 0.0.0.0:8453 \
    --fleet 100.64.0.2:8454,100.64.0.3:8454 \
    --fleet-auth /etc/superfluid/fleet.key

At start the head connects to every node and reports each one, or why it cannot reach it:

superfluid: fleet node 'box-a' at 100.64.0.2:8454: qwen, context window 32768 tokens, 8 lanes
superfluid: fleet node 100.64.0.3:8454: closed the connection at the handshake; check that --fleet-auth holds the key in the node's --auth-file; retrying it every 10 s
superfluid: context window 32768 tokens (the largest node's)
flag default meaning
--fleet <host:port>[,...] the nodes
--fleet-auth <path> none credential sent to every node
--fleet-policy load-aware load-aware, least-loaded or round-robin
--fleet-pool-high <pct> 90 load-aware avoids nodes at or above this KV occupancy
--fleet-conns-per-node <N> 1 connections per node, and so concurrent generations there

The head loads no engine but still reads the model for its tokenizer and chat template; its model id must match what the nodes advertise (the head warns about a node that serves another). Its context window is the nodes', so it needs no --max-context. It logs where each session is placed and every move to another node. --dialect, --tool-call-parser and --temperature/--top-p/--top-k/--min-p apply on the head. --api-key, --key-policy and --rate-limit protect its listener; a policy's class and max_class do not apply, since the head has no scheduler.

Refused on the head (set them on each node's own superfluid serve): --class-lanes, --http-default-qos, --no-http-qos-header, --http-allow-batch-invariant, --spec-max-temperature, --fim-model, --repeat-penalty, --otlp-endpoint.

Placement policies

policy how it picks
load-aware (default) prefers nodes whose window fits the prompt, skips nodes above --fleet-pool-high, then the lowest estimated completion time from each node's reported rates. A node measures its rates only while it is generating, so an idle node keeps the speed it showed when busy; one that has not generated yet counts at the mean of the others
least-loaded fewest sessions in progress plus requests in flight on that node, whatever its speed
round-robin rotate through capable nodes; use it to prove a fleet is distributing

A prompt no node can hold is refused with context_length_exceeded.

What the head serves

  • POST /v1/chat/completions, streaming or not, with tools (OpenAI API).
  • GET /v1/models: the one model the fleet serves.

Refused with 400: non-zero presence_penalty/frequency_penalty, a repeat_penalty other than 1.0, ignore_eos, stream_options.continuous_usage_stats, and image or audio parts. Not read: logprobs, logit_bias, response_format, tool_choice. There is no /health, /metrics, Anthropic, Ollama, embeddings, files or batches surface on the head. A generation has a ten-minute wall-clock ceiling.

Failure handling

event what happens
node fails mid-stream delivered tokens become context; the session is re-placed on the next capable node, which generates the rest. No token is delivered twice
node unreachable skipped; the head tries it again every 10 s and uses it once it answers
node restarts comes back empty; the head reconnects and new placements use it
connection idle the head sends a keepalive every 15 s, under the node's --idle-timeout-ms; a connection a node closed anyway is replaced before it is used
head restarts in-flight requests fail; placement state is lost
head misses a lease renewal the node fences the session and emits nothing more

Limits

  • One model per fleet.
  • No KV transfer between nodes: a failover re-prefills.
  • --fleet-conns-per-node is the only concurrency lever on a node, whatever its lane count.
  • No QoS class crosses the wire.