Benchmarks¶
A serving layer should cost nothing over the runtime it serves. These numbers compare superfluid against each runtime's own server on the same machine, model and runtime build.
Setup¶
| machine | Apple M1 Max (64 GB), macOS |
| model | Qwen3-4B, 4-bit: Q4_K_M GGUF for llama.cpp, mlx-community/Qwen3-4B-4bit for MLX |
| runtime builds | llama.cpp b11284; mlx-lm 0.31.3 (mlx 0.32.2), on both sides |
| superfluid | superfluid serve --max-batch 8 --max-context 4096 |
| baselines | llama-server -np 8 -c 32768, mlx_lm.server |
| date | 2026-10-06 |
Workload. Eight coding and explanation prompts, 256-token replies, at 1, 2, 4 and 8 concurrent requests. Throughput is completion tokens over wall time for the whole batch (/v1/completions, raw ChatML, so no chat template is in the comparison), the median of three rounds, each prompt distinct so no prefix cache answers. First-token time and stream smoothness are from /v1/chat/completions streams.
Throughput (aggregate tokens per second)¶
Greedy (temperature 0):
| 1 | 2 | 4 | 8 | |
|---|---|---|---|---|
| superfluid, llama.cpp | 76.7 | 87.6 | 126.5 | 150.8 |
| llama-server | 76.8 | 86.2 | 125.4 | 145.1 |
| superfluid, MLX | 100.2 | 117.5 | 127.8 | 130.2 |
| mlx_lm.server | 96.6 | 96.8 | 96.8 | 96.8 |
Sampled (temperature 0.7, top-p 0.8, top-k 20):
| 1 | 2 | 4 | 8 | |
|---|---|---|---|---|
| superfluid, llama.cpp | 76.3 | 87.2 | 125.3 | 148.8 |
| llama-server | 77.7 | 86.3 | 125.9 | 145.3 |
| superfluid, MLX | 97.0 | 114.7 | 125.4 | 128.3 |
| mlx_lm.server | 90.0 | 90.1 | 90.2 | 90.2 |
mlx_lm.server serves one request at a time, so its total does not grow with concurrency.
Chat streams¶
At eight concurrent chat streams the first token arrives in 0.70 s through superfluid on MLX against 9.42 s from mlx_lm.server, which queues requests; on llama.cpp it is 0.57 s against llama-server's 0.58 s.
Serving semantics¶
Throughput does not show what happens when a worker is killed, the KV pool is oversubscribed, or a chat arrives behind eight agents. tools/suite runs eight such scenarios, each with a stated pass condition (see its README). Apple M5 Pro (64 GB), Qwen3-4B 4-bit, llama.cpp b11284, mlx-lm 0.31.3, Ollama v0.35.1; ✓ met the pass condition, ✗ did not, n/a has no way to run the scenario.
| scenario | superfluid (llama.cpp) | superfluid (MLX) | llama-server | mlx_lm.server | Ollama |
|---|---|---|---|---|---|
| worker kill | ✓ | ✓ | n/a | n/a | ✗ |
| daemon restart | ✓ | ✓ | ✓ | ✗ | ✓ |
| agents + interactive chat | — | ✓ | ✗ | ✗ | ✗ |
| KV exhaustion | ✓ | ✓ | ✓ | ✗ | ✓ |
| unload under traffic | ✓ | ✓ | n/a | n/a | ✓ |
| 200-turn tool session | ✓ | ✓ | ✓ | ✓ | ✓ |
| client disconnect | ✓ | ✓ | ✗ | ✗ | ✓ |
| shared-prefix burst | ✓ | ✓ | ✗ | ✗ | ✗ |
- Agents + interactive chat: with eight agents prefilling ~6k-token prompts, a chat in the
interactiveclass gets its first token in 0.97 s on MLX; 67 s on llama-server, 64 s on Ollama, 120 s on mlx_lm.server. - Shared-prefix burst: 32 requests sharing a 4k-token prompt prefill it 0.91 times in total (llama-server and Ollama: about 8); median first token 6.5 s against llama-server's 30.3 s.
- KV exhaustion: 32 requests against 8 lanes of KV all complete; the burst finishes in 149 s against llama-server's 165 s, and 280 s against mlx_lm.server's 435 s (which resets three connections).
Agreement¶
At temperature 0 a reply is the same text as the runtime's own server gives until two candidate tokens are close enough for rounding to decide. Against llama-server, 6 of 8 replies were identical with each prompt alone and 8 of 8 at four and eight lanes. On MLX the 4-bit model's logits tie exactly at many steps, and a hand-written loop over mlx-lm's own model call parts ways with mlx_lm.generate the same way (1 of 4 identical); the adapter's sampler is checked against the host sampler's rules row by row instead.
Batch scaling¶
| lanes | 1 | 2 | 4 | 8 |
|---|---|---|---|---|
| llama.cpp aggregate tok/s | 76.7 | 87.6 | 126.5 | 150.8 |
| llama.cpp per-lane tok/s | 76.7 | 43.8 | 31.6 | 18.9 |
| MLX aggregate tok/s | 100.2 | 117.5 | 127.8 | 130.2 |
- Aggregate throughput rises with lanes until the engine is compute-bound; the per-lane rate falls.
- Median TTFT roughly doubles from one to eight lanes on both runtimes.
- A larger
--max-batchcosts nothing per tick, but the KV pool is sized for every lane.
How to measure¶
The bench harness is not shipped; the method is.
- Throughput. Send N concurrent requests with N distinct prompts and a fixed reply length (
max_tokens, plusignore_eos: trueon/v1/completions). Divide the sum ofusage.completion_tokensby the wall time from the first send to the last response. Take the median of several rounds. - Templates. Use
/v1/completionswith the prompt pre-formatted to exclude template rendering; use chat streams for TTFT and smoothness. - Access log rates are per request.
decode=is one lane's rate (about an eighth of the aggregate at eight lanes);prefill=includes queueing. The engine's rates are on/slots. - TTFT.
superfluid_ttft_secondson/metricsstarts at admission; the access log'sttftincludes the wait before it. - Warm-up. Discard the first round after a model loads. Make prompts distinct to measure prefill; share prefixes to measure agent workloads. Say which.
- Thermals. Interleave systems round by round and report medians, not best runs.
Tuning¶
| flag | default | effect |
|---|---|---|
--max-batch N |
8 | lanes decoding at once; aggregate throughput rises until compute-bound |
--max-context N |
auto | per-lane window and KV pool size; costs memory, not speed, until the pool no longer fits |
--prefill-budget N |
4096 | 0 trades prompt throughput for decode continuity on slow models |
--tick-target-ms N |
2000 | longer ticks amortize per-tick cost; shorter admit and cancel sooner (capped at 250 on llama.cpp and MLX) |
--class-lanes |
uncapped | reserve lanes between QoS classes |
--speculate |
off | basert only; a low-concurrency win the gates turn off under load |
--kv-bits |
0 | basert only; narrower KV fits more context |
--park |
off | does not raise throughput; sealing copies a finishing session's whole KV on the tick thread |
basert, sampled vs greedy at one lane. A lone sampled lane is sampled on the host (cheaper than one GPU sampling pass); from two lanes rows are sampled on the GPU. Expect one sampled stream to run slightly behind one greedy stream, closing at two lanes.
llama.cpp. No llama.cpp flags pass through; the adapter chooses:
- A KV buffer of
max-contextcells per sequence for plain-attention models, so a step attends only over its own sequence; one shared pool for recurrent and sliding-window models. Per-sequence buffers number the lanes plus two (at least 4, at most 256), so 8 lanes of 8192 hold 10 x 8192 cells, against llama-server's 8 x 8192 for-np 8. (The next power of two above twice the lanes, the earlier choice, was 1.6x the KV for no decode gain: Qwen3-4B Q4_K_M, 8 lanes, M4 Max: 181.3 tok/s at 9 buffers, 180.6 at 12, 178.3 at 16.) - The shared pool is
max-batch × max-contextcells capped at 60% of the device memory the weights leave free. - The prefix cache keeps one lane's context of a shared pool and holds the rest in host memory, so its entries cost the lanes nothing while they wait. On gemma-3-4b (M5 Pro,
--max-context 8192 --max-batch 4), after six 2,200-token documents were cached a cold 28 to 94-token prompt's first token took 110 to 118 ms, and 67 to 77 ms with the cache bounded and the batch filled to a multiple of 8 (the Metal kernel computes a short last group of queries against every cell in use). - An entry pushed out for a sequence is kept in host memory as well. Six conversations returning to a Qwen3-4B daemon with
--max-batch 2(four buffers) were all cold, 887 ms to a second turn's first token; with the entries held in host memory all six were warm, 98 ms. n_batch2048,n_ubatch512, every layer on the GPU.
MLX. The mlx and mlx-lm versions in the install set the kernels; pin them when comparing. The adapter prefills in 2048-token passes and sizes its pool as the llama.cpp adapter above does.