Speculative decoding¶
A drafter proposes several tokens per round and the target model verifies them in one forward pass. Every draft is verified, so a poor drafter costs speed, never output. Every strategy runs on the basert runtime; prompt-lookup also runs on llamacpp and mlx. Measured gates turn speculation off where it does not pay.
superfluid serve ~/models/Qwen3-4B-Q4.base --speculate auto
superfluid serve ~/models/Qwen3-4B-Q4.base --speculate dflash:~/models/drafters/Qwen3-4B-DFlash-Q4.base
superfluid serve ~/models/GLM-5.2-Q4.base --speculate mtp-head
The startup log names the strategy it registered (superfluid: speculation: mtp-head).
Strategies¶
| directive | drafter | needs |
|---|---|---|
off (default) |
none | |
auto |
the bundle's own MTP head, else the best installed drafter that fits, else plain decoding | |
mtp-head |
the multi-token-prediction head in the target bundle | a bundle converted with the head |
prompt-lookup |
n-gram drafting from the lane's own context | nothing; drafts only where text repeats |
dflash:<path> |
a DFlash block drafter | a .base drafter |
dspark:<path> |
a DSpark drafter | a .base drafter |
eagle3:<path> |
an EAGLE-3 head | a .base drafter |
draft-model:<path> |
a smaller model with the same tokenizer | a .base model |
<path> |
whatever the file's header says | a .base file |
On llamacpp and mlx, prompt-lookup is the one strategy, and auto means prompt-lookup. It drafts the tokens that followed the last earlier occurrence of the lane's last three (else two) tokens, so it pays where a reply copies its context: code edits, quoting, structured repeats. Verification is distribution-exact, greedy and sampled. MLX drafts while one lane decodes.
Note
A model that cannot cut a rejected draft off its KV (recurrent or sliding-window state) refuses prompt-lookup at startup in its runtime's words, and auto decodes plainly. Any other strategy named on llamacpp or mlx is refused.
With several models, the directive is resolved per model.
How auto picks a drafter¶
- The target bundle's own head, if it has one.
- Otherwise installed DSpark, DFlash and EAGLE-3 drafters whose header fits the target (same hidden size, vocabulary and backend), searched in the target's directory, a
drafters/directory beside it or one level up, and the baseRT models cache. Nothing is downloaded. A plain draft model is never picked. - Among fitting drafters: one the hub catalog pairs with the target, then one whose name contains the target's, then DSpark, DFlash, EAGLE-3, then the target's quantization, then the newest.
If the engine refuses the first pick, auto tries the next and finally decodes plainly. An explicit directive has no fallback: a refusal is a startup error.
The hub catalog is fetched from https://raw.githubusercontent.com/basecompute/baseRT/main/base-convert/crates/base-hub/catalog.json into $SUPERFLUID_HOME/cache/hub-catalog.json (refreshed daily, never with --offline), plus the one baseRT caches at $BASERT_MODELS_DIR/.catalog-cache.json. A drafter it lists (a pulled <org>/<repo>/<variant>/model.base, or a file under a catalog file name) must match the catalog's size and sha256: auto skips one that does not and says so, and an explicit directive naming it is a startup error. Each file is hashed once and its digest kept in $SUPERFLUID_HOME/cache/sha256/ until the file changes.
Gates¶
| gate | flag | default | behaviour |
|---|---|---|---|
| adaptive depth | --spec-adaptive |
on | picks the fastest draft depth per lane count |
| yield floor | --spec-min-yield, --spec-yield-rounds |
0.75, 24 | a request whose drafter returns fewer accepted tokens per round stops drafting; two failing requests and none passing abandon the strategy for good |
| throughput gate | --spec-throughput-gate, --spec-min-speedup |
on, 1.08 | measures speculating against plain decoding per lane count and decodes plainly where speculation is not 1.08x faster |
| re-probe | --spec-gate-reprobe, --spec-gate-reprobe-max |
32, 1024 | after a throughput abandonment, re-measure after 32 requests, doubling up to 1024 |
Speculation is usually a low-concurrency win that can invert under load; the gates keep their verdicts per lane count. All flags are in the CLI reference.
Default draft depth: mtp-head lets the engine choose (3; 1 on GLM-DSA), eagle3 2, dflash/dspark the drafter's block width, prompt-lookup and draft-model 4.
Interactions¶
- Exactness. Greedy output with speculation equals greedy output without. Sampled output follows the same distribution, but seeded sampled text can differ.
- Grammars,
logit_bias,logprobs. These lanes never speculate. - Penalties. Allowed, but drafters predict the unpenalized distribution, so acceptance drops.
- Temperature. A flat distribution rejects most drafts.
--spec-max-temperaturecaps a temperature that came from the model's default (never one a caller or--temperatureset). - Streaming. An accepted draft arrives as one chunk carrying several tokens;
stream_options.continuous_usage_statsreports cumulative token counts per chunk.
What is reported¶
superfluid.spec(proposed,accepted,acceptance_rate) on chat replies and the final stream chunk.superfluid_spec_proposed_totalandsuperfluid_spec_accepted_totalon/metrics.spec acceptin the terminal monitor.
Expected speedups¶
With Qwen3-4B q4 at one lane, DSpark, DFlash and EAGLE-3 drafters measured 130, 123 and 98 tokens per second (the order auto prefers them in). The Qwen3.8-27B MTP head accepts 41% of its drafts at temperature 1.0 against 59% at temperature 0, which is the measurement behind --spec-max-temperature. Measure on your own hardware before relying on a figure.