Ollama-compatible API¶
superfluid serves a subset of Ollama's API (inference, discovery and embeddings) on its HTTP listener, for any runtime. It does not read Ollama's model store, pull models or interpret Modelfiles.
Point an Ollama client at the server root:
superfluid serve unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M --port 11434
curl http://127.0.0.1:11434/api/tags
Note
Use a different port if Ollama itself is listening on 11434.
from ollama import Client
client = Client(host="http://127.0.0.1:11434")
model = client.list().models[0].model
for part in client.chat(model=model, messages=[{"role": "user", "content": "Hello"}], stream=True):
print(part.message.content, end="", flush=True)
superfluid launch ollama runs the ollama command line against the server; see Ollama CLI.
Model names are superfluid's model ids; id:latest is accepted. Load models at startup or through POST /v1/models/load; arbitrary paths and remote names are refused.
Routes¶
| method | path | behaviour |
|---|---|---|
GET |
/ |
superfluid is running |
GET |
/api/version |
{"version", "server": "superfluid", "compatibility": "ollama-inference"} |
GET |
/api/tags |
every known model id |
GET |
/api/ps |
loaded models and their context length |
POST |
/api/show |
model info; for a loaded model, capabilities and chat template |
POST |
/api/chat |
chat with history, tools, thinking, images, JSON output |
POST |
/api/generate |
templated prompt; raw for plain completion; suffix for fill-in-the-middle |
POST |
/api/embed |
normalized embeddings (embedding models only) |
POST |
/api/embeddings |
legacy single-prompt form, unnormalized |
| any | /api/pull, /api/push, /api/create, /api/copy, /api/delete, /api/blobs/* |
501 |
Discovery routes never load a cold model.
Chat and generate¶
curl -N http://127.0.0.1:11434/api/chat -d '{
"model": "unsloth/Qwen3.8-27B-GGUF",
"messages": [{"role": "user", "content": "Explain prefix caching briefly."}],
"options": {"num_predict": 128, "temperature": 0.2}
}'
curl http://127.0.0.1:11434/api/chat -d '{
"model": "unsloth/Qwen3.8-27B-GGUF", "stream": false,
"messages": [{"role": "user", "content": "Return a planet name as JSON."}],
"format": {"type": "object", "properties": {"name": {"type": "string"}}, "required": ["name"]}
}'
| field | notes |
|---|---|
model |
required |
messages (chat) |
system, user, assistant, tool; images (base64), thinking, tool_calls, tool_name, tool_call_id. Empty = preload |
prompt, system, images, suffix, raw (generate) |
empty prompt with nothing else = preload; an empty suffix is none |
template (generate) |
empty only (the model's own template); any other is a 400 |
tools |
Ollama function tools |
stream |
default true (NDJSON) |
format |
"json" or a JSON schema |
think |
true/false maps to enable_thinking; minimal, low, medium, high to reasoning_effort |
logprobs, top_logprobs |
chat only; top_logprobs 0-20 |
options |
see below; null is none |
keep_alive |
any non-null value is a 400 |
Any other non-null field is a 400 unsupported parameter: <name>. raw and suffix requests refuse system, images, format, think and options.stop.
Options¶
| option | constraint |
|---|---|
num_predict |
positive, or -1/-2 to fill the context |
num_ctx |
must equal the server's context window |
temperature |
≥ 0 |
top_p, min_p |
0-1 |
top_k |
≥ 0 |
seed |
≥ 0, or -1 for random |
repeat_penalty |
> 0 |
presence_penalty, frequency_penalty |
-2 to 2 |
stop |
up to four non-empty strings |
Any other option (num_gpu, num_thread, mirostat, ...) is a 400 unsupported option: <name>.
Records¶
Each streamed line is one JSON object. Reasoning comes back separately as message.thinking (chat) or thinking (generate). The terminal record:
{"model":"Qwen3-4B","created_at":"...","message":{"role":"assistant","content":""},
"done":true,"done_reason":"stop","total_duration":812345678,
"prompt_eval_count":31,"prompt_eval_cached_count":16,"eval_count":42}
done_reasonisstop,length, orloadfor a preload.prompt_eval_cached_count(extension) is the prefix served from the cache.load_duration,prompt_eval_durationandeval_durationare not reported for generation.- Tool calls arrive whole on one record with
done: false, just before the terminal record;function.argumentsis an object. - Errors before the body keep their HTTP status with
{"error": "message"}; an error mid-stream is one{"error": "..."}line.
Embeddings¶
| field | default | notes |
|---|---|---|
input |
string or array; empty = preload | |
truncate |
true |
false refuses oversized input |
dimensions |
the model's | truncate and re-normalize |
llamacpp and mlx serve no embeddings.
Not supported¶
keep_alive(residency is the server's:--idle-timeoutand the model routes)- model distribution and Modelfiles (501)
- Ollama metadata and phase timings:
/api/tagslists a digest of the model id, size 0 and emptydetails;/api/showignoressystem,templateandoptions - the fleet head does not serve these routes