curl, or your own HTTP client — works against it after you change the
base URL and the API key.
This page covers that HTTP surface. It is separate from the
speech-to-speech agent path, where a model is attached to a call through
the Voice surface instead of being called directly.
Base URL and authentication
There is a single, operator-global base URL. Per-tenant URL routing is not implemented inmod_inference today, so every tenant sends requests
to the same host and the API key identifies the caller. The console’s
Developers → Base URL card shows the exact value and has a copy
button next to it; treat that card as the source of truth rather than
hardcoding a URL from documentation.
Store the value alongside your key:
model id that is not deployed for the
tenant behind the key is not addressable.
Endpoint reference
Five OpenAI-compatible paths are exposed under the base URL:
The
model field in every request body takes a deployed model id.
Read the ids from /models (or from the console’s model list) rather
than guessing them — the snippets in the console ship with a
<model-id> placeholder for exactly this reason.
A minimal non-streaming chat request:
Streaming over HTTP/3
Setstream: true and the response is delivered as Server-Sent Events,
carried over QUIC / HTTP/3. Each event is an OpenAI-shaped chunk, so
chunk.choices[0].delta.content is where incremental text arrives.
Two consequences worth knowing:
- No client changes are required. The OpenAI SDKs iterate the SSE stream exactly as they do against any other OpenAI-compatible server. HTTP/3 negotiation happens underneath the HTTP client.
- 0-RTT resumption cuts reconnect latency. When a client reconnects to a host it has already talked to, the QUIC session can resume without a full handshake round trip.
Enabling tool calling per model
Tool / function calling is enabled server-side, per model, in the deploy spec. vLLM needs two flags:<parser> is the tool-call parser appropriate to the model you are
serving. Set both on the model’s entry in the deploy spec; they are not
request-level options, so a client cannot turn tool calling on for a
model that was deployed without them.
Once the flags are set, the OpenAI tools field works unchanged:
--enable-auto-tool-choice presents as a plain text answer.
Using the official OpenAI SDKs
Point the SDK’s base URL at the inference base URL and pass your ClutchCall API key. Nothing else changes. These are the snippets the console’s SDK quickstart card emits. PythonRelated
- Authentication — API keys and how they scope a caller
- Modalities overview

