Deployed inference models are reachable over an OpenAI-compatible HTTP API. Any client that speaks the OpenAI REST shape — the official SDKs, curl, or your own HTTP client — works against it after you change the base URL and the API key. This page covers that HTTP surface. It is separate from the speech-to-speech agent path, where a model is attached to a call through the Voice surface instead of being called directly.

Base URL and authentication

There is a single, operator-global base URL. Per-tenant URL routing is not implemented in mod_inference today, so every tenant sends requests to the same host and the API key identifies the caller. The console’s Developers → Base URL card shows the exact value and has a copy button next to it; treat that card as the source of truth rather than hardcoding a URL from documentation. Store the value alongside your key:
The key travels as a bearer token:
Because the key selects the tenant, it also selects which deployed models you can address. A model id that is not deployed for the tenant behind the key is not addressable.

Endpoint reference

Five OpenAI-compatible paths are exposed under the base URL: The model field in every request body takes a deployed model id. Read the ids from /models (or from the console’s model list) rather than guessing them — the snippets in the console ship with a <model-id> placeholder for exactly this reason. A minimal non-streaming chat request:

Streaming over HTTP/3

Set stream: true and the response is delivered as Server-Sent Events, carried over QUIC / HTTP/3. Each event is an OpenAI-shaped chunk, so chunk.choices[0].delta.content is where incremental text arrives. Two consequences worth knowing:
  • No client changes are required. The OpenAI SDKs iterate the SSE stream exactly as they do against any other OpenAI-compatible server. HTTP/3 negotiation happens underneath the HTTP client.
  • 0-RTT resumption cuts reconnect latency. When a client reconnects to a host it has already talked to, the QUIC session can resume without a full handshake round trip.

Enabling tool calling per model

Tool / function calling is enabled server-side, per model, in the deploy spec. vLLM needs two flags:
<parser> is the tool-call parser appropriate to the model you are serving. Set both on the model’s entry in the deploy spec; they are not request-level options, so a client cannot turn tool calling on for a model that was deployed without them. Once the flags are set, the OpenAI tools field works unchanged:
If tool calls never come back for a model that you expect to support them, check the deploy spec for that model before debugging the client — a missing --enable-auto-tool-choice presents as a plain text answer.

Using the official OpenAI SDKs

Point the SDK’s base URL at the inference base URL and pass your ClutchCall API key. Nothing else changes. These are the snippets the console’s SDK quickstart card emits. Python
Node
Keep the API key on your backend. It is a long-lived, tenant-scoped credential; see Authentication for the pattern to use when a browser needs access.