# OpenAI-Compatible HTTP API

> Call deployed inference models over the OpenAI-compatible HTTP surface: base URL, endpoints, SSE streaming over HTTP/3, and per-model tool calling.

Deployed inference models are reachable over an OpenAI-compatible HTTP
API. Any client that speaks the OpenAI REST shape — the official SDKs,
`curl`, or your own HTTP client — works against it after you change the
base URL and the API key.

This page covers that HTTP surface. It is separate from the
speech-to-speech agent path, where a model is attached to a call through
the Voice surface instead of being called directly.

## Base URL and authentication

There is a single, operator-global base URL. Per-tenant URL routing is
not implemented in `mod_inference` today, so every tenant sends requests
to the same host and the **API key identifies the caller**. The console's
**Developers → Base URL** card shows the exact value and has a copy
button next to it; treat that card as the source of truth rather than
hardcoding a URL from documentation.

Store the value alongside your key:

```bash
export INFERENCE_BASE_URL=…        # copy from Developers → Base URL
export CLUTCHCALL_CREDENTIALS=…       # your ClutchCall API key
```

The key travels as a bearer token:

```
Authorization: Bearer $​CLUTCHCALL_CREDENTIALS
Content-Type: application/json
```

Because the key selects the tenant, it also selects which deployed
models you can address. A `model` id that is not deployed for the
tenant behind the key is not addressable.

## Endpoint reference

Five OpenAI-compatible paths are exposed under the base URL:

| Path                | Purpose                                                                 |
| ------------------- | ----------------------------------------------------------------------- |
| `/chat/completions` | Chat-style generation. Supports `stream: true` and the `tools` field.   |
| `/completions`      | Legacy text completion for models that expose a raw prompt interface.    |
| `/embeddings`       | Vector embeddings from a deployed embedding model.                       |
| `/models`           | Lists the model ids addressable with the presented API key.              |
| `/health`           | Liveness of the inference surface. Use it for probes and uptime checks.  |

The `model` field in every request body takes a deployed **model id**.
Read the ids from `/models` (or from the console's model list) rather
than guessing them — the snippets in the console ship with a
`<model-id>` placeholder for exactly this reason.

A minimal non-streaming chat request:

```bash
curl $INFERENCE_BASE_URL/chat/completions \
  -H "Authorization: Bearer $​CLUTCHCALL_CREDENTIALS" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model-id>",
    "messages": [{ "role": "user", "content": "Hello" }]
  }'
```

## Streaming over HTTP/3

Set `stream: true` and the response is delivered as Server-Sent Events,
carried over QUIC / HTTP/3. Each event is an OpenAI-shaped chunk, so
`chunk.choices[0].delta.content` is where incremental text arrives.

Two consequences worth knowing:

- **No client changes are required.** The OpenAI SDKs iterate the SSE
  stream exactly as they do against any other OpenAI-compatible server.
  HTTP/3 negotiation happens underneath the HTTP client.
- **0-RTT resumption cuts reconnect latency.** When a client reconnects
  to a host it has already talked to, the QUIC session can resume
  without a full handshake round trip.

```bash
curl $INFERENCE_BASE_URL/chat/completions \
  -H "Authorization: Bearer $​CLUTCHCALL_CREDENTIALS" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model-id>",
    "stream": true,
    "messages": [{ "role": "user", "content": "Hello" }]
  }'
```

## Enabling tool calling per model

Tool / function calling is enabled **server-side, per model**, in the
deploy spec. vLLM needs two flags:

```
--enable-auto-tool-choice
--tool-call-parser <parser>
```

`<parser>` is the tool-call parser appropriate to the model you are
serving. Set both on the model's entry in the deploy spec; they are not
request-level options, so a client cannot turn tool calling on for a
model that was deployed without them.

Once the flags are set, the OpenAI `tools` field works unchanged:

```ts
const res = await client.chat.completions.create({
  model: "<model-id>",
  messages: [{ role: "user", content: "What is the weather in Lisbon?" }],
  tools: [
    {
      type: "function",
      function: {
        name: "get_weather",
        parameters: {
          type: "object",
          properties: { city: { type: "string" } },
          required: ["city"],
        },
      },
    },
  ],
});
```

If tool calls never come back for a model that you expect to support
them, check the deploy spec for that model before debugging the client —
a missing `--enable-auto-tool-choice` presents as a plain text answer.

## Using the official OpenAI SDKs

Point the SDK's base URL at the inference base URL and pass your
ClutchCall API key. Nothing else changes. These are the snippets the
console's **SDK quickstart** card emits.

**Python**

```python
from openai import OpenAI

client = OpenAI(
    base_url=os.environ["INFERENCE_BASE_URL"],
    api_key=os.environ["CLUTCHCALL_CREDENTIALS"],
)

stream = client.chat.completions.create(
    model="<model-id>",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")
```

**Node**

```ts
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: process.env.INFERENCE_BASE_URL,
  apiKey: process.env.CLUTCHCALL_CREDENTIALS,
});

const stream = await client.chat.completions.create({
  model: "<model-id>",
  messages: [{ role: "user", content: "Hello" }],
  stream: true,
});
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
```

Keep the API key on your backend. It is a long-lived, tenant-scoped
credential; see [Authentication](/concepts/authentication) for the
pattern to use when a browser needs access.

## Related

- [Authentication](/concepts/authentication) — API keys and how they scope a caller
- [Modalities overview](/modalities/overview)
