curl, and stops there. Completions run over the
OpenAI-compatible endpoint with your own key — the console never sends the
request, never streams a response, and never reports timing or session
metrics for one.
The screen has two tabs:
What the builder emits
For the current control values the screen renders two blocks, plus a Copy curl button that writes the samecurl string to your clipboard.
The curl block:
- The base URL is the fabric’s OpenAI-compatible base,
https://fabric.clutchcall.dev/v1, with the builder’s path/v1/chat/completionsappended. The path and the transport (QUIC/H3 · SSE) are shown in the Endpoint section of the parameter sidebar. - The credential is emitted as a shell variable reference, not a literal. The builder interpolates the platform’s credential environment variable name into the header, so the copied command carries no secret. Export the variable in the shell where you run the command.
system message is present only when the System prompt field is
non-empty; an empty field omits the entry rather than sending an empty
string. The user message is a fixed "Hello!" placeholder — the builder has
no composer, so the user turn is not editable on this screen. Replace it
after you copy the command.
Choosing a model from the registry
The Model dropdown is populated from the live registry via themodels.local procedure — the set of models actually deployed to the
inference fabric. The builder then filters that list to entries whose
capability set includes chat, and offers each one by its registry name.
The first matching model is selected automatically when the query resolves.
The header strip above the request blocks reflects this:
If the registry returns no chat-capable model, the dropdown is disabled
(
No models deployed) and the right-hand pane replaces the request blocks
with an empty state — No models deployed yet — plus a Go to Models
button. Deploy a chat-capable model first; there is nothing to compose a
request against until then.
Parameters and where they are validated
The sidebar exposes exactly the fields that appear in the emitted body.
Those slider bounds and steps are properties of the console’s controls. They
are not per-model limits and they are not enforced by the console beyond
constraining what you can drag to. The builder performs no validation of
the composed request: it does not check the prompt against the model’s
context window, does not reconcile
max_tokens with that window, and does
not verify your credential. Validation happens at the endpoint when you
actually send the request, and the endpoint’s answer is the authoritative
one. If a value is rejected, you will see it in the HTTP response from your
shell or SDK, not on this screen.
Running the request from a shell or an SDK
Copy the command, export your credential, and run it. The-sN flags are
deliberate: -N disables curl’s output buffering so that a streaming
response arrives incrementally instead of in one block at the end.
The gateway streams Server-Sent Events back over QUIC/H3 and routes the
selected model to its self-hosted engine. With stream toggled off, the
same request returns a single non-streaming response body.
For the SDK equivalent of the same request, use the Developers screen — the
builder points there rather than duplicating client snippets, and the request
body above maps field-for-field onto any OpenAI-compatible client.
Voice pipeline: what is not wired yet
The Voice pipeline tab is an empty state in this build, by design. No live voice-pipeline telemetry reaches the operator console, so the screen renders an honest placeholder instead of a synthetic mic level, waveform, or stage waterfall. Concretely, the tab contains:- A Voice pipeline tester panel explaining that voice pipelines (ASR → LLM → TTS) are assembled from the provider vendors the fabric routes to, and that live mic testing appears here once a pipeline and its runtime telemetry are connected. It links to the Providers screen.
- A Waterfall trace card whose subtitle is
No voice pipeline trace, holding a No trace samples empty state. Stage timings will populate it when runtime telemetry is connected.
Related
- Telemetry — the streams the gateway does emit
- Authentication — API keys and how the credential variable is populated

