# Inference overview metrics

> What the six KPI tiles on the inference Overview screen measure, the window they cover, how to read the Active models list, and what authenticates a request to the global gateway endpoint.

The Overview screen of the ClutchCall inference console is a read-only
summary of the fabric. It has three panels: a six-tile KPI strip, an
OpenAI-compatible endpoint card with a cURL quickstart, and the Active
models list. Both the KPI strip and the models list are filled from the
server — nothing on this screen is computed in the browser.

## The KPI strip, tile by tile

The strip is not hardcoded in the console. It renders whatever
`overview.kpis` returns: one tile per element of the returned array, laid
out in a fixed six-column grid, so a response of six entries fills the
row exactly.

| Part of the entry | What the console does with it |
| ----------------- | ----------------------------- |
| `key`             | Identity of the tile. Keep it stable across refetches so a tile does not remount and lose its animation. |
| every other field | Spread straight into the tile component. The caption, the figure, and any trend indicator or sparkline you see come verbatim from the procedure. |

Three consequences worth knowing before you file a bug:

- **The tile's own caption is the authoritative name of the measure.**
  The console has no local table of tile names, units, or "good" and
  "bad" colours.
- **No client-side formatting.** If a figure is rounded, scaled, or
  suffixed, `overview.kpis` did that. The console prints what it is
  given.
- **A missing tile is a missing array element.** If a measure you expect
  is absent, the procedure stopped returning that key. It is not a
  render failure, and the grid simply shows fewer than six columns
  filled.

## How each figure is derived and over what window

`overview.kpis` serves a single **fixed live window**. The console sends
no range argument, and there is no range selector on this screen — an
earlier range control was removed because changing it changed nothing in
the response. You therefore cannot widen, narrow, or back-date the
figures from Overview, and any comparison or trend shown on a tile is
computed server-side inside that same fixed window.

The window only advances when the query actually runs:

- on screen mount, and
- when you press **Refresh**, which refetches `overview.kpis`.

The screen sets no refresh interval of its own, so a strip left open may
be showing the result of an old fetch. Press Refresh before reading a
figure you intend to act on.

**Refresh refetches the KPI strip only.** The Active models list comes
from a separate query (`models.local`) and that button does not re-run
it. If the strip and the list look inconsistent, reload the screen so
both queries run again.

## Reading the Active models list

Each row in the Active models card is one record from `models.local`.
The card is a summary; **Manage** opens the Models screen.

| Row element | Source field | Notes |
| ----------- | ------------ | ----- |
| Name        | `name`       | Rendered in monospace and truncated with an ellipsis when it does not fit. |
| Engine tag  | `engine`     | Shown only if the field is present. The tag colour is derived from the engine name; the card subtitle describes these as vLLM / SGLang runtimes. |
| Namespace   | `namespace`  | Shown only if the field is present, right-aligned in muted monospace. |

**Served state.** The small square to the left of every name is a fixed
accent marker, identical on every row. It is not a per-model health
readout, and this card exposes no health field. The signal here is
membership: a model appears in this list because it is registered with
the fabric and served on the operator's GPUs. For anything beyond
"present / not present", use the Models screen.

**Empty states.** While the query is in flight the card reads
*Loading models…*. Once it resolves with an empty array it reads
*No models deployed*, with the hint that models registered with the
fabric appear here and can be deployed from the Models screen. An empty
list is a legitimate server answer, not an error state.

## One global gateway endpoint and what authenticates a request

Inference is **operator-global**. There is one base URL for the whole
fabric:

```
https://fabric.<brand-domain>/v1
```

- The console derives the domain from the hostname it is served on: a
  `telequick.dev` host yields `fabric.telequick.dev`, and anything else
  falls back to `fabric.clutchcall.dev`. The Base URL row on the card is
  the value you should copy; it is correct for the console you are
  looking at.
- **There is no per-tenant URL path.** The gateway's `mod_inference`
  registers flat, OpenAI-compatible routes on the engine's QUIC/H3
  listener, and per-tenant routing is a later phase in `mod_inference.cc`.
  A URL carrying an org, tenant, or project segment is wrong.
- **The API key authenticates the caller.** Identity comes from the
  bearer token on the request, not from the URL. This is what the card's
  "the API key authenticates the caller" label means.

The routes advertised on the card are:

`/chat/completions` · `/completions` · `/embeddings` · `/models` ·
`/health`

The quickstart the card copies out:

```bash
curl https://fabric.<brand-domain>/v1/chat/completions \
  -H "Authorization: Bearer $CLUTCHCALL_CREDENTIALS" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model-id>",
    "stream": true,
    "messages": [
      { "role": "user", "content": "Hello" }
    ]
  }'
```

`$CLUTCHCALL_CREDENTIALS` is the single credential variable name used by
every snippet in the console. `model` takes an id from the Active models
list or from `GET /v1/models` on the same base URL. With
`"stream": true` the endpoint responds as SSE, which is what the card's
"Drop-in · QUIC/H3 · SSE streaming" subtitle refers to.

## What to check when a tile looks wrong

1. **Refetch first.** The window is fixed and only advances when the
   query runs. A frozen-looking tile is often just the last fetch.
2. **Refresh does not reload the models list.** If the strip disagrees
   with Active models, reload the screen so both queries re-run.
3. **There is no filter behind these figures.** `overview.kpis` takes no
   arguments, so you cannot scope the strip to a model, a namespace, or a
   time range from this screen. Do not read a tile as if it were filtered
   to the model you are testing.
4. **Confirm you are calling the endpoint the card shows.** Hit
   `/health` and `/models` on that exact base URL. Because there is one
   global endpoint, a URL from another environment — or one with a tenant
   path in it — will not be the fabric these tiles describe.
5. **Traffic on the tiles but nothing in Active models** (or the
   reverse) means the model registry and the metrics source disagree.
   Open the Models screen to see the model's own state.
6. **Odd units or captions are payload, not presentation.** The console
   does no formatting, so report the raw `overview.kpis` response rather
   than a screenshot of the tile.

## Related

- [Telemetry](/platform/telemetry) — the gateway's metrics, trace, and
  CDR streams, if you need figures over a window you choose.
- [Telephony Metrics](/glossary/metrics) — standard definitions, formulas,
  and healthy ranges for the numbers the platform reports.
- [Authentication](/concepts/authentication) — how API keys are issued and
  rotated.
