# Inference Runtime and Routing

> How the inference runtime publishes its health blob, what the resource gauges and boot self-tests show, and how routing strategies, weighted scorers, and metric profiles are published to a live module.

The **Runtime** screen is the operator view of a single inference host: the vLLM
runtime running under AIBrix, its GPU/CPU/memory state, its boot self-tests, the
server arguments it was started with, and the live routing strategy that
`mod_inference` applies when it picks a pod for a request.

Everything on the screen comes from two read procedures — `runtime.health` and
`runtime.routing` — plus one write, `runtime.setRouting`. Nothing on the screen
reaches into the cluster directly.

## What the runtime publishes and where it is stored

The runtime does not answer console queries. It **publishes a blob**, and the
console reads it.

| Data | Source | Read by |
| ---- | ------ | ------- |
| Host identity, version, uptime, engine | operator-published runtime blob in Redis (`clutchcall:inference:runtime`) | `runtime.health` |
| GPU / CPU / memory gauges | same blob | `runtime.health` |
| Self-test results | same blob | `runtime.health` |
| vLLM server arguments | same blob | `runtime.health` |
| Desired routing strategy | routing key published by the console | `runtime.routing` → `desired` |
| Effective routing strategy | `<routing key>.effective`, republished by a live module | `runtime.routing` → `effective` |

Two consequences follow from this shape, and they explain most of the screen's
behaviour:

- A field shown as `—` is a field that was **absent from the published blob**,
  not a zero. The console does not substitute a default.
- Freshness is the publisher's property. The console renders what is in Redis at
  the moment of the query.

## Host banner fields

The banner across the top of the screen is read straight from the health blob,
with one exception.

| Field | Meaning |
| ----- | ------- |
| Host | The runtime's own reported hostname. |
| Version | The runtime version string it published. |
| Uptime | Process uptime as reported by the runtime. |
| Engine | The engine the runtime identifies itself as. |
| Listener | Rendered as a fixed label (`:4433 · QUIC/H3`) to match the port in the server arguments. It is not read from the blob. |
| Systems nominal pill | A count: self-tests with status `pass` over total self-tests in the blob. |

## Reading the resource gauges

Three cards show host resources. The donut in each card is a percentage; the
bars beside it are meters.

**GPU** — device name, compute utilization (donut), then VRAM used against VRAM
total, temperature, and power draw against the reported power cap.

**CPU** — model and core count, load (donut), then 1-minute loadavg, the
configured tensor-parallel width, and `max-model-len` rendered in thousands of
tokens.

**System memory** — host RAM total, used as a share of total (donut), then the
number of memory-mapped models, page cache, and swap.

Two rendering rules matter when you are reading a number off this screen:

- **The memory donut needs a total.** If the blob reports no memory total, the
  ratio is undefined, so the donut shows `0%` and the sub-label reads
  `not reported` rather than a computed `NaN%`.
- **Some bars are relative, not capacity limits.** For rows where the blob
  carries no natural maximum — tensor-parallel width, `max-model-len`, models
  mapped, page cache, swap, loadavg — the bar's full scale is the larger of the
  reported value and a fixed display floor. The bar is there so you can see
  movement; read the number in the right-hand label, not the fill fraction.

VRAM, temperature, and power are the rows that do have a real denominator
(`vramTotal`, a fixed temperature scale, `powerCap`), so their fill fractions
are meaningful.

## Runtime config and what each vLLM argument means

The **Runtime config** card lists the vLLM server arguments for this runtime.
The first five rows come from the health blob; the last two are rendered as
fixed rows in the console.

| Argument | From blob | What it controls |
| -------- | --------- | ---------------- |
| `--served-model-name` | yes (`servedModel`) | The model name clients address in requests, and the name that scraped metrics are keyed by. |
| `--tensor-parallel-size` | yes (`tensorParallel`) | How many GPUs the model's weights are sharded across within one runtime instance. Also shown as the CPU card's `tensor-parallel` row. |
| `--gpu-memory-utilization` | yes (`gpuMemUtil`) | The fraction of GPU memory vLLM is allowed to pre-allocate, which sets the size of the KV cache pool. |
| `--max-model-len` | yes (`maxModelLen`) | The maximum context length in tokens. Also drives the CPU card's `max-model-len` row. |
| `--dtype` | yes (`dtype`) | The weight/activation datatype the runtime loaded with. |
| `--enable-auto-tool-choice` | no — fixed `true` | Lets the model select tools itself rather than requiring the caller to name one. |
| `--host` / `--port` | no — fixed `0.0.0.0 --port 4433` | The bind address and the port the banner's listener label refers to. |

Because the last two rows are constants in the console, they describe how the
runtime is expected to be launched — they are not evidence of what this
particular process was given.

## Self-tests and how to read them

Each entry in the self-test list is a boot diagnostic that the runtime
published, with four fields:

| Field | Meaning |
| ----- | ------- |
| `name` | The diagnostic's name, as the runtime named it. |
| `detail` | The runtime's own one-line result string. This is where the substance is: which device, which path, which endpoint. |
| `ms` | How long that check took, in milliseconds. |
| `status` | `pass`, `warn`, or a failure. `pass` renders green with a check; everything else renders with an alert icon — `warn` amber, failures red. |

The set of tests is **not fixed by the console**. The screen renders whatever
tests are present in the blob, in the order published. If a check you expect is
missing, the runtime did not report it; that is a publisher-side gap, not a pass.

Use `ms` as a signal in its own right. A check that passes but takes far longer
than its neighbours usually means the resource it touched was slow to respond,
which tends to show up later as latency rather than as a failure.

### What "Run self-test" actually does

The **Run self-test** button calls `refetch()` on `runtime.health`. It re-reads
the published runtime/self-test blob. It does **not** instruct the runtime to
re-run its boot diagnostics, and it does not fabricate a progress animation.

So:

- If the runtime has re-published since your last look, you will see new results.
- If it has not, you will see the same results again, and the timestamp label
  still reads `last run just now` — that label describes the fetch, not the
  diagnostic run.

To genuinely re-run boot diagnostics you restart the runtime, which is an
operator action (see below).

## Routing strategies and the metrics they require

`runtime.routing` returns the strategy list for the org; the console renders one
button per strategy and does not hardcode names. Selecting a button only sets a
local draft — nothing is published until you press **Apply strategy**.

The important field alongside the list is `requires`: a per-strategy list of the
**scraped metrics that strategy needs in order to score at all**. The console
shows it inline under the strategy picker:

- If the selected strategy requires metrics, the hint names them and warns that
  without them the router **falls back to a random ready pod**. This is the
  failure mode worth internalising: a KV-aware or load-aware strategy whose
  metrics are not being scraped does not error. It silently degrades to random
  selection, and your latency profile quietly becomes that of no routing at all.
- If the strategy requires no scraped metrics, the hint says so — that strategy
  routes from prompt affinity and live in-flight load, both of which the module
  observes itself.

Before you adopt a metrics-dependent strategy, confirm the metrics named in the
hint are actually being scraped for the pods you are routing across.

## Weighted multi-scorer routing

The `multi (weighted)` option composes several scorers instead of picking one.
Its strategy string is `multi:` followed by comma-separated `name:weight` pairs.
Each scorer is normalized to `[0,1]` and the normalized scores are combined by
weight. The console seeds a starting spec of:

```
multi:least-request:3,least-kv-cache:1
```

`runtime.routing` returns `multiScorers` — the set of scorer names that can be
weighted. This set is narrower than the full strategy list: not every strategy
is weightable.

The spec you type goes to the live router **verbatim**, so the console parses it
before publishing and shows the first problem instead of publishing a broken
spec:

| Condition | What you get |
| --------- | ------------ |
| Spec is empty | Asked to enter at least one `name:weight` pair. |
| An empty entry (e.g. a double comma) | Asked to remove it and separate pairs with single commas. |
| An entry with no `:`, or a missing name or weight | Told the entry is not a `name:weight` pair, with an example. |
| A name that is not in `multiScorers` | Told that name is not weightable, and given the list that is. |
| A weight that is not a finite number above zero | Told the weight must be above zero. |

Parsing splits each pair at its **last** colon, so the name is everything before
it and the weight everything after. Validation errors surface when you press
**Apply strategy**; the input turns invalid and the message appears rather than
the mutation firing.

## Metric profiles per engine

The **Metric profile (static routes)** control selects **which engine's
`/metrics` naming the scraper parses for pods in the static route table**.
Pods discovered dynamically use their own engine label instead, so this setting
only governs statically-routed pods.

| Profile | What the scraper reads, and what routes well |
| ------- | ------------------------------------------- |
| `vLLM` | The vLLM metric names. The default profile. |
| `Triton` | `tritonserver` core metrics (`nv_inference_*`). KV-aware routing needs the TensorRT-LLM backend; with other backends the KV-aware strategies stay cold. |
| `Triton+vLLM` | `tritonserver` with the vLLM backend: KV occupancy from `vllm:*`, GPU utilization from Triton core. Metrics are keyed by the routed model name, which assumes single-model pods. |
| `FreeToken` | Read via the sidecar exporter at `/v1/stats`. Requests cap at `--max-running-requests` (default 4), so pods saturate early — spread with `least-request` across them. No queue depth is exposed. |
| `colibri` | Served by `coli_gateway` (native C, no sidecar). One request per engine, so `least-request`/`least-busy` is the fit; `num_requests_waiting` is the gateway's live queue depth. No KV or throughput rows — tiering makes them ill-defined. |

The console publishes `routesEngine` **only when you actually pick a profile**.
The value displayed is the effective profile falling back to the desired one,
and re-publishing a displayed value would rewrite a field you never touched —
worse, it would coerce an engine string the console does not model (for example
`sglang`) into `vllm`. Touch the control and it is sent; leave it alone and
`setRouting` carries the strategy only.

## Desired state, effective state, and the poll cycle

The routing card keeps two facts apart and never conflates them:

| Shown as | Field | Meaning |
| -------- | ----- | ------- |
| **In force (engine)** | `effective.strategy` | What a running `mod_inference` reports it has **applied**. |
| **Requested** | `desired.strategy` | What was last **published** from this console. `— (module default)` when nothing has been published. |
| **Metric profile** | `effective.routes_engine`, falling back to `desired.routes_engine` | The metric naming in use for static routes. |
| **Source** | `effective.source` | Where the reporting module says the strategy in force came from. |

The effective record is a **TTL'd beacon**: the module republishes it on every
poll. `engineReporting` is true only while that beacon is present. The status
pill reflects the combination:

| Pill | Condition |
| ---- | --------- |
| `engine in sync` | The engine is reporting and `inSync` is true (or nothing has been requested yet). |
| `pending pickup` | The engine is reporting, but what it has applied differs from what was requested. |
| `engine not reporting` | No beacon. |

When there is no beacon, the card explains that no live `mod_inference` is
publishing routing state at `<key>.effective`, and that a strategy published
here is **stored as desired state and adopted when a module starts polling**.
ClutchCall deliberately does not show a green confirmed strategy in that
situation, because a strategy no engine has acknowledged may be steering nothing.

Switching strategies takes effect on the module's **next poll (~5 s)**. There is
no restart, and no request in flight is re-routed.

Two details of the control's behaviour follow from this model:

- The strategy picker initialises from **what is actually running** (`effective`),
  falling back to what was requested. It shows reality first, intent second.
- **Apply strategy** is enabled only when your draft differs from the applied
  strategy, or when you have explicitly picked a metric profile. Otherwise
  there is nothing to publish. **Reset** discards the draft and returns the
  picker to the applied state.

If the org has no strategies at all, the card replaces the controls with an
empty state: routing data could not be loaded, check that inference routing is
enabled for the org, then retry.

## Read-only deployments and operator-only actions

Two different restrictions apply on this screen, and they have different causes.

**Routing writes can be disabled per deployment.** `runtime.routing` returns
`writable`. When it is false, every strategy button, the weights input, the
metric-profile selector, and **Apply strategy** are disabled, the card shows
`Read-only on this deployment`, and the apply button's tooltip names the gate:
`INFERENCE_DEPLOY_ENABLED`. Reads continue to work — you can still see what is
in force, what was requested, and which metrics each strategy needs.

**Restarting a worker is not a console capability at all.** The **Restart
worker** button is permanently disabled, with the tooltip *"Restart is
