The Runtime screen is the operator view of a single inference host: the vLLM runtime running under AIBrix, its GPU/CPU/memory state, its boot self-tests, the server arguments it was started with, and the live routing strategy that mod_inference applies when it picks a pod for a request. Everything on the screen comes from two read procedures — runtime.health and runtime.routing — plus one write, runtime.setRouting. Nothing on the screen reaches into the cluster directly.

What the runtime publishes and where it is stored

The runtime does not answer console queries. It publishes a blob, and the console reads it. Two consequences follow from this shape, and they explain most of the screen’s behaviour:
  • A field shown as — is a field that was absent from the published blob, not a zero. The console does not substitute a default.
  • Freshness is the publisher’s property. The console renders what is in Redis at the moment of the query.

Host banner fields

The banner across the top of the screen is read straight from the health blob, with one exception.

Reading the resource gauges

Three cards show host resources. The donut in each card is a percentage; the bars beside it are meters. GPU — device name, compute utilization (donut), then VRAM used against VRAM total, temperature, and power draw against the reported power cap. CPU — model and core count, load (donut), then 1-minute loadavg, the configured tensor-parallel width, and max-model-len rendered in thousands of tokens. System memory — host RAM total, used as a share of total (donut), then the number of memory-mapped models, page cache, and swap. Two rendering rules matter when you are reading a number off this screen:
  • The memory donut needs a total. If the blob reports no memory total, the ratio is undefined, so the donut shows 0% and the sub-label reads not reported rather than a computed NaN%.
  • Some bars are relative, not capacity limits. For rows where the blob carries no natural maximum — tensor-parallel width, max-model-len, models mapped, page cache, swap, loadavg — the bar’s full scale is the larger of the reported value and a fixed display floor. The bar is there so you can see movement; read the number in the right-hand label, not the fill fraction.
VRAM, temperature, and power are the rows that do have a real denominator (vramTotal, a fixed temperature scale, powerCap), so their fill fractions are meaningful.

Runtime config and what each vLLM argument means

The Runtime config card lists the vLLM server arguments for this runtime. The first five rows come from the health blob; the last two are rendered as fixed rows in the console. Because the last two rows are constants in the console, they describe how the runtime is expected to be launched — they are not evidence of what this particular process was given.

Self-tests and how to read them

Each entry in the self-test list is a boot diagnostic that the runtime published, with four fields: The set of tests is not fixed by the console. The screen renders whatever tests are present in the blob, in the order published. If a check you expect is missing, the runtime did not report it; that is a publisher-side gap, not a pass. Use ms as a signal in its own right. A check that passes but takes far longer than its neighbours usually means the resource it touched was slow to respond, which tends to show up later as latency rather than as a failure.

What “Run self-test” actually does

The Run self-test button calls refetch() on runtime.health. It re-reads the published runtime/self-test blob. It does not instruct the runtime to re-run its boot diagnostics, and it does not fabricate a progress animation. So:
  • If the runtime has re-published since your last look, you will see new results.
  • If it has not, you will see the same results again, and the timestamp label still reads last run just now — that label describes the fetch, not the diagnostic run.
To genuinely re-run boot diagnostics you restart the runtime, which is an operator action (see below).

Routing strategies and the metrics they require

runtime.routing returns the strategy list for the org; the console renders one button per strategy and does not hardcode names. Selecting a button only sets a local draft — nothing is published until you press Apply strategy. The important field alongside the list is requires: a per-strategy list of the scraped metrics that strategy needs in order to score at all. The console shows it inline under the strategy picker:
  • If the selected strategy requires metrics, the hint names them and warns that without them the router falls back to a random ready pod. This is the failure mode worth internalising: a KV-aware or load-aware strategy whose metrics are not being scraped does not error. It silently degrades to random selection, and your latency profile quietly becomes that of no routing at all.
  • If the strategy requires no scraped metrics, the hint says so — that strategy routes from prompt affinity and live in-flight load, both of which the module observes itself.
Before you adopt a metrics-dependent strategy, confirm the metrics named in the hint are actually being scraped for the pods you are routing across.

Weighted multi-scorer routing

The multi (weighted) option composes several scorers instead of picking one. Its strategy string is multi: followed by comma-separated name:weight pairs. Each scorer is normalized to [0,1] and the normalized scores are combined by weight. The console seeds a starting spec of:
runtime.routing returns multiScorers — the set of scorer names that can be weighted. This set is narrower than the full strategy list: not every strategy is weightable. The spec you type goes to the live router verbatim, so the console parses it before publishing and shows the first problem instead of publishing a broken spec: Parsing splits each pair at its last colon, so the name is everything before it and the weight everything after. Validation errors surface when you press Apply strategy; the input turns invalid and the message appears rather than the mutation firing.

Metric profiles per engine

The Metric profile (static routes) control selects which engine’s /metrics naming the scraper parses for pods in the static route table. Pods discovered dynamically use their own engine label instead, so this setting only governs statically-routed pods. The console publishes routesEngine only when you actually pick a profile. The value displayed is the effective profile falling back to the desired one, and re-publishing a displayed value would rewrite a field you never touched — worse, it would coerce an engine string the console does not model (for example sglang) into vllm. Touch the control and it is sent; leave it alone and setRouting carries the strategy only.

Desired state, effective state, and the poll cycle

The routing card keeps two facts apart and never conflates them: The effective record is a TTL’d beacon: the module republishes it on every poll. engineReporting is true only while that beacon is present. The status pill reflects the combination: When there is no beacon, the card explains that no live mod_inference is publishing routing state at <key>.effective, and that a strategy published here is stored as desired state and adopted when a module starts polling. ClutchCall deliberately does not show a green confirmed strategy in that situation, because a strategy no engine has acknowledged may be steering nothing. Switching strategies takes effect on the module’s next poll (~5 s). There is no restart, and no request in flight is re-routed. Two details of the control’s behaviour follow from this model:
  • The strategy picker initialises from what is actually running (effective), falling back to what was requested. It shows reality first, intent second.
  • Apply strategy is enabled only when your draft differs from the applied strategy, or when you have explicitly picked a metric profile. Otherwise there is nothing to publish. Reset discards the draft and returns the picker to the applied state.
If the org has no strategies at all, the card replaces the controls with an empty state: routing data could not be loaded, check that inference routing is enabled for the org, then retry.

Read-only deployments and operator-only actions

Two different restrictions apply on this screen, and they have different causes. Routing writes can be disabled per deployment. runtime.routing returns writable. When it is false, every strategy button, the weights input, the metric-profile selector, and Apply strategy are disabled, the card shows Read-only on this deployment, and the apply button’s tooltip names the gate: INFERENCE_DEPLOY_ENABLED. Reads continue to work — you can still see what is in force, what was requested, and which metrics each strategy needs. Restarting a worker is not a console capability at all. The Restart worker button is permanently disabled, with the tooltip *“Restart is