mod_inference applies when it picks a pod for a request.
Everything on the screen comes from two read procedures — runtime.health and
runtime.routing — plus one write, runtime.setRouting. Nothing on the screen
reaches into the cluster directly.
What the runtime publishes and where it is stored
The runtime does not answer console queries. It publishes a blob, and the console reads it.
Two consequences follow from this shape, and they explain most of the screen’s
behaviour:
- A field shown as
—is a field that was absent from the published blob, not a zero. The console does not substitute a default. - Freshness is the publisher’s property. The console renders what is in Redis at the moment of the query.
Host banner fields
The banner across the top of the screen is read straight from the health blob, with one exception.Reading the resource gauges
Three cards show host resources. The donut in each card is a percentage; the bars beside it are meters. GPU — device name, compute utilization (donut), then VRAM used against VRAM total, temperature, and power draw against the reported power cap. CPU — model and core count, load (donut), then 1-minute loadavg, the configured tensor-parallel width, andmax-model-len rendered in thousands of
tokens.
System memory — host RAM total, used as a share of total (donut), then the
number of memory-mapped models, page cache, and swap.
Two rendering rules matter when you are reading a number off this screen:
- The memory donut needs a total. If the blob reports no memory total, the
ratio is undefined, so the donut shows
0%and the sub-label readsnot reportedrather than a computedNaN%. - Some bars are relative, not capacity limits. For rows where the blob
carries no natural maximum — tensor-parallel width,
max-model-len, models mapped, page cache, swap, loadavg — the bar’s full scale is the larger of the reported value and a fixed display floor. The bar is there so you can see movement; read the number in the right-hand label, not the fill fraction.
vramTotal, a fixed temperature scale, powerCap), so their fill fractions
are meaningful.
Runtime config and what each vLLM argument means
The Runtime config card lists the vLLM server arguments for this runtime. The first five rows come from the health blob; the last two are rendered as fixed rows in the console.
Because the last two rows are constants in the console, they describe how the
runtime is expected to be launched — they are not evidence of what this
particular process was given.
Self-tests and how to read them
Each entry in the self-test list is a boot diagnostic that the runtime published, with four fields:
The set of tests is not fixed by the console. The screen renders whatever
tests are present in the blob, in the order published. If a check you expect is
missing, the runtime did not report it; that is a publisher-side gap, not a pass.
Use
ms as a signal in its own right. A check that passes but takes far longer
than its neighbours usually means the resource it touched was slow to respond,
which tends to show up later as latency rather than as a failure.
What “Run self-test” actually does
The Run self-test button callsrefetch() on runtime.health. It re-reads
the published runtime/self-test blob. It does not instruct the runtime to
re-run its boot diagnostics, and it does not fabricate a progress animation.
So:
- If the runtime has re-published since your last look, you will see new results.
- If it has not, you will see the same results again, and the timestamp label
still reads
last run just now— that label describes the fetch, not the diagnostic run.
Routing strategies and the metrics they require
runtime.routing returns the strategy list for the org; the console renders one
button per strategy and does not hardcode names. Selecting a button only sets a
local draft — nothing is published until you press Apply strategy.
The important field alongside the list is requires: a per-strategy list of the
scraped metrics that strategy needs in order to score at all. The console
shows it inline under the strategy picker:
- If the selected strategy requires metrics, the hint names them and warns that without them the router falls back to a random ready pod. This is the failure mode worth internalising: a KV-aware or load-aware strategy whose metrics are not being scraped does not error. It silently degrades to random selection, and your latency profile quietly becomes that of no routing at all.
- If the strategy requires no scraped metrics, the hint says so — that strategy routes from prompt affinity and live in-flight load, both of which the module observes itself.
Weighted multi-scorer routing
Themulti (weighted) option composes several scorers instead of picking one.
Its strategy string is multi: followed by comma-separated name:weight pairs.
Each scorer is normalized to [0,1] and the normalized scores are combined by
weight. The console seeds a starting spec of:
runtime.routing returns multiScorers — the set of scorer names that can be
weighted. This set is narrower than the full strategy list: not every strategy
is weightable.
The spec you type goes to the live router verbatim, so the console parses it
before publishing and shows the first problem instead of publishing a broken
spec:
Parsing splits each pair at its last colon, so the name is everything before
it and the weight everything after. Validation errors surface when you press
Apply strategy; the input turns invalid and the message appears rather than
the mutation firing.
Metric profiles per engine
The Metric profile (static routes) control selects which engine’s/metrics naming the scraper parses for pods in the static route table.
Pods discovered dynamically use their own engine label instead, so this setting
only governs statically-routed pods.
The console publishes
routesEngine only when you actually pick a profile.
The value displayed is the effective profile falling back to the desired one,
and re-publishing a displayed value would rewrite a field you never touched —
worse, it would coerce an engine string the console does not model (for example
sglang) into vllm. Touch the control and it is sent; leave it alone and
setRouting carries the strategy only.
Desired state, effective state, and the poll cycle
The routing card keeps two facts apart and never conflates them:
The effective record is a TTL’d beacon: the module republishes it on every
poll.
engineReporting is true only while that beacon is present. The status
pill reflects the combination:
When there is no beacon, the card explains that no live
mod_inference is
publishing routing state at <key>.effective, and that a strategy published
here is stored as desired state and adopted when a module starts polling.
ClutchCall deliberately does not show a green confirmed strategy in that
situation, because a strategy no engine has acknowledged may be steering nothing.
Switching strategies takes effect on the module’s next poll (~5 s). There is
no restart, and no request in flight is re-routed.
Two details of the control’s behaviour follow from this model:
- The strategy picker initialises from what is actually running (
effective), falling back to what was requested. It shows reality first, intent second. - Apply strategy is enabled only when your draft differs from the applied strategy, or when you have explicitly picked a metric profile. Otherwise there is nothing to publish. Reset discards the draft and returns the picker to the applied state.
Read-only deployments and operator-only actions
Two different restrictions apply on this screen, and they have different causes. Routing writes can be disabled per deployment.runtime.routing returns
writable. When it is false, every strategy button, the weights input, the
metric-profile selector, and Apply strategy are disabled, the card shows
Read-only on this deployment, and the apply button’s tooltip names the gate:
INFERENCE_DEPLOY_ENABLED. Reads continue to work — you can still see what is
in force, what was requested, and which metrics each strategy needs.
Restarting a worker is not a console capability at all. The Restart
worker button is permanently disabled, with the tooltip *“Restart is
