inference/models) is a view over the self-hosted
model registry. The registry is a desired-configuration store: each record
is a model spec that describes how ClutchCall should serve one model
on AIBrix. Writing a spec does not start a pod by itself — the serving
fabric picks the spec up on its next hydrate.
Four procedures back the screen:
What the registry holds
Each registry record carries the deploy configuration and nothing that requires live Kubernetes or engine truth. The screen renders these fields:
The stat strip above the table counts models, sums requested GPUs and
replicas, and lists the distinct engines present. Rescan registry
re-runs
models.local. The search box filters client-side over
name, served/weights path, engine and namespace; the stat strip
always reflects the full list, not the filtered one.
The model spec field by field
The Deploy model drawer builds the payload sent tomodels.deploy.
Validation errors stay hidden until the first Deploy click, then
update live. Leaving
gpu, replicas or port blank is allowed — the
payload then falls back to the deployer’s defaults shown above — but a
non-empty value must parse as a usable number.
If models.deploy rejects the write, the drawer renders the mutation’s
error message inline and stays open. On success the drawer closes and
the registry list refetches.
Engines and serving images
The registry serves four engines: vLLM, SGLang, whisper and kokoro.engine and image are independent fields — the spec does
not derive one from the other, so pair them yourself. The form defaults
to engine: vllm with image: vllm/vllm-openai:latest; an empty engine
field falls back to vllm.
Engine values are rendered as colour-coded tags in the table, the drawer
and the “Engines” stat card.
Sourcing weights from HuggingFace
The drawer’s Source toggle picks between two sourcing modes:- HuggingFace Hub — vLLM pulls the repo directly from
huggingface.co.modelis the repo id inorg/modelform, for examplemeta-llama/Llama-3.1-8B-Instruct. - Custom path — the model is served from a weights path already
present on the node or baked into the image, for example
/models/llama-3-8b.
hfToken and revision only appear, and only travel in the payload,
when sourcing from the Hub. Set revision to pin a Hub commit or branch;
leave it blank to track the repo’s default. The table shows the pinned
revision per model, or — when nothing is pinned.
Tenant sealing for HuggingFace tokens
Supply an HF token for gated or private repos. A token is sealed per tenant, so the drawer requires a tenant / org id as soon as the token field is non-empty; submitting without one raises “Enter a tenant to seal the HF token under”. The sealing path:- The token is submitted with
tenantIdtomodels.deploy. - The BFF seals it with
cred_seal(AES-256-GCM, keyed per tenant) before the value reaches Redis. - The sealed value is decrypted only inside the serving pod, at deploy time.
sealed creds tag in its drawer
(driven by a non-empty sealedEnv). The token itself is never read back
into the console.
Namespaces, GPUs, replicas and ports
namespaceselects the Kubernetes namespace the AIBrix deployment lands in; blank meansdefault.gpuis the GPU request per replica, accepted in the range 0–8. Zero is valid, for CPU-served engines.replicasis the desired replica count, accepted in the range 1–64.portis the service port exposed by the serving container; blank means8000.
gpu and replicas across every registered model,
which is the quickest read on how much capacity the registry is asking
for in aggregate.
Status and what the registry does not track
Each row and drawer shows astatus pill derived from the spec’s
status field. This value comes from the registry record, so it tracks
the spec’s own lifecycle — it is not a live readback from Kubernetes or
from the engine.
The registry deliberately carries no live serving truth. Ready replica
counts and throughput are not part of a registry record and are not
displayed; the drawer states that they appear only once the deployment is
live and serving traffic. To reason about a running deployment, use the
serving-side signals, not this screen.
When a registry write takes effect
models.deploy saves desired configuration. It does not restart or
reload anything on its own:
- The spec is written into the registry that
mod_inferencereads on its next hydrate. - The change applies on the next deployment restart.
- Running requests are not reloaded. In-flight traffic continues against the previously loaded configuration.
image, revision, gpu,
replicas or port: the registry reflects them immediately, the serving
fabric does not until it next hydrates.
Enabling and disabling registry writes
Writes are gated on the BFF by theINFERENCE_DEPLOY_ENABLED environment
variable, set to 1 to enable them. The console reads the resulting flag
through models.deployStatus, which returns { enabled }.
When writes are disabled:
- A warning banner appears under the page header: “Registry writes are disabled on this deployment — models can be viewed but not deployed or removed.”
- Deploy model is disabled, with the tooltip “Registry writes are disabled on this deployment”.
- Every Remove button, in both the table and the drawer, is disabled.
models.localstill runs, so the registry remains fully readable and Rescan registry still works.
Removing a model
models.remove takes a single argument, { name }, and deletes the
matching registry record. Two controls call it:
- The Remove button on each table row. Its click handler stops propagation, so it removes the model without opening the drawer.
- Remove from registry in the model drawer’s footer.
Related
- Telemetry — metrics and traces for serving traffic
- Authentication — API keys and per-tenant credential scope

