The Models screen (inference/models) is a view over the self-hosted model registry. The registry is a desired-configuration store: each record is a model spec that describes how ClutchCall should serve one model on AIBrix. Writing a spec does not start a pod by itself — the serving fabric picks the spec up on its next hydrate. Four procedures back the screen:

What the registry holds

Each registry record carries the deploy configuration and nothing that requires live Kubernetes or engine truth. The screen renders these fields: The stat strip above the table counts models, sums requested GPUs and replicas, and lists the distinct engines present. Rescan registry re-runs models.local. The search box filters client-side over name, served/weights path, engine and namespace; the stat strip always reflects the full list, not the filtered one.

The model spec field by field

The Deploy model drawer builds the payload sent to models.deploy. Validation errors stay hidden until the first Deploy click, then update live. Leaving gpu, replicas or port blank is allowed — the payload then falls back to the deployer’s defaults shown above — but a non-empty value must parse as a usable number. If models.deploy rejects the write, the drawer renders the mutation’s error message inline and stays open. On success the drawer closes and the registry list refetches.

Engines and serving images

The registry serves four engines: vLLM, SGLang, whisper and kokoro. engine and image are independent fields — the spec does not derive one from the other, so pair them yourself. The form defaults to engine: vllm with image: vllm/vllm-openai:latest; an empty engine field falls back to vllm. Engine values are rendered as colour-coded tags in the table, the drawer and the “Engines” stat card.

Sourcing weights from HuggingFace

The drawer’s Source toggle picks between two sourcing modes:
  • HuggingFace Hub — vLLM pulls the repo directly from huggingface.co. model is the repo id in org/model form, for example meta-llama/Llama-3.1-8B-Instruct.
  • Custom path — the model is served from a weights path already present on the node or baked into the image, for example /models/llama-3-8b.
hfToken and revision only appear, and only travel in the payload, when sourcing from the Hub. Set revision to pin a Hub commit or branch; leave it blank to track the repo’s default. The table shows the pinned revision per model, or — when nothing is pinned.

Tenant sealing for HuggingFace tokens

Supply an HF token for gated or private repos. A token is sealed per tenant, so the drawer requires a tenant / org id as soon as the token field is non-empty; submitting without one raises “Enter a tenant to seal the HF token under”. The sealing path:
  1. The token is submitted with tenantId to models.deploy.
  2. The BFF seals it with cred_seal (AES-256-GCM, keyed per tenant) before the value reaches Redis.
  3. The sealed value is decrypted only inside the serving pod, at deploy time.
A spec with sealed credentials shows a sealed creds tag in its drawer (driven by a non-empty sealedEnv). The token itself is never read back into the console.

Namespaces, GPUs, replicas and ports

  • namespace selects the Kubernetes namespace the AIBrix deployment lands in; blank means default.
  • gpu is the GPU request per replica, accepted in the range 0–8. Zero is valid, for CPU-served engines.
  • replicas is the desired replica count, accepted in the range 1–64.
  • port is the service port exposed by the serving container; blank means 8000.
The stat strip totals gpu and replicas across every registered model, which is the quickest read on how much capacity the registry is asking for in aggregate.

Status and what the registry does not track

Each row and drawer shows a status pill derived from the spec’s status field. This value comes from the registry record, so it tracks the spec’s own lifecycle — it is not a live readback from Kubernetes or from the engine. The registry deliberately carries no live serving truth. Ready replica counts and throughput are not part of a registry record and are not displayed; the drawer states that they appear only once the deployment is live and serving traffic. To reason about a running deployment, use the serving-side signals, not this screen.

When a registry write takes effect

models.deploy saves desired configuration. It does not restart or reload anything on its own:
  • The spec is written into the registry that mod_inference reads on its next hydrate.
  • The change applies on the next deployment restart.
  • Running requests are not reloaded. In-flight traffic continues against the previously loaded configuration.
The same applies to edits that change image, revision, gpu, replicas or port: the registry reflects them immediately, the serving fabric does not until it next hydrates.

Enabling and disabling registry writes

Writes are gated on the BFF by the INFERENCE_DEPLOY_ENABLED environment variable, set to 1 to enable them. The console reads the resulting flag through models.deployStatus, which returns { enabled }. When writes are disabled:
  • A warning banner appears under the page header: “Registry writes are disabled on this deployment — models can be viewed but not deployed or removed.”
  • Deploy model is disabled, with the tooltip “Registry writes are disabled on this deployment”.
  • Every Remove button, in both the table and the drawer, is disabled.
  • models.local still runs, so the registry remains fully readable and Rescan registry still works.
This is a per-deployment gate, not a per-user permission. Read-only consoles pointed at a production BFF are the usual reason to leave it off.

Removing a model

models.remove takes a single argument, { name }, and deletes the matching registry record. Two controls call it:
  • The Remove button on each table row. Its click handler stops propagation, so it removes the model without opening the drawer.
  • Remove from registry in the model drawer’s footer.
Both fire immediately — there is no confirmation step. While the mutation is pending the button shows a spinner and is disabled. On success the registry list refetches and the drawer closes. Removal edits desired configuration only, on the same timing as a deploy: it takes the spec out of the registry, and the serving fabric reconciles on its next hydrate.