# Self-hosted model registry

> What the model registry stores, how a model spec maps to an AIBrix deployment, and when a registry write reaches the serving fabric.

The **Models** screen (`inference/models`) is a view over the self-hosted
model registry. The registry is a desired-configuration store: each record
is a **model spec** that describes how ClutchCall should serve one model
on AIBrix. Writing a spec does not start a pod by itself — the serving
fabric picks the spec up on its next hydrate.

Four procedures back the screen:

| Procedure              | Kind     | Purpose                                        |
| ---------------------- | -------- | ---------------------------------------------- |
| `models.local`         | query    | Lists every model spec in the registry.        |
| `models.deployStatus`  | query    | Returns whether registry writes are enabled.   |
| `models.deploy`        | mutation | Writes a model spec.                           |
| `models.remove`        | mutation | Deletes a spec by `name`.                      |

## What the registry holds

Each registry record carries the deploy configuration and nothing that
requires live Kubernetes or engine truth. The screen renders these fields:

| Field       | Shown as                              | Notes                                                   |
| ----------- | ------------------------------------- | ------------------------------------------------------- |
| `name`      | Table row, drawer title               | Registry key. `models.remove` addresses a model by it.   |
| `served`    | Sub-line, drawer "Model" block        | The HuggingFace repo id or weights path being served.    |
| `image`     | Drawer deploy config                  | Serving container image.                                 |
| `engine`    | Engine tag                            | Colour-coded by `backendTone(engine)`.                   |
| `namespace` | Table column                          | Kubernetes namespace the deployment targets.             |
| `gpu`       | Table column, "GPUs requested" stat   | Summed across all models in the stat strip.              |
| `replicas`  | Table column, "Replicas" stat         | Summed across all models in the stat strip.              |
| `revision`  | Table column, drawer "Pinned revision"| Renders `—` when unset.                                  |
| `port`      | Drawer deploy config                  | Service port for the serving container.                  |
| `status`    | Status pill                           | See [Status and what the registry does not track](#status-and-what-the-registry-does-not-track). |
| `caps`      | Drawer tags                           | Capability tags on the spec.                             |
| `sealedEnv` | `sealed creds` tag in the drawer      | Non-empty when the spec carries sealed credentials.      |

The stat strip above the table counts models, sums requested GPUs and
replicas, and lists the distinct engines present. **Rescan registry**
re-runs `models.local`. The search box filters client-side over
`name`, `served`/weights path, `engine` and `namespace`; the stat strip
always reflects the full list, not the filtered one.

## The model spec field by field

The **Deploy model** drawer builds the payload sent to `models.deploy`.

| Field       | Required | Validation                                                                 | Sent as                       |
| ----------- | -------- | -------------------------------------------------------------------------- | ----------------------------- |
| `name`      | yes      | DNS-1123 label: `^[a-z0-9]([a-z0-9-]*[a-z0-9])?$`. Used for the service and label. | trimmed string           |
| `model`     | yes      | With HuggingFace sourcing, must look like `org/model` (no spaces, one slash). With a custom path, any non-empty value. | trimmed string |
| `image`     | yes      | Non-empty.                                                                  | trimmed string                |
| `engine`    | no       | —                                                                           | trimmed value, or `vllm`      |
| `namespace` | no       | —                                                                           | trimmed value, or `default`   |
| `gpu`       | no       | If set, an integer between 0 and 8.                                         | `Number(gpu)`, or `0`         |
| `replicas`  | no       | If set, an integer between 1 and 64.                                        | `Number(replicas)`, or `0`    |
| `port`      | no       | If set, a valid port number.                                                | `Number(port)`, or `8000`     |
| `tenantId`  | conditional | Required when an HF token is supplied.                                   | omitted when blank            |
| `hfToken`   | no       | —                                                                           | omitted unless sourcing from the Hub |
| `revision`  | no       | —                                                                           | omitted unless sourcing from the Hub |

Validation errors stay hidden until the first **Deploy** click, then
update live. Leaving `gpu`, `replicas` or `port` blank is allowed — the
payload then falls back to the deployer's defaults shown above — but a
non-empty value must parse as a usable number.

If `models.deploy` rejects the write, the drawer renders the mutation's
error message inline and stays open. On success the drawer closes and
the registry list refetches.

## Engines and serving images

The registry serves four engines: **vLLM**, **SGLang**, **whisper** and
**kokoro**. `engine` and `image` are independent fields — the spec does
not derive one from the other, so pair them yourself. The form defaults
to `engine: vllm` with `image: vllm/vllm-openai:latest`; an empty engine
field falls back to `vllm`.

Engine values are rendered as colour-coded tags in the table, the drawer
and the "Engines" stat card.

## Sourcing weights from HuggingFace

The drawer's **Source** toggle picks between two sourcing modes:

- **HuggingFace Hub** — vLLM pulls the repo directly from
  `huggingface.co`. `model` is the repo id in `org/model` form, for
  example `meta-llama/Llama-3.1-8B-Instruct`.
- **Custom path** — the model is served from a weights path already
  present on the node or baked into the image, for example
  `/models/llama-3-8b`.

`hfToken` and `revision` only appear, and only travel in the payload,
when sourcing from the Hub. Set `revision` to pin a Hub commit or branch;
leave it blank to track the repo's default. The table shows the pinned
revision per model, or `—` when nothing is pinned.

## Tenant sealing for HuggingFace tokens

Supply an HF token for gated or private repos. A token is **sealed
per tenant**, so the drawer requires a tenant / org id as soon as the
token field is non-empty; submitting without one raises
*"Enter a tenant to seal the HF token under"*.

The sealing path:

1. The token is submitted with `tenantId` to `models.deploy`.
2. The BFF seals it with `cred_seal` (AES-256-GCM, keyed per tenant)
   **before** the value reaches Redis.
3. The sealed value is decrypted only inside the serving pod, at deploy
   time.

A spec with sealed credentials shows a `sealed creds` tag in its drawer
(driven by a non-empty `sealedEnv`). The token itself is never read back
into the console.

## Namespaces, GPUs, replicas and ports

- `namespace` selects the Kubernetes namespace the AIBrix deployment
  lands in; blank means `default`.
- `gpu` is the GPU **request** per replica, accepted in the range 0–8.
  Zero is valid, for CPU-served engines.
- `replicas` is the desired replica count, accepted in the range 1–64.
- `port` is the service port exposed by the serving container; blank
  means `8000`.

The stat strip totals `gpu` and `replicas` across every registered model,
which is the quickest read on how much capacity the registry is asking
for in aggregate.

## Status and what the registry does not track

Each row and drawer shows a `status` pill derived from the spec's
`status` field. This value comes from the registry record, so it tracks
the spec's own lifecycle — it is not a live readback from Kubernetes or
from the engine.

The registry deliberately carries no live serving truth. Ready replica
counts and throughput are **not** part of a registry record and are not
displayed; the drawer states that they appear only once the deployment is
live and serving traffic. To reason about a running deployment, use the
serving-side signals, not this screen.

## When a registry write takes effect

`models.deploy` saves desired configuration. It does not restart or
reload anything on its own:

- The spec is written into the registry that `mod_inference` reads on its
  **next hydrate**.
- The change applies on the **next deployment restart**.
- **Running requests are not reloaded.** In-flight traffic continues
  against the previously loaded configuration.

The same applies to edits that change `image`, `revision`, `gpu`,
`replicas` or `port`: the registry reflects them immediately, the serving
fabric does not until it next hydrates.

## Enabling and disabling registry writes

Writes are gated on the BFF by the `INFERENCE_DEPLOY_ENABLED` environment
variable, set to `1` to enable them. The console reads the resulting flag
through `models.deployStatus`, which returns `{ enabled }`.

When writes are disabled:

- A warning banner appears under the page header: *"Registry writes are
  disabled on this deployment — models can be viewed but not deployed or
  removed."*
- **Deploy model** is disabled, with the tooltip *"Registry writes are
  disabled on this deployment"*.
- Every **Remove** button, in both the table and the drawer, is disabled.
- `models.local` still runs, so the registry remains fully readable and
  **Rescan registry** still works.

This is a per-deployment gate, not a per-user permission. Read-only
consoles pointed at a production BFF are the usual reason to leave it off.

## Removing a model

`models.remove` takes a single argument, `{ name }`, and deletes the
matching registry record. Two controls call it:

- The **Remove** button on each table row. Its click handler stops
  propagation, so it removes the model without opening the drawer.
- **Remove from registry** in the model drawer's footer.

Both fire immediately — there is no confirmation step. While the mutation
is pending the button shows a spinner and is disabled. On success the
registry list refetches and the drawer closes.

Removal edits desired configuration only, on the same timing as a deploy:
it takes the spec out of the registry, and the serving fabric reconciles
on its next hydrate.

## Related

- [Telemetry](/platform/telemetry) — metrics and traces for serving traffic
- [Authentication](/concepts/authentication) — API keys and per-tenant credential scope
