# Metering & Billing

> How the gateway meters inference usage: occupancy, turns, tokens, cancelled turns, quota denials, and per-tenant attribution.

The **Usage & billing** screen reads the same records the invoice rollup
bills from. It does not recompute a bill from analytics data — it queries
`engine_billing_event`, the table `mod_inference` writes each time it
meters a turn. That is why the screen and the invoice cannot disagree.

This page defines what each metered dimension measures, why occupancy is
the billable unit, and why some rows show `(unattributed)`.

## The billing dimensions

Every metered event is a row in `engine_billing_event` carrying a
`tenant_id`, a `dimension`, a `units` value, and a timestamp. The screen
aggregates the `inference.*` dimensions over the selected window:

| Dimension                    | What `units` holds                                     | Shown as              |
| ---------------------------- | ------------------------------------------------------ | --------------------- |
| `inference.occupancy_ms`     | Milliseconds this turn held the model.                 | Model occupancy (min) |
| `inference.units`            | Output units relayed to the caller.                    | Output units          |
| `inference.tokens_in`        | Input tokens, as reported by the engine.               | Input tokens          |
| `inference.tokens_out`       | Output tokens, as reported by the engine.              | Output tokens         |
| `inference.cancelled_turns`  | Turns stopped before they finished.                    | Cancelled             |
| `inference.quota_denied`     | Turns refused at admission.                            | Quota denied          |

**Turns served** is not its own dimension. It is the *count* of
`inference.occupancy_ms` rows — one row per metered turn. So a turn always
contributes exactly one turn and some amount of occupancy, even if it
produced no output at all.

Occupancy is stored in milliseconds and displayed in minutes, rounded to
one decimal place. The underlying record keeps the millisecond value.

## Why occupancy, not tokens

The billable unit is **model occupancy** — wall-clock time the turn held
the model — not output tokens.

The reason is cancellation. ClutchCall propagates caller cancellation
down into the engine so that a turn the caller interrupted stops
generating immediately. That work already happened: the model was loaded,
scheduled, and running for however long the turn lasted. If billing
counted output tokens, the turns that cancel propagation is specifically
designed to cut short would be the cheapest turns on the invoice, and the
cost of the GPU time they consumed would be unrecovered.

Occupancy charges for the resource actually held. A turn that ran for
800 ms and was cancelled before a single output unit reached the caller
is charged for those 800 ms.

## Tokens vs output units

The screen shows one of two output columns, depending on what the engine
reported:

- **Tokens** — when the engine reports token counts, `tokens_in` and
  `tokens_out` are metered and displayed. These are upstream-reported
  numbers, passed through as the engine gave them.
- **Output units** — when the engine reports no tokens, `tokens_out`
  aggregates to `0` and the screen falls back to `inference.units`, the
  count of chunks relayed to the caller.

Which one you see is decided per window, not per row: the screen shows
token tiles when total `tokens_out` is greater than zero, and the units
tile otherwise. The per-tenant table always shows the **Units** column,
so you can read relayed output for every tenant regardless.

Output units drive the analytics breakdowns. Occupancy is what the
invoice is computed from. When an engine does report tokens, the rate
card can price against them — but the occupancy record is still written
for the same turn, and it is still the unit the rollup bills.

## Cancelled turns and what they cost

A cancelled turn is a normal metered turn. It appears in:

- **Turns served**, because it wrote an `inference.occupancy_ms` row,
- **Model occupancy**, for the time it actually ran,
- **Cancelled**, via `inference.cancelled_turns`.

Cancelled is therefore a subset of turns, not a separate class of usage,
and it is **not** an additional charge. Occupancy already reflects the
shortened run. If you cancel aggressively, you should see cancelled rise
while average occupancy per turn falls — that is the intended shape.

Output for a cancelled turn is whatever reached the caller before the
cancel landed, which may be nothing.

## Quota denial at admission

`inference.quota_denied` counts turns that were **refused at admission** —
the gateway declined to run them. A denied turn never occupied the model,
so it contributes no `occupancy_ms` row and does not count toward turns
served.

The **Quota denied** tile only renders when the total is above zero. A
non-zero value means requests were rejected rather than served; treat it
as a capacity or entitlement signal, not as usage you were charged for.

## Per-tenant attribution and unattributed usage

Every metered turn is attributed to the tenant that the **API key on the
request resolves to**. Attribution never comes from a tenant identifier
named in the request path or body. A caller cannot bill another tenant by
addressing it.

The table groups by `tenant_id` and orders by occupancy, descending,
returning up to 200 tenant rows for the window.

A row rendered as `(unattributed)` is a real group in the data: metered
events whose `tenant_id` is empty. The screen shows the label rather than
a blank cell so that the usage is visible instead of silently folded into
another tenant. The occupancy, turns, and units on that row are genuine
metered usage — they simply carry no resolved tenant. Investigate
unattributed occupancy before reconciling, because it is usage the
rollup has recorded that no tenant line item will obviously explain.

## Reconciling this screen with an invoice

The screen and the invoice read the same source. To reconcile:

1. **Match the window.** The screen queries a rolling window of
   `windowDays` (default 30) ending now — `ts > now() - INTERVAL <days> DAY`.
   An invoice covers a fixed billing period. A rolling window and a
   calendar period will not contain the same rows.
2. **Compare occupancy, not output.** Occupancy is the billed dimension.
   Token and unit columns are informational unless your rate card prices
   tokens.
3. **Account for `(unattributed)`.** Its occupancy is in the screen total
   but will not appear under any tenant's line item.
4. **Ignore denials.** Quota-denied turns were never served and are not
   billed.

If the query fails — ClickHouse unreachable, or the table absent — the
procedure returns `ok: false` with an empty tenant list and the screen
renders the empty state. It does not fall back to partial or synthesised
figures, so an empty screen means "no data was read", not "no usage
occurred". Distinguish that from a genuinely idle deployment, where the
gateway has simply not served traffic yet.

## Related

- [Telemetry](/platform/telemetry) — metrics, traces, and CDRs alongside billing events
- [Authentication](/concepts/authentication) — the API keys that attribution resolves from
- [Telephony Metrics](/glossary/metrics)
