The Usage & billing screen reads the same records the invoice rollup bills from. It does not recompute a bill from analytics data — it queries engine_billing_event, the table mod_inference writes each time it meters a turn. That is why the screen and the invoice cannot disagree. This page defines what each metered dimension measures, why occupancy is the billable unit, and why some rows show (unattributed).

The billing dimensions

Every metered event is a row in engine_billing_event carrying a tenant_id, a dimension, a units value, and a timestamp. The screen aggregates the inference.* dimensions over the selected window: Turns served is not its own dimension. It is the count of inference.occupancy_ms rows — one row per metered turn. So a turn always contributes exactly one turn and some amount of occupancy, even if it produced no output at all. Occupancy is stored in milliseconds and displayed in minutes, rounded to one decimal place. The underlying record keeps the millisecond value.

Why occupancy, not tokens

The billable unit is model occupancy — wall-clock time the turn held the model — not output tokens. The reason is cancellation. ClutchCall propagates caller cancellation down into the engine so that a turn the caller interrupted stops generating immediately. That work already happened: the model was loaded, scheduled, and running for however long the turn lasted. If billing counted output tokens, the turns that cancel propagation is specifically designed to cut short would be the cheapest turns on the invoice, and the cost of the GPU time they consumed would be unrecovered. Occupancy charges for the resource actually held. A turn that ran for 800 ms and was cancelled before a single output unit reached the caller is charged for those 800 ms.

Tokens vs output units

The screen shows one of two output columns, depending on what the engine reported:
  • Tokens — when the engine reports token counts, tokens_in and tokens_out are metered and displayed. These are upstream-reported numbers, passed through as the engine gave them.
  • Output units — when the engine reports no tokens, tokens_out aggregates to 0 and the screen falls back to inference.units, the count of chunks relayed to the caller.
Which one you see is decided per window, not per row: the screen shows token tiles when total tokens_out is greater than zero, and the units tile otherwise. The per-tenant table always shows the Units column, so you can read relayed output for every tenant regardless. Output units drive the analytics breakdowns. Occupancy is what the invoice is computed from. When an engine does report tokens, the rate card can price against them — but the occupancy record is still written for the same turn, and it is still the unit the rollup bills.

Cancelled turns and what they cost

A cancelled turn is a normal metered turn. It appears in:
  • Turns served, because it wrote an inference.occupancy_ms row,
  • Model occupancy, for the time it actually ran,
  • Cancelled, via inference.cancelled_turns.
Cancelled is therefore a subset of turns, not a separate class of usage, and it is not an additional charge. Occupancy already reflects the shortened run. If you cancel aggressively, you should see cancelled rise while average occupancy per turn falls — that is the intended shape. Output for a cancelled turn is whatever reached the caller before the cancel landed, which may be nothing.

Quota denial at admission

inference.quota_denied counts turns that were refused at admission — the gateway declined to run them. A denied turn never occupied the model, so it contributes no occupancy_ms row and does not count toward turns served. The Quota denied tile only renders when the total is above zero. A non-zero value means requests were rejected rather than served; treat it as a capacity or entitlement signal, not as usage you were charged for.

Per-tenant attribution and unattributed usage

Every metered turn is attributed to the tenant that the API key on the request resolves to. Attribution never comes from a tenant identifier named in the request path or body. A caller cannot bill another tenant by addressing it. The table groups by tenant_id and orders by occupancy, descending, returning up to 200 tenant rows for the window. A row rendered as (unattributed) is a real group in the data: metered events whose tenant_id is empty. The screen shows the label rather than a blank cell so that the usage is visible instead of silently folded into another tenant. The occupancy, turns, and units on that row are genuine metered usage — they simply carry no resolved tenant. Investigate unattributed occupancy before reconciling, because it is usage the rollup has recorded that no tenant line item will obviously explain.

Reconciling this screen with an invoice

The screen and the invoice read the same source. To reconcile:
  1. Match the window. The screen queries a rolling window of windowDays (default 30) ending now — ts > now() - INTERVAL <days> DAY. An invoice covers a fixed billing period. A rolling window and a calendar period will not contain the same rows.
  2. Compare occupancy, not output. Occupancy is the billed dimension. Token and unit columns are informational unless your rate card prices tokens.
  3. Account for (unattributed). Its occupancy is in the screen total but will not appear under any tenant’s line item.
  4. Ignore denials. Quota-denied turns were never served and are not billed.
If the query fails — ClickHouse unreachable, or the table absent — the procedure returns ok: false with an empty tenant list and the screen renders the empty state. It does not fall back to partial or synthesised figures, so an empty screen means “no data was read”, not “no usage occurred”. Distinguish that from a genuinely idle deployment, where the gateway has simply not served traffic yet.