engine_billing_event, the table mod_inference writes each time it
meters a turn. That is why the screen and the invoice cannot disagree.
This page defines what each metered dimension measures, why occupancy is
the billable unit, and why some rows show (unattributed).
The billing dimensions
Every metered event is a row inengine_billing_event carrying a
tenant_id, a dimension, a units value, and a timestamp. The screen
aggregates the inference.* dimensions over the selected window:
Turns served is not its own dimension. It is the count of
inference.occupancy_ms rows — one row per metered turn. So a turn always
contributes exactly one turn and some amount of occupancy, even if it
produced no output at all.
Occupancy is stored in milliseconds and displayed in minutes, rounded to
one decimal place. The underlying record keeps the millisecond value.
Why occupancy, not tokens
The billable unit is model occupancy — wall-clock time the turn held the model — not output tokens. The reason is cancellation. ClutchCall propagates caller cancellation down into the engine so that a turn the caller interrupted stops generating immediately. That work already happened: the model was loaded, scheduled, and running for however long the turn lasted. If billing counted output tokens, the turns that cancel propagation is specifically designed to cut short would be the cheapest turns on the invoice, and the cost of the GPU time they consumed would be unrecovered. Occupancy charges for the resource actually held. A turn that ran for 800 ms and was cancelled before a single output unit reached the caller is charged for those 800 ms.Tokens vs output units
The screen shows one of two output columns, depending on what the engine reported:- Tokens — when the engine reports token counts,
tokens_inandtokens_outare metered and displayed. These are upstream-reported numbers, passed through as the engine gave them. - Output units — when the engine reports no tokens,
tokens_outaggregates to0and the screen falls back toinference.units, the count of chunks relayed to the caller.
tokens_out is greater than zero, and the units
tile otherwise. The per-tenant table always shows the Units column,
so you can read relayed output for every tenant regardless.
Output units drive the analytics breakdowns. Occupancy is what the
invoice is computed from. When an engine does report tokens, the rate
card can price against them — but the occupancy record is still written
for the same turn, and it is still the unit the rollup bills.
Cancelled turns and what they cost
A cancelled turn is a normal metered turn. It appears in:- Turns served, because it wrote an
inference.occupancy_msrow, - Model occupancy, for the time it actually ran,
- Cancelled, via
inference.cancelled_turns.
Quota denial at admission
inference.quota_denied counts turns that were refused at admission —
the gateway declined to run them. A denied turn never occupied the model,
so it contributes no occupancy_ms row and does not count toward turns
served.
The Quota denied tile only renders when the total is above zero. A
non-zero value means requests were rejected rather than served; treat it
as a capacity or entitlement signal, not as usage you were charged for.
Per-tenant attribution and unattributed usage
Every metered turn is attributed to the tenant that the API key on the request resolves to. Attribution never comes from a tenant identifier named in the request path or body. A caller cannot bill another tenant by addressing it. The table groups bytenant_id and orders by occupancy, descending,
returning up to 200 tenant rows for the window.
A row rendered as (unattributed) is a real group in the data: metered
events whose tenant_id is empty. The screen shows the label rather than
a blank cell so that the usage is visible instead of silently folded into
another tenant. The occupancy, turns, and units on that row are genuine
metered usage — they simply carry no resolved tenant. Investigate
unattributed occupancy before reconciling, because it is usage the
rollup has recorded that no tenant line item will obviously explain.
Reconciling this screen with an invoice
The screen and the invoice read the same source. To reconcile:- Match the window. The screen queries a rolling window of
windowDays(default 30) ending now —ts > now() - INTERVAL <days> DAY. An invoice covers a fixed billing period. A rolling window and a calendar period will not contain the same rows. - Compare occupancy, not output. Occupancy is the billed dimension. Token and unit columns are informational unless your rate card prices tokens.
- Account for
(unattributed). Its occupancy is in the screen total but will not appear under any tenant’s line item. - Ignore denials. Quota-denied turns were never served and are not billed.
ok: false with an empty tenant list and the screen
renders the empty state. It does not fall back to partial or synthesised
figures, so an empty screen means “no data was read”, not “no usage
occurred”. Distinguish that from a genuinely idle deployment, where the
gateway has simply not served traffic yet.
Related
- Telemetry — metrics, traces, and CDRs alongside billing events
- Authentication — the API keys that attribution resolves from
- Telephony Metrics

