The two panels and their two data sources
The gateway panel is a billing-meter rollup.
mod_inference calls
meter() for every completion it serves, which writes rows into the
billing stream. Each row is a (dimension, units) pair; the panel
aggregates those rows by dimension name.
The turn panel is per-turn telemetry. mod_inference /
mod_agentruntime emit one row per agent turn into
agent_turn_sample, carrying the full stage breakdown for that turn.
Benchmarks can write the same shape by posting to
/v0/events?name=agent_turn_sample.
Because the two panels read different tables written by different code
paths, they can disagree in coverage. That is expected, not a bug — see
Fixed windows and why one panel can be empty.
Gateway meters: requests, occupancy, units
Every KPI in the top panel is a sum or average over rows whosedimension starts with inference.:
Requests being defined as “rows carrying an occupancy meter” is worth
remembering when you reconcile against your own logs: a completion that
never reached the point of holding the model does not emit
inference.occupancy_ms and therefore is not counted as a request here.
Occupancy is also summed (occupancy_ms_total) by the procedure, which is
the figure to use if you want total GPU-time held across the window rather
than a per-turn average.
The abort breakdown: superseded, caller, deadline
An abort is a turn the gateway stopped generating on purpose. The gateway meters a separate dimension per reason, so the three counters are disjoint and sum towardinference.cancelled_turns:
These three are diagnostic, not interchangeable:
- A high caller share is a turn-detection / pacing signal: the agent is speaking when the human still wants the floor.
- A high superseded share means turns are being issued faster than they complete — state churns under the generation.
- A high deadline share is a latency signal: generation is slow enough, or started late enough, that answers stop being timely.
How GPU saved and tokens reclaimed are derived
There are two different savings figures on this screen, computed two different ways. They are not the same number and should not be compared directly. Gateway panel —Abort rate as a proxy. The gateway panel has no
per-turn detail, so it derives a rate, not a saving:
cancelled is the summed inference.cancelled_turns units and requests
is the occupancy-row count. The hint on the card reads “turns cancelled
early → GPU reclaimed” because a cancelled turn stops holding the model
early; the share of cancelled turns is a stand-in for reclaimed GPU time.
An exact milliseconds-saved figure is not available from billing meters
alone — it needs per-turn rows.
Turn panel — measured per turn. agent_turn_sample carries
gpu_saved_pct and tokens_avoided on each row, so the savings are
measured rather than inferred:
Two consequences follow from the
avgIf shape. First, Abort GPU saved
does not move when abort volume moves — it describes the average abort, not
how many there were. Read it next to Abort rate. Second, with very few
aborted turns in the window, it is an average over a very small sample and
will look jumpy.
Per-turn stage latencies: EOU, LLM, TTS first audio
The second turn card breaks a turn into the stages you can act on separately. All four come fromagent_turn_sample:
tokens_per_s_avg (shown as Throughput on the first turn card) is the
same shape.
Note the nullIf(x, 0) wrapper on tokens_per_s, llm_ms and
tts_ttfb_ms: a zero in those columns is treated as absent and dropped
from the average rather than pulling it down. A turn that never reached the
LLM stage, or never produced audio, does not depress the stage averages for
turns that did. eou_ms and eou_prob are not null-guarded — every
sampled turn is expected to carry them.
Aborted turns still appear in these averages. An aborted turn has a real
EOU time and a real (truncated) LLM time, so a window with many aborts will
show shorter llm_ms_avg than a window without.
Fixed windows and why one panel can be empty
Both windows are fixed by the screen, not chosen by you. The screen callsturnAnalytics({ windowDays: 30 }) and gatewayStats({ windowDays: 7 }).
Each procedure accepts a windowDays input and clamps it to the 1–90 range,
so other windows are reachable via the API, but the panels as rendered are
30 days and 7 days respectively.
The split is deliberate:
- The gateway panel is a live operational view. It lights up as soon as a call runs through the ClutchCall inference gateway, because billing meters are emitted by the gateway itself on every completion. A short window keeps it reflective of current behaviour.
- The turn panel needs per-turn telemetry, which is sampled. A longer window gives the averages and the aborted-turn subset enough rows to mean something.
Four states you can see
The screen gates each panel on its own row count —hasGw requires
requests > 0, hasData requires turns > 0:
The panels also differ in how they treat non-traffic failures. Both
procedures wrap their query in a
try and return ok: false with a null
agg if the table does not exist yet or ClickHouse is unreachable. The
screen reads agg only when ok is true, so an unavailable table renders
as an honest empty panel rather than an error. An empty panel therefore
does not distinguish “no traffic” from “cannot read the table.” If a
panel stays empty on a deployment you know is busy, check that its table
exists and is reachable before concluding there was no traffic.
One historical instance of exactly that: the gateway panel originally read
the metric table, but mod_inference emits these dimensions through
meter() (the billing stream) and never through metric(). The query
matched zero rows and the panel rendered empty on a busy gateway. The
procedure now reads the billing-event table.
Related
- Telemetry — the metric, trace, and CDR streams the gateway emits
- Telephony Metrics — standard definitions for the call-level numbers

