# Inference — Analytics

> How the Analytics screen derives gateway occupancy, abort reasons, GPU savings, and per-turn stage latencies — and why each panel has its own window.

The Analytics screen answers one question: what did a turn cost, and what
did you get back when the turn was abandoned. It does that with two
independent panels, each backed by a different table, each with its own
fixed time window. Reading a number off the wrong panel is the most common
mistake here, so start with where the numbers come from.

## The two panels and their two data sources

| Panel | Procedure | Table | Window |
| ----- | --------- | ----- | ------ |
| Inference gateway (requests, occupancy, aborts) | `inference.gatewayStats` | `agent_turn_sample`'s sibling billing table — `engine_billing_event` | 7 days |
| Turn analytics (GPU saved, stage latencies) | `inference.turnAnalytics` | `agent_turn_sample` | 30 days |

The gateway panel is a **billing-meter rollup**. `mod_inference` calls
`meter()` for every completion it serves, which writes rows into the
billing stream. Each row is a `(dimension, units)` pair; the panel
aggregates those rows by dimension name.

The turn panel is **per-turn telemetry**. `mod_inference` /
`mod_agentruntime` emit one row per agent turn into
`agent_turn_sample`, carrying the full stage breakdown for that turn.
Benchmarks can write the same shape by posting to
`/v0/events?name=agent_turn_sample`.

Because the two panels read different tables written by different code
paths, they can disagree in coverage. That is expected, not a bug — see
[Fixed windows and why one panel can be empty](#fixed-windows-and-why-one-panel-can-be-empty).

## Gateway meters: requests, occupancy, units

Every KPI in the top panel is a sum or average over rows whose
`dimension` starts with `inference.`:

| KPI | Derived from | Meaning |
| --- | ------------ | ------- |
| **Requests** | count of rows with `dimension = 'inference.occupancy_ms'` | One row per completion served through the ClutchCall gateway's vLLM. The occupancy dimension is emitted once per request, so counting it counts requests. |
| **Avg occupancy** | average `units` on `inference.occupancy_ms` | GPU wall time held per turn, in milliseconds. This is *held* time, not generation time — a turn that holds the model while waiting still accrues occupancy. |
| **Tokens** | sum of `units` on `inference.units` | Billable units served. |
| **Abort rate** | `cancelled / requests`, computed in the procedure | Share of requests that were cancelled before they finished. |
| **Cancelled total** | sum of `units` on `inference.cancelled_turns` | Turn count, not milliseconds. |

`Requests` being defined as "rows carrying an occupancy meter" is worth
remembering when you reconcile against your own logs: a completion that
never reached the point of holding the model does not emit
`inference.occupancy_ms` and therefore is not counted as a request here.

Occupancy is also summed (`occupancy_ms_total`) by the procedure, which is
the figure to use if you want total GPU-time held across the window rather
than a per-turn average.

## The abort breakdown: superseded, caller, deadline

An abort is a turn the gateway stopped generating on purpose. The gateway
meters a separate dimension per reason, so the three counters are disjoint
and sum toward `inference.cancelled_turns`:

| KPI | Dimension | What happened |
| --- | --------- | ------------- |
| **Abort · superseded** | `inference.cancel.superseded` | A newer turn replaced this one. The conversation moved on while the model was still producing an answer to an older state, so the older generation was dropped. |
| **Abort · caller** | `inference.cancel.caller` | The caller barged in. The human started talking over the agent, so the in-flight answer is no longer wanted. |
| **Abort · deadline** | `inference.cancel.deadline` | The answer would have arrived too late to be useful. The gateway gave up rather than pay for output that would land after its usefulness window. |

These three are diagnostic, not interchangeable:

- A high **caller** share is a turn-detection / pacing signal: the agent is
  speaking when the human still wants the floor.
- A high **superseded** share means turns are being issued faster than they
  complete — state churns under the generation.
- A high **deadline** share is a latency signal: generation is slow enough,
  or started late enough, that answers stop being timely.

Each abort is also a saving, which is what the next section quantifies.

## How GPU saved and tokens reclaimed are derived

There are two different savings figures on this screen, computed two
different ways. They are not the same number and should not be compared
directly.

**Gateway panel — `Abort rate` as a proxy.** The gateway panel has no
per-turn detail, so it derives a rate, not a saving:

```
cancel_rate_pct = round( cancelled / requests * 100, 1 )
```

`cancelled` is the summed `inference.cancelled_turns` units and `requests`
is the occupancy-row count. The hint on the card reads "turns cancelled
early → GPU reclaimed" because a cancelled turn stops holding the model
early; the share of cancelled turns is a stand-in for reclaimed GPU time.
An exact milliseconds-saved figure is not available from billing meters
alone — it needs per-turn rows.

**Turn panel — measured per turn.** `agent_turn_sample` carries
`gpu_saved_pct` and `tokens_avoided` on each row, so the savings are
measured rather than inferred:

| KPI | Expression | Notes |
| --- | ---------- | ----- |
| **Abort GPU saved** | `avgIf(gpu_saved_pct, aborted)` | Averaged **only over aborted turns**. Non-aborted turns are excluded from the denominator, so this reads as "when a turn was aborted, this fraction of the generation was avoided" — not as a fleet-wide saving. |
| **Tokens reclaimed** | `sum(tokens_avoided)` | Tokens that were not generated because the turn stopped. Summed across the whole window, aborted and not. |
| **Abort rate** | `100 * countIf(aborted) / greatest(count(), 1)` | Share of sampled turns flagged `aborted`. The `greatest(…, 1)` guard keeps the expression defined when the window has no rows. |

Two consequences follow from the `avgIf` shape. First, **Abort GPU saved**
does not move when abort volume moves — it describes the average abort, not
how many there were. Read it next to **Abort rate**. Second, with very few
aborted turns in the window, it is an average over a very small sample and
will look jumpy.

## Per-turn stage latencies: EOU, LLM, TTS first audio

The second turn card breaks a turn into the stages you can act on
separately. All four come from `agent_turn_sample`:

| KPI | Column | Aggregate | Meaning |
| --- | ------ | --------- | ------- |
| **EOU detect** | `eou_ms` | `avg` | End-of-utterance detection latency — how long after the human stopped that the turn was declared finished. This is turn-detection time, before any model work starts. |
| **EOU confidence** | `eou_prob` | `avg`, 3 decimals | The turn-detection model's average completion probability for the utterance it cut on. Low average confidence with a healthy `eou_ms` means cuts are being made on weak evidence. |
| **LLM turn** | `llm_ms` | `avg` | Model time for the turn. |
| **TTS first audio** | `tts_ttfb_ms` | `avg` | Time to first audio byte out of speech synthesis. This is what the caller actually waits through after the model commits. |

`tokens_per_s_avg` (shown as **Throughput** on the first turn card) is the
same shape.

Note the `nullIf(x, 0)` wrapper on `tokens_per_s`, `llm_ms` and
`tts_ttfb_ms`: a zero in those columns is treated as *absent* and dropped
from the average rather than pulling it down. A turn that never reached the
LLM stage, or never produced audio, does not depress the stage averages for
turns that did. `eou_ms` and `eou_prob` are **not** null-guarded — every
sampled turn is expected to carry them.

Aborted turns still appear in these averages. An aborted turn has a real
EOU time and a real (truncated) LLM time, so a window with many aborts will
show shorter `llm_ms_avg` than a window without.

## Fixed windows and why one panel can be empty

Both windows are fixed by the screen, not chosen by you. The screen calls
`turnAnalytics({ windowDays: 30 })` and `gatewayStats({ windowDays: 7 })`.
Each procedure accepts a `windowDays` input and clamps it to the 1–90 range,
so other windows are reachable via the API, but the panels as rendered are
30 days and 7 days respectively.

The split is deliberate:

- The **gateway panel** is a live operational view. It lights up as soon as
  a call runs through the ClutchCall inference gateway, because billing
  meters are emitted by the gateway itself on every completion. A short
  window keeps it reflective of current behaviour.
- The **turn panel** needs per-turn telemetry, which is sampled. A longer
  window gives the averages and the aborted-turn subset enough rows to mean
  something.

### Four states you can see

The screen gates each panel on its own row count — `hasGw` requires
`requests > 0`, `hasData` requires `turns > 0`:

| State | What it tells you |
| ----- | ----------------- |
| Gateway panel only | Traffic is going through the ClutchCall gateway and being metered, but per-turn rows do not exist yet. This is the normal state before turn telemetry is on. |
| Turn panel only | `agent_turn_sample` has rows — from the agent runtime or from benchmarks posting to `/v0/events?name=agent_turn_sample` — while the gateway's billing rows for the last 7 days are absent. |
| Both panels | Full picture: metered gateway usage plus measured per-turn detail. |
| Neither panel | The empty state: "No usage data yet." |

The panels also differ in how they treat non-traffic failures. Both
procedures wrap their query in a `try` and return `ok: false` with a null
`agg` if the table does not exist yet or ClickHouse is unreachable. The
screen reads `agg` only when `ok` is true, so an unavailable table renders
as an honest empty panel rather than an error. **An empty panel therefore
does not distinguish "no traffic" from "cannot read the table."** If a
panel stays empty on a deployment you know is busy, check that its table
exists and is reachable before concluding there was no traffic.

One historical instance of exactly that: the gateway panel originally read
the metric table, but `mod_inference` emits these dimensions through
`meter()` (the billing stream) and never through `metric()`. The query
matched zero rows and the panel rendered empty on a busy gateway. The
procedure now reads the billing-event table.

## Related

- [Telemetry](/platform/telemetry) — the metric, trace, and CDR streams the gateway emits
- [Telephony Metrics](/glossary/metrics) — standard definitions for the call-level numbers
