# Streams metrics reference

> How the Stream overview numbers are computed: windows, bucket sizes, viewer sampling, glass-to-glass timestamps, cache hit ratio, and metered usage.

The Stream **Overview** screen reads from six analytics procedures. Each
one has its own window, its own bucket size, and its own refresh
interval, so two tiles on the same screen can legitimately disagree.
This page states what each number actually measures.

| Tile | Procedure | Window |
| ---- | --------- | ------ |
| KPI strip (viewers now, live streams, P50 glass-to-glass, 24h egress, cache hit) | `streams.analytics.overviewKpis` | mixed — see below |
| Live now list | `streams.liveInputs.list` | instantaneous snapshot |
| Concurrent viewers chart | `streams.analytics.viewerSeries` | 24 h |
| Glass-to-glass histogram + P50/P95/P99 | `streams.analytics.g2gHistogram`, `streams.analytics.qoeKpis` | 60 min |
| This billing cycle | `streams.analytics.usage`, `streams.analytics.billing` | current billing cycle |
| Recent events | `streams.eventDeliveries.list` | last 6 deliveries |

## Windows, buckets, and refresh intervals

The screen passes the window and bucket size explicitly, so you can
reproduce any tile by calling the same procedure with the same
arguments.

| Query | Arguments the Overview sends | Refetch |
| ----- | ---------------------------- | ------- |
| `overviewKpis` | `{ orgId }` | 30 s |
| `liveInputs.list` (live snapshot) | `{ orgId, page: 1, perPage: 100 }` | 30 s |
| `liveInputs.list` (total count only) | `{ orgId, page: 1, perPage: 1 }` | `staleTime` 60 s |
| `eventDeliveries.list` | `{ orgId, limit: 6 }` | 30 s |
| `viewerSeries` | `{ orgId, windowMinutes: 1440, bucketSeconds: 600 }` | 60 s |
| `qoeKpis` | `{ orgId, windowMinutes: 60 }` | 30 s |
| `g2gHistogram` | `{ orgId, windowMinutes: 60 }` | 60 s |
| `usage` | `{ orgId }` | 60 s |
| `billing` | `{ orgId }` | 120 s |

Two consequences worth internalising:

- **The chart is coarser than the KPI.** The viewer series uses 600-second
  buckets over 24 hours. The "Concurrent viewers" KPI is a current
  reading from `overviewKpis`. A spike shorter than a bucket shows up in
  the KPI and is flattened in the chart.
- **The refresh intervals are staggered.** The glass-to-glass KPI in the
  strip refreshes on the 30 s `overviewKpis` cycle; the histogram beside
  it refreshes on its own 60 s cycle. The P50 in the strip and the P50
  under the histogram come from *different queries* (`overviewKpis` vs
  `qoeKpis`) and can be up to one refresh interval apart.

`viewerSeries` returns one row per bucket with `bucket`, `viewers`, and
`g2g_p50` — i.e. the series carries a per-bucket latency median
alongside the viewer count, even though the Overview chart only plots
`viewers`.

## Viewer counts: concurrent, unique, and how they are sampled

Two different viewer numbers appear on the screen.

**Org-wide concurrent viewers** comes from `overviewKpis.viewers_now`.
It is a current reading, not an average over the refresh interval: the
tile re-reads it every 30 s and replaces the displayed value.

**Per-stream concurrent viewers** comes from each live input row's
`viewers_now` field in `liveInputs.list`, rendered in the "Live now"
list next to that stream's `g2g_p50_ms`. A row with no reading is
displayed as `0`.

The 24-hour chart plots `viewers` per 600-second bucket from
`viewerSeries`. Because it is a bucketed rollup, each plotted point is
one value for a ten-minute span, not a continuous trace. The chart
renders **"No viewer samples yet."** rather than a flat line whenever
fewer than two buckets have arrived, or when every bucket in the window
is zero — a relay rollup that has not produced two buckets yet cannot be
drawn as a line.

## Live stream counts and the trailing "+"

The "Live streams" KPI and the "across N live streams" hint both come
from client-side filtering, and both can be a floor rather than an exact
count.

The Overview fetches one page of inputs (`perPage: 100`) **without a
status filter**, then keeps the rows where `is_live` is true or `status`
is `active`. The status filter is deliberately omitted because the
relay-derived `is_live` reconcile happens server-side *after* the
Postgres status filter — filtering in the query would drop rows the
reconcile would have marked live.

That client-side filter can only see the rows on the page it fetched. So
when `total` from the response exceeds the number of rows returned, the
count is a lower bound and the console appends a `+`:

```
Live streams   7+ / 213
```

Read `7+` as "at least 7 of the inputs on the first page are live; there
are more inputs than were inspected." The denominator (`/ 213`) is the
org's total input count, taken from a separate `perPage: 1` query that
exists only to read `total`. If you need an exact live count for an org
with many inputs, page through `streams.liveInputs.list` yourself and
apply the same `is_live || status === 'active'` predicate across all
pages.

## Glass-to-glass: the target line and the percentiles

Glass-to-glass latency is surfaced three ways:

- `overviewKpis.g2g_p50_ms` — the single number in the KPI strip.
- `qoeKpis.g2g_p50_ms` / `g2g_p95_ms` / `g2g_p99_ms` — percentiles over
  the trailing 60 minutes, shown under the histogram.
- `g2gHistogram.buckets` — the distribution over the same 60-minute
  window, drawn as the bar chart. The Overview renders the first 14
  buckets, with axis labels at `80`, `180`, `280`, `420`, and `600+`
  milliseconds, and places a marker at the bucket derived from the
  current P50.

The percentiles are labelled **all PoPs** — they are aggregated across
points of presence, not per-PoP. A regional problem shows up in P95/P99
while P50 stays healthy.

The `250ms` shown on the KPI card is a **display target**, not a quota or
a service commitment. The card renders in its "ok" state when P50 is
greater than zero and at or below that target; above it, the card simply
loses the ok styling. Nothing is throttled, rejected, or billed
differently on either side of the line.

A reported value of `0` means "no reading", not "zero latency" — every
tile treats `g2g_p50_ms > 0` as the precondition for rendering a value,
and shows `—` otherwise.

## Egress and cache hit

`overviewKpis` carries two delivery-side numbers:

| Field | Tile | Notes |
| ----- | ---- | ----- |
| `bytes_egress_24h` | Egress · 24h | Raw bytes. The console divides by 1024⁴ and prints two decimals as TB. |
| `cache_hit_ratio` | Cache hit · edge | Already a percentage; printed to one decimal. |

Note the window mismatch: egress in the KPI strip is a rolling 24-hour
figure, while the egress meter in the **This billing cycle** panel covers
the billing period. They are not the same measurement and should not be
expected to match.

The cache hit card renders its "ok" state at 95% or above. As with the
glass-to-glass target, that is a threshold for the card's colour only.

## Metered usage vs plan pools

The billing panel draws three meters, and each one prefers the plan pool
over raw metered bytes when a pool exists.

`streams.analytics.usage` returns raw meters:

| Field | Unit |
| ----- | ---- |
| `delivery_minutes` | minutes |
| `egress_bytes` | bytes |
| `storage_bytes` | bytes |

`streams.analytics.billing` returns `pools`, keyed by meter, where each
pool has a `limit` and a `remaining`. The Overview uses two pools:
`stream_egress_gb` and `vod_storage_gb`.

The resolution rule for the egress and storage meters is:

1. If the pool exists, **used = `limit − remaining`**, clamped at zero,
   and the meter's limit is the pool's `limit`.
2. If the pool is absent, fall back to the raw meter from `usage`
   (`egress_bytes` or `storage_bytes`, converted to GB), and the meter is
   drawn with no limit.

Both values are rounded to one decimal place for display.

Deriving "used" by subtraction is why a pool whose `remaining` briefly
exceeds its `limit` (for example after a credit) shows as `0` rather than
a negative number.

**Delivery minutes has no pool.** It is drawn straight from
`usage.delivery_minutes` with no limit, so it reports consumption only —
it is not a progress bar against an allowance.

## When a number is stale or incomplete

Work through these in order when a tile disagrees with what you expect:

- **`—` instead of a value.** The query has not resolved, or the
  underlying field is zero/absent. Every KPI on this screen prints `—`
  when its source is missing, and the glass-to-glass tiles additionally
  print `—` when the value is `0`.
- **"No viewer samples yet."** Fewer than two buckets in the series, or
  all buckets zero. Not an error — a stream that started less than two
  buckets ago cannot be plotted.
- **"No active inputs".** `liveInputs.list` resolved with nothing
  matching `is_live || status === 'active'`. The first encoder publish
  populates the list.
- **A trailing `+` on a count.** The count is a floor computed from one
  page. See [Live stream counts and the trailing
  "+"](#live-stream-counts-and-the-trailing).
- **Two tiles showing different latency.** Different queries, different
  refresh intervals. Compare only after both have refreshed.
- **Egress in the KPI strip ≠ egress in the billing panel.** Different
  windows (24 h vs billing cycle) and, when a pool exists, a different
  derivation (pool subtraction vs raw bytes).
- **A telemetry health banner at the top of the screen.** The Overview
  renders one above the KPIs. When it is showing, treat every number on
  the page as suspect until it clears.

Because the whole screen is polled, a value is at most one refresh
interval old: 30 s for the KPI strip, the live snapshot, the QoE
percentiles and the event list; 60 s for the viewer series, the
histogram and usage; 120 s for billing pools.

## Reproducing these numbers from the SDK

Every tile is a plain tRPC call against the same control-plane origin the
console uses, so you can pull the same figures into your own dashboard
with the ClutchCall SDK or over HTTPS. Pass the same `windowMinutes`
and `bucketSeconds` the console passes if you want the values to line up
with what the screen shows.

```ts
const kpis   = await streams.analytics.overviewKpis({ orgId });
const series = await streams.analytics.viewerSeries({
  orgId, windowMinutes: 24 * 60, bucketSeconds: 600,
});
const qoe    = await streams.analytics.qoeKpis({ orgId, windowMinutes: 60 });
const hist   = await streams.analytics.g2gHistogram({ orgId, windowMinutes: 60 });
const usage  = await streams.analytics.usage({ orgId });
const bill   = await streams.analytics.billing({ orgId });
```

## Related

- [Telemetry](/platform/telemetry) — gateway-side metric families, traces, and CDRs
- [Telephony Metrics](/glossary/metrics) — standard definitions for latency, jitter, and MOS
- [SDK methods](/modalities/streams/sdk-methods)
