# Realtime engine telemetry

> Where the console's live counters and history charts come from: engine-written Redis keys, the wired heartbeat, the minute sampler, coverage, and why values render as em-dashes.

The realtime console shows two kinds of numbers, produced by two
different mechanisms:

| Surface | Source | Refresh |
| ------- | ------ | ------- |
| Live counters (Connections, Active channels, Messages / min) | Engine-written Redis keys, read by the ClutchCall backend | Polled every 5s |
| Channel mix, top channels, per-channel table | The engine's per-app channel registry in Redis | Polled every 5s |
| History charts, period peaks, coverage | Backend sampler rows in `public.realtime_app_stat` | Polled every 60s |

Live values are a snapshot of what the engine wrote most recently.
History values are what the sampler managed to record. The two can
disagree, and the console does not hide it — see
[Coverage, gaps, and engine-offline buckets](#coverage-gaps-and-engine-offline-buckets).

## Live counters and the engine heartbeat

`realtime.overview` and `realtime.stats` read a small set of Redis keys
for the selected app. The app is addressed by its **public `app_id`**,
not its internal row id:

| Key | Field | Meaning |
| --- | ----- | ------- |
| `clutchcall:realtime:wired` | `wired` | Engine heartbeat. `true` when the value is `1`. |
| `clutchcall:realtime:app:<appId>:conns` | `connections` | Current connections for the app. |
| `clutchcall:realtime:app:<appId>:channels` | `channels` | Currently occupied channels for the app. |
| `clutchcall:realtime:app:<appId>:msgs:1m` | `messagesPerMin` | Messages counted over the last 60s. |

All four are read in a single `MGET`. The backend first checks that the
requested app belongs to the calling organization; a mismatch is
rejected before any Redis read happens.

Missing keys are read as absent, not as errors. If the Redis read itself
throws, the procedure returns `wired: false` with all counters at zero
rather than failing the request — so a Redis problem looks like "engine
not connected" in the console, not a broken page.

## The wired heartbeat: what engine connected does and does not prove

The `engine connected` pill reflects exactly one thing: the presence of
the **global** key `clutchcall:realtime:wired` with the value
`1`. It is written by the engine (`mod_realtime`), not by your app.

What it proves:

- An engine node is running and heartbeating telemetry into Redis.

What it does **not** prove:

- That *your* app has any traffic. The heartbeat is global, not
  per-app. An engine with zero connections still reports
  `engine connected`.
- That your app's counters exist. The engine writes nothing for an idle
  app, so `connections`, `channels`, and `msgs:1m` can all be absent
  while the heartbeat is present.
- That your specific client is connected. Use the live counters and the
  per-channel table for that, not the pill.

Conversely, `engine not connected` means the heartbeat key is missing.
While it is missing, the live counters and live breakdowns are not
trustworthy, so the console suppresses them and shows an explicit
"Engine not connected" panel instead of zeros. History remains readable
in that state, because history comes from the sampler table rather than
from Redis.

## Counter windows: connections, channels, messages in the last 60s

The three live counters have different semantics, and the console
labels them accordingly:

- **Connections** — an instantaneous gauge. It is the count at the
  moment the engine last wrote the key, labelled `live now`.
- **Active channels** — also instantaneous: channels currently
  occupied, labelled `live now` on the overview and `occupied` on
  Stats.
- **Messages / min** — a **window**, not a gauge. It is the engine's
  count over the trailing 60 seconds, labelled `last 60s`. It is read
  from `…:msgs:1m`.

Because the console polls every 5s and the engine writes on its own
cadence, consecutive polls can return the same value. The
`updated <n> ago` caption on the overview strip shows when the console last
received
data, and `refreshing…` replaces it while a fetch is in flight. The
refresh button next to it forces an immediate re-read and is disabled
while a fetch is already running.

## Why a value shows an em-dash instead of zero

On the overview strip a counter renders as `—` whenever the console
cannot assert a real number. Formally, a numeric value is rendered only
when the heartbeat is present *and* the field came back defined.
Everything else renders `—`.

This is deliberate: `0` is a claim that there is no traffic. `—` is a
statement that there is no telemetry. Distinguishing them is the whole
point of the em-dash.

The sub-caption under each counter tells you which case you are in, and
takes priority over the normal `live now` / `last 60s` labels.

## Reading the four non-live states (no app, loading, backend unreachable, engine offline)

The overview strip has four distinct non-live states. Each has its own
pill and its own counter sub-caption:

| State | Pill | Sub-caption | What it means |
| ----- | ---- | ----------- | ------------- |
| No app selected | `select an app` | `select an app` | No org + `app_id` pair yet. The query is **idle** — the console is not polling at all. |
| Loading | `checking…` | `checking…` | First fetch for this app is in flight. |
| Backend unreachable | `console couldn't reach the backend` | `backend unreachable` | The query errored. This is a console-to-backend problem, not a statement about the engine. |
| Engine offline | `engine not connected` | `engine offline` | The request succeeded, but the heartbeat key is absent or not `1`. |

Only in the fifth, live state (`engine connected`) do the counters show
numbers.

The Stats screen collapses these slightly differently: with no app
selected it renders a "No app selected" empty state instead of a
telemetry strip, and it distinguishes `loading…`, `couldn't load stats`
(with the underlying request message), and `engine not connected`.

## The minute sampler, buckets, and 30-day retention

History does not come from Redis. A backend sampler records **one row
per app per minute** into `public.realtime_app_stat`, and those rows are
kept for **30 days**. `realtime.history` reads them back.

Consequences worth knowing before you interpret a chart:

- History begins when an app first connects. There is no backfill.
- The first minute takes up to 60s to appear, so a brand-new app shows
  "No samples in this period" for a short while even while it is
  serving traffic.
- Nothing older than the 30-day retention window is available, which is
  also why `30d` is the longest period offered.

The period picker offers `1h`, `6h`, `24h`, `7d`, and `30d`. The
response tells the console the bucket width it used (`bucketSeconds`),
which the console renders as the bucket label — e.g. `5m` or `1h`
buckets — and reuses in every chart sub-caption and in the messages
peak label. The selected period and scope are remembered per viewer in
`localStorage`; they are a UI convenience and are not part of the data.

Each series is aggregated to match its meaning: the connections and
channels charts plot the **average per bucket**, and the messages chart
plots the **total per bucket**. The summary alongside them carries
`connectionsPeak` / `connectionsAvg`, `channelsPeak` / `channelsAvg`,
`messagesTotal` with `messagesPeakPerBucket`, and `firstSampleAt`.

## Coverage, gaps, and engine-offline buckets

Coverage is the fraction of buckets in the window that actually contain
samples, reported as `summary.coverage` and rendered as a percentage.
The console also reports `summary.sampledBuckets`; when that is zero it
treats the whole window as empty and shows a "No samples" empty state
rather than a flat-line chart.

Each history point carries two flags that you need to read together:

| Point flags | Reading |
| ----------- | ------- |
| `sampled: false` | Nobody recorded this bucket. It is a **gap**, not zero traffic. The chart leaves it as a gap. |
| `sampled: true`, `wired: false` | The bucket was recorded while the engine was **not** heartbeating. The console counts these and appends `<n> engine-offline` to the history sub-caption. |
| `sampled: true`, `wired: true` | A normal bucket. |

The header line summarises all of this as: window `from → to`, bucket
width, `<n>% sampled`, and the engine-offline count when there is one.
Low coverage is surfaced rather than smoothed — the Coverage KPI
switches to a warning tone below 90% sampled, and its sub-caption shows
`since <firstSampleAt>` so you can tell "the window predates this app"
apart from "the sampler missed minutes".

If the history request fails outright, the console says so explicitly
and notes that the sampler's table may not be migrated yet — a distinct
condition from an empty window.

## What is counted per app and what is counted per channel

The engine keeps two different granularities, and the console's shape
follows them exactly.

**Per app** (from the counter keys above): connections, occupied
channels, and messages in the last 60s.

**Per channel** (from the engine's per-channel registry): the channel
name, its type, and its subscriber count. The registry lives at
`clutchcall:realtime:app:<appId>:channels:list` as a JSON array
with a short TTL, and `realtime.channels` reports `wired` as simply
"that key is present".

Two things follow from this split:

- **There is no per-channel message rate.** The engine tracks message
  counts per app only. The per-channel table therefore has no
  messages column, and the per-channel `msgsPerMin` field that the API
  returns is an honest zero kept only so the table's row shape stays
  stable. Do not read it as "this channel is silent".
- **`wired: false` from `realtime.channels` is not an outage.** The
  engine writes nothing for an idle app, so the absence of the
  `channels:list` key usually means "no channels right now". The
  console gates the whole per-channel section on the *global* engine
  heartbeat instead, and inside that section an empty registry is shown
  as "No active channels" — not as a disconnected engine.

The live breakdowns derived from the registry are the channel-type
donut (`channelMix`, with zero-valued types dropped) and the top
channels list ranked by subscriber count. The per-channel table can be
filtered by type and by name substring. Busy apps can have thousands of
channels, so the table is capped; when the filtered set exceeds the cap
the console sorts by subscribers and shows the top rows, and says so in
the card sub-caption.

## App scope versus organization scope

Live telemetry is always for one app: the counter keys and the channel
registry are keyed by that app's public `app_id`, and an app that does
not belong to the calling organization is rejected.

History has two scopes, chosen with the **This app** / **All apps**
toggle:

- **This app** — the request carries the `app_id`, and the sampler rows
  for that app alone are returned.
- **All apps** — the request omits the `app_id`, and the organization's
  whole realtime footprint is returned: every app the sampler records,
  combined. The response reports how many apps were combined
  (`apps`), which the console prints as `<n> apps combined` in the
  history sub-caption.

Use **All apps** when a single app shows no samples but you want to know
whether the organization has any realtime traffic at all. An `apps`
count of zero means the organization has no realtime apps yet, which the
console states rather than showing an empty chart.

## Related

- [Authentication](/concepts/authentication) — org-scoped API keys behind these procedures
- [Telemetry](/platform/telemetry) — gateway-side metrics, traces, and CDRs
