# Connection telemetry

> How the connection telemetry screen derives RTT percentiles, congestion window, path migrations, and byte counters from a robot's relay session.

The **Connection telemetry** screen is the transport-level view of a single
robot. Where the robotics details page describes namespaces, the CDR frame,
and QoS lanes — what you send — this screen describes the QUIC connection
those frames travel over: round-trip time, congestion window, path
migrations, and byte counters.

The screen is scoped to one robot at a time and one fixed window: **the last
15 minutes**. Every number on the page is derived from two queries.

| Query | Input | Used for |
| ----- | ----- | -------- |
| `fleetRobotics.robotsList` | `{ orgId, page: 1, perPage: 200 }` | The robot picker in the header, and the default focus when no robot is in the URL. |
| `fleetRobotics.connectionTelemetry` | `{ orgId, robotId, windowMinutes: 15 }` | Everything else on the page. |

## What the robot reports and how often

`connectionTelemetry` returns two arrays of **per-minute aggregate rows** for
the requested window. Each row is one minute of the robot's relay session,
already reduced server-side — the screen never sees individual RTT samples or
individual packets.

| Array | Fields the screen reads |
| ----- | ----------------------- |
| `rtt` | `rtt_p50`, `rtt_p95`, `rtt_p99` |
| `cwnd` | `cwnd_avg`, `migrations`, `bytes_in`, `bytes_out` |

Refresh behaviour:

- The telemetry query refetches every **30 seconds**.
- The fleet list behind the picker refetches every **60 seconds** and is
  considered stale after 30 seconds.
- The refresh button in the header forces a refetch and is disabled while one
  is in flight. The header shows `Refreshing…` during a fetch, otherwise
  `Updated <relative time>` based on when the telemetry query last resolved.

Because rows are one-minute aggregates, a 15-minute window yields roughly 15
points. Two panel titles and the congestion-window axis are labelled
`last 60s` / `−60 s … now`; the underlying data is the 15-minute window at
one-minute resolution. Read the axis as "oldest sample → newest sample", not
as literal seconds.

## Choosing the focused robot

The focused robot comes from the route parameter `/telemetry/:id`. If the
route has no id and the fleet list has loaded, the screen redirects to the
first robot in the list. The header `<select>` navigates to
`/telemetry/<id>` on change.

Each option is labelled `<name> · <id>`, where the name falls back through
`name` → `display_name` → `external_robot_id`, and the id falls back from
`id` to `external_robot_id`.

If the id in the URL is **not** present in the fleet list, the picker adds a
synthetic option labelled `<id> (not in active fleet)` so the selection stays
visible. The telemetry query still runs for that id — a robot that has been
removed from the active fleet can still have telemetry rows in the window.

Telemetry is only queried when both an org id and a focused robot id are
present.

## RTT percentiles: where they are measured

The RTT percentiles describe the **QUIC connection between the robot and the
relay**, for the duration of the robot's relay session. They are a property of
that session, not of any single track or namespace: when there is no active
relay session for the robot, `rtt` comes back empty.

Three percentiles are surfaced, and they are aggregated differently:

| Displayed value | Derivation |
| --------------- | ---------- |
| `P50` | The **median across minutes** of the per-minute `rtt_p50`. |
| `P95` | The **median across minutes** of the per-minute `rtt_p95`. |
| `P99` | The **maximum across minutes** of the per-minute `rtt_p99` — the worst tail anywhere in the window. |

P50 and P95 are therefore "a typical minute", while P99 is "the worst minute".
A single bad minute moves P99 and leaves P50 and P95 alone. That is
intentional: the tail is the thing you want to see even if it happened once.

The tag in the panel corner is a display label derived from the worst P99:
above `1000 ms` it reads `tail expanded`, otherwise `within band`. This is a
colouring rule for this screen only. It is not a service commitment, and it
says nothing about the other percentiles.

When `rtt` is empty the tag reads `no data` and all three percentages render
as `—` rather than `0`. This matters: the medians of an empty array evaluate
to zero, and a zero RTT must never be read as a healthy result.

## How the RTT histogram is built

The large histogram is not a distribution of RTT samples. It is a
distribution of **minutes**, bucketed by that minute's median RTT:

- 20 buckets, **75 ms wide**, covering 0–1500 ms.
- Each row in `rtt` increments exactly one bucket, chosen by
  `floor(rtt_p50 / 75)`, clamped to the last bucket.
- The bar count under the chart (`N minutes`) is the sum of all buckets, which
  equals the number of rows returned.
- The marker line is placed on the bucket containing the **worst P99** of the
  window, i.e. `floor(worst_p99 / 75)`, also clamped to the last bucket.

So a tall bar means "many minutes had a median RTT in this range". The marker
sits to the right of the mass whenever the tail is worse than the typical
minute; the distance between them is the spread you care about.

Anything at or above 1500 ms lands in the final bucket. The histogram cannot
distinguish 1.6 s from 6 s — use the numeric P99 readout for that.

## Congestion window and path migration

The congestion-window chart plots `cwnd_avg` from each `cwnd` row, converted
from bytes to **KB** (divided by 1024). The footer reports:

| Readout | Meaning |
| ------- | ------- |
| `Peak` | Largest `cwnd_avg` in the window. |
| `Min` | Smallest `cwnd_avg` in the window. |
| `Now` | The most recent sample — the last row in the array. |
| `Migrations in window` | Sum of `migrations` across all rows. |
| `Samples` | Number of rows plotted. |

A **path migration** is counted by the relay whenever the robot's QUIC
connection continues over a new network path — the connection ID survives the
handover, so there is no new TLS handshake and no session restart. A robot
moving between access points, or switching from Wi-Fi to cellular, produces
migrations without producing a reconnect. The counter is per-minute: a row
with `migrations: 2` means two migrations occurred during that minute.

On the chart, the dashed red `MIG` marker is drawn at the **first** bucket
where `migrations > 0`. Later migrations in the same window are counted in the
`Migrations in window` total but are not drawn. Read the marker as "migration
activity started here", not as "the only migration".

Expect the congestion window to collapse and re-grow around a migration: the
new path has a fresh estimate. A cwnd dip that lines up with the `MIG` marker
is the path change, not congestion on the old path.

The separate **Path migrations** panel is a placeholder — see
[Panels with no data source yet](#panels-with-no-data-source-yet).

## Bytes in and bytes out

`bytes_in` and `bytes_out` are reported **cumulatively per connection**, not
as a per-minute delta. The screen therefore does not sum them. It walks the
rows in order and keeps the last non-zero value of each, which is the running
total for the connection as of the newest bucket in the window.

Consequences worth knowing:

- The figure is a connection lifetime total, so it can exceed what the robot
  actually transferred during the 15-minute window.
- If the connection is replaced, the counter restarts from the new
  connection's baseline, and the tile will appear to drop.
- Zero and missing are treated identically: a non-finite or zero value never
  overwrites the running total, and a total of zero renders as `—` rather than
  `0 B`.

Values are formatted with binary units (B, KB, MB, GB, TB at 1024 steps),
rounded to whole units below 1 KB or at 100 and above, otherwise to one
decimal.

## Per-track delivery and where drops come from

The **Per-track frame ledger** exists because a drop is usually a property of
one track, not of the connection. The connection-level panels above it can
look entirely healthy — steady cwnd, no migrations, flat RTT — while a single
high-rate track is shedding frames. The ledger is the panel that attributes
loss to a track.

Its columns are:

| Column | Meaning |
| ------ | ------- |
| `Track` | Track name. |
| `Direction` | Which way the track flows. |
| `Delivered` | Frames delivered. |
| `Dropped` | Frames dropped. Highlighted when non-zero. |
| `Late` | Frames that arrived but missed their deadline. |
| `FEC recovered` | Frames reconstructed by forward error correction. Highlighted when non-zero. |
| `Drop rate` | `dropped / (delivered + dropped)`, as a percentage, with a bar alongside. |

Because `Late` and `FEC recovered` are tracked separately from `Dropped`, they
do **not** contribute to the drop rate. A track can show a 0.00% drop rate and
still be unusable if its frames are consistently late; read all four counters
together.

This table currently renders no rows — see below.

## Panels with no data source yet

Several panels are laid out but not wired to `connectionTelemetry`. They
render placeholder text or `—` rather than being hidden, so that the shape of
the screen stays stable once the data lands. Do not read a placeholder as a
measurement of zero.

| Panel | Status |
| ----- | ------ |
| **Connection identity** — SNI, ALPN, Cipher, Cert, 0-RTT, Connection, Session | Not wired. All seven tiles render `—`, and the panel hint reads `no active relay session`. |
| **Connection identity** — Bytes in, Bytes out | Wired, from the `cwnd` rows. See [Bytes in and bytes out](#bytes-in-and-bytes-out). |
| **RTT distribution** | Wired, from the `rtt` rows. |
| **Path migrations** (the standalone panel) | Not wired. Always renders `No path-migration events in the active window.` The migration **count** in the congestion-window footer is wired and is the number to trust. |
| **Congestion window** | Wired, from the `cwnd` rows. |
| **Per-track frame ledger** | Not wired. The table renders its header with no rows. |
| **Bandwidth · last 60s** | Not wired. Renders `No bandwidth samples yet.`, and the inbound peak / outbound peak / stream count readouts render `—`. |

## Reading an empty panel: no relay session vs no samples

Two different conditions produce a blank panel, and they call for different
responses.

**No relay session.** `connectionTelemetry` returns empty `rtt` and `cwnd`
arrays when the robot has no active relay session in the window. The RTT panel
shows the `no data` tag with `—` for all three percentiles, and the
congestion-window chart shows `No congestion-window samples yet.` This is a
connectivity question, not a performance question: the robot has not been
attached to the relay during the window. Check that the robot is online and
authenticated before reading anything else on the page.

**No samples for a panel that has no source.** The placeholders in the table
above look similar but mean only that ClutchCall does not yet populate that
panel. They will read empty even for a robot with a fully healthy, actively
transmitting session.

To tell them apart: if the RTT and congestion-window panels have data, the
session exists, and any other empty panel is an unwired panel. If those two
are also empty, there is no session.

The one trap to avoid in both cases is the zero. An empty `rtt` array produces
median and maximum values of `0`. The screen suppresses them and renders `—`
instead, precisely so that "no session" is never displayed as "0 ms RTT". If
you consume `connectionTelemetry` directly rather than through this screen,
apply the same rule: check the array length before reading percentiles, and
never treat a zero-length result as a pass.
