# Fleet health and the degraded telemetry state

> How the robotics Overview counters are derived, why liveness can read as unknown, and which numbers stay authoritative when the telemetry rollup is down.

The robotics **Overview** screen opens on a four-card strip: Robots,
Online, Errored, and Fleet rate. Every number on that strip comes from a
single ClutchCall control-plane call, `fleetRobotics.kpiStrip`, scoped
to the current org. The same procedure backs the Fleet directory, so the
two screens never disagree.

This page explains what each counter counts, where it is sourced from,
and what happens to the strip when the telemetry rollup cannot be
reached.

## What the four counters measure

| Card | Field on `kpiStrip` | Source | Blanked when degraded |
| ---- | ------------------- | ------ | --------------------- |
| Robots | `total_robots` | Robot registry | No |
| Online | `online_robots` | Telemetry rollup | **Yes** |
| Errored | `errored_robots` | Robot registry | No |
| Fleet rate | `fleet_fps` | Telemetry rollup | **Yes** |

**Robots** is the number of robots registered to the org. It is a
registry count, not a connectivity count — a robot that has never
published a topic still appears here.

**Online** is the number of robots the telemetry rollup currently
considers live. See the next section.

**Errored** is the number of robots the registry currently holds in an
errored state. The card renders in the `bad` severity when the value is
above zero, and renders plain at zero.

**Fleet rate** is the aggregate topic publish rate across the fleet,
reported in hertz. It is a rollup-derived number, so it reflects traffic
that actually reached the WAN transport, not what a robot believes it is
publishing locally.

While the query is in flight, every card shows `…` rather than a zero.

## Liveness comes from telemetry, not from a heartbeat

There is no separate liveness ping for robots. A robot is online because
its telemetry is arriving and being rolled up — that stream is the only
liveness signal the console has.

The practical consequence is that **Online** and **Fleet rate** are
downstream of the same pipeline. If the rollup stops, both lose their
input at the same moment, and neither one can distinguish "the fleet
went quiet" from "we stopped being able to see the fleet." That is why
the console refuses to guess.

## The degraded state: unknown versus zero

When the ClickHouse rollup behind the strip is unreachable,
`kpiStrip` returns `source: 'degraded'`. In that state the Overview:

- renders **Online** and **Fleet rate** as an em dash (`—`), not as `0`;
- drops the `ok` severity from the Online card and attaches the hint
  *Telemetry unavailable*;
- shows a banner explaining that liveness and topic rate are unknown and
  that the robot count below is still authoritative;
- keeps **Robots** and **Errored** rendered as normal numbers.

Printing `0 online` during an ingest outage would make a broken pipeline
look like a dead fleet, which is the more expensive mistake to make at
3 a.m. An em dash means *we do not know*; a zero means *we looked and
there were none*.

Registry-sourced counters are unaffected by the degraded state. If you
are triaging during an outage, `Robots` and `Errored` are still safe to
act on. Anything sourced from the rollup is not.

The Fleet directory applies the same rule, so a degraded strip on the
Overview and a degraded directory are the same condition, not two
separate faults.

## The errored counter

`errored_robots` survives the degraded state, which tells you what it is
and what it is not: it is registry state, not a telemetry-derived
inference. A robot counted here is recorded as errored regardless of
whether its telemetry is currently flowing, and a telemetry outage will
not inflate or deflate the number.

The strip only gives you the count. To see *which* robots are errored,
open the Fleet directory — it is backed by the same procedure and filters
down to individual rows.

## How often the numbers refresh

The Overview polls `fleetRobotics.kpiStrip` every 30 seconds. There is no
push channel behind the strip, so a value can be up to one polling
interval stale, and recovery from the degraded state becomes visible on
the next successful poll rather than immediately.

The query does not run at all until an org is selected; with no org id
the cards stay in their loading state.

## Where to go next when a counter looks wrong

The Overview's second panel links the four screens that answer the
follow-up questions:

- **Robots is zero, or lower than expected.** Nothing is registered yet,
  or the robot you expected was never added. Start at the quickstart for
  the middleware swap and the install command, then register the robot
  from the Fleet directory.
- **Online is lower than Robots.** Some robots are registered but their
  telemetry is not arriving. Open Topics to confirm the graph is
  publishing over the WAN rather than only on the local network.
- **Online and Fleet rate are both em dashes.** This is the degraded
  state, not a fleet problem. Treat it as a telemetry-pipeline
  investigation and rely on Robots and Errored until the rollup returns.
- **Fleet rate is low but Online looks correct.** Robots are connected
  but publishing little. Topics shows which topics are flowing; the
  connection-quality screen shows per-link rate and loss, which is the
  pair of numbers teleoperation actually depends on.
- **Errored is above zero.** Open the Fleet directory to identify the
  specific robots; the count alone does not carry a reason.
