# Fleet & Latency

> How the Fleet & Latency screen builds its two fleet views: the static GPU fleet list with measured and target RTTs, and the mod_compute beacon that reports dial-out nodes.

The **Fleet & Latency** screen shows two unrelated inventories on one
page. Reading it correctly means knowing which numbers are measured,
which are constants baked into the code, and which arrive from a live
beacon that can go stale.

## The two fleets on this screen

| Card | Source | Direction |
| ---- | ------ | --------- |
| **Compute fleet** (top card) | `compute.fleetStatus` | Nodes that dial *out* and connect to ClutchCall, then run workloads. |
| **Provisioned & available GPUs** (table, latency bars, thesis KPIs) | `inference.gpuFleet` | Rented pods and vendor endpoints that ClutchCall routes inference *to*. |

They do not share rows, statuses, or a refresh mechanism. The compute
fleet card polls every 10 seconds. The GPU fleet is fetched once per
screen load and is not a live measurement at all — see
[Where the RTT numbers come from](#where-the-rtt-numbers-come-from-and-how-to-refresh-them).

A third card, **Per-call metrics**, comes from `inference.gatewayStats`
and is genuine query output over ClickHouse — it is not part of either
fleet.

## GPU fleet rows: statuses, measured vs target RTTs

Every row in the GPU table carries a `status` that says what the row is
for, not how healthy it is:

| Status | Pill | Meaning |
| ------ | ---- | ------- |
| `live` | LIVE (green) | Currently serving inference. These rows drive the latency-attribution bars and `liveWorstRttMs`. |
| `available` | AVAILABLE (blue) | In-region capacity that exists and could be provisioned, but is not serving. |
| `vendor` | VENDOR (grey) | A managed vendor option under consideration; not provisioned. |
| `target` | TARGET | A placement target — where capacity is wanted, not where it is. |

The `rttMs` column is the round trip from the serving region
(`servingRegion`, shown in the header) to that row's region. It is
annotated two ways:

- **Unmarked** — `measured: true`. An ICMP average taken from the
  serving-region engine on the date in `measuredAt`.
- **Marked with `*`** — `measured: false`. A placement target or
  estimate derived from verified datacenter availability. No packet was
  sent for this number.

Colour on the RTT cell is a pure threshold on the value the procedure
returned: green at or below 5 ms, red at or above 200 ms, default in
between. It is not a health signal.

The **region matrix** (`regionMatrix`) is derived, not stored: rows are
deduplicated by region code, and when two rows share a region the
measured one wins. It is sorted ascending by RTT and supplies the
maximum used to scale the latency bars.

## Where the RTT numbers come from and how to refresh them

`inference.gpuFleet` returns a **hardcoded array**. There is no probe
loop behind this screen. The procedure builds `fleet`, derives
`regionMatrix`, computes `liveWorstRttMs` as the maximum RTT among
`live` rows, and returns a fixed `inRegionRttMs` and `measuredAt`.

Consequences worth stating plainly:

- `measuredAt` is the date the ICMP averages were taken, not the date
  the page was loaded. A row marked as measured is measured *as of that
  date*.
- Adding or removing a rented pod does not change this screen. The list
  changes when the list is edited.
- `inRegionRttMs` — the green "in-region" comparison bar and KPI — is a
  constant in the procedure, not a probe of an in-region GPU. The
  in-region row in the table is a `target`, so there is nothing in the
  fleet to probe.

The procedure is baked in for demonstration purposes and is designed to
move behind a Redis key (`clutchcall:inference:fleet`) for live
operations. Until that key is wired up, treat the table as a curated
inventory document rendered as a table, and refresh it by editing the
procedure.

If the query fails or returns `ok: false`, the whole screen collapses to
a single **Fleet unavailable** card. The compute fleet card is not
rendered in that state either.

## The 8 ms stack allowance and what it does not include

`STACK_MS = 8` is a constant declared in the screen component. It
represents the engine-side QUIC transport and framing cost per round
trip. It is not read from any procedure, and no telemetry feeds it.

It is used in three places:

- **"Our stack overhead"** KPI renders it directly as `~8 ms`.
- **"Attributable to ocean"** KPI is
  `round((liveWorstRttMs - 8) / liveWorstRttMs * 100)`.
- **Latency bars** split each bar into a red ocean segment sized
  `(rttMs - 8) / max` and a green stack segment sized `8 / max`. Rows at
  or below 5 ms RTT are drawn as an all-green bar with no ocean segment.

So the red/green split on every bar is arithmetic against a constant,
not two separately measured quantities. If the real transport cost
changes, nothing on this screen will notice.

What the 8 ms figure does **not** cover: model execution time, queueing
or admission at the GPU, ASR/TTS/LLM generation, or anything on the
caller's access network. Those show up, blended, in the per-call
occupancy figures below.

## The compute fleet beacon: TTL, draining, reporting failures

`compute.fleetStatus` reads a single Redis key,
`clutchcall:compute:fleet`, and parses it as a flat string map. It
returns one of two shapes:

- `{ reporting: false }` — the key is absent **or** the payload failed
  to parse.
- `{ reporting: true, ... }` — the numeric fields below, each coerced
  from the string map with a default of `0` for missing keys.

| Field | Rendered as |
| ----- | ----------- |
| `nodesAlive` | **Nodes alive** value; tinted green when greater than zero. |
| `nodes` | `"<n> enrolled"` hint. |
| `nodesDraining` | Appended to the hint as `"· <n> draining"`, only when non-zero. |
| `cpuMillis` / `cpuMillisUsed` | **CPU committed**, converted to cores, against available. |
| `memMb` / `memMbUsed` | **Memory committed**, converted to GB, against available. |
| `gpus` / `gpusUsed` | **GPUs committed** against fleet total. |
| `nodesCommunity` / `cpuMillisCommunity` | **Community nodes** tile — see below. |

mod_compute **republishes** these totals on a TTL. That is why absence
is meaningful: when nothing has published recently, the key expires and
the card shows *"No engine is reporting a fleet."* This is a reporting
outage, not a fleet of zero nodes.

A malformed payload is treated identically. The procedure catches the
parse error and returns `reporting: false` rather than a record of
zeroes, precisely so that a corrupt beacon cannot be mistaken for an
empty fleet. If you see the empty state, check that an engine is
publishing before you conclude that nodes are gone.

**Enrolled vs alive vs draining** are three different counts and the
card shows all three. `nodes` is enrollment. `nodesAlive` is what is
currently reporting in. `nodesDraining` is a subset that is being taken
out of service and should be expected to fall out of `nodesAlive`.

## Community capacity and workload opt-in

`nodesCommunity` and `cpuMillisCommunity` are the contributed share of
what is alive — capacity on hardware ClutchCall does not operate.

They are deliberately **broken out rather than folded into the totals**.
Capacity that a workload cannot reach unless it opted in is not
interchangeable with capacity that it can, and a single combined number
would assert that it is.

The tile renders only when `nodesCommunity > 0`. On an all-secure fleet
a permanent zero tile would be noise; the moment it is non-zero it is
the most consequential number on the card, because only workloads that
opted in can be scheduled onto it. If you are sizing against the CPU,
memory, and GPU tiles, subtract the community share unless every
workload you care about has opted in.

## Reading per-call occupancy against leg RTT

The **Per-call metrics** card comes from
`inference.gatewayStats({ windowDays })`, defaulting to a 7-day window.
Unlike the GPU fleet, this is a real aggregation — over
`engine_billing_event`, filtered to `dimension LIKE 'inference.%'`.
mod_inference emits these names through the meter path, so the metric
table holds none of them; querying it returns nothing.

| KPI | Aggregation |
| --- | ----------- |
| **Requests** | Count of rows with dimension `inference.occupancy_ms`. |
| **Avg generation** | Average `units` on `inference.occupancy_ms`, in ms. |
| **Total GPU time** | Sum of `inference.occupancy_ms`, rendered in seconds. |
| **Units** | Sum of `inference.units` — tokens and audio units generated. |

The critical joint reading: **avg generation includes the ocean RTT of
whichever leg served the turn.** GPU time held per turn is wall time on
a leg whose round trip you can read off the table above. A turn served
from a `live` row is paying that row's `rttMs` inside the occupancy
figure; the same model co-located would not.

So the comparison the screen invites is: take **Avg generation**, and
subtract the RTT of the leg for the role in question to reason about how
much is generation and how much is distance. The screen does not do that
subtraction for you, and it cannot — occupancy is aggregated fleet-wide
across all roles and legs, with no per-leg breakdown in this query.

`gatewayStats` also computes cancellation counters and a cancel rate
from the same window. Those are returned by the procedure but are not
displayed on this screen.

When the query throws, the procedure returns `ok: false` with a null
aggregate and the card shows **"No calls in window."** An empty result
set produces the same card. The two are not distinguished here.

## Related

- [Telemetry](/platform/telemetry) — the metric, trace, and CDR streams the gateway emits
- [Telephony Metrics](/glossary/metrics) — RTT, one-way latency, and jitter definitions
