The Fleet & Latency screen shows two unrelated inventories on one page. Reading it correctly means knowing which numbers are measured, which are constants baked into the code, and which arrive from a live beacon that can go stale.

The two fleets on this screen

They do not share rows, statuses, or a refresh mechanism. The compute fleet card polls every 10 seconds. The GPU fleet is fetched once per screen load and is not a live measurement at all — see Where the RTT numbers come from. A third card, Per-call metrics, comes from inference.gatewayStats and is genuine query output over ClickHouse — it is not part of either fleet.

GPU fleet rows: statuses, measured vs target RTTs

Every row in the GPU table carries a status that says what the row is for, not how healthy it is: The rttMs column is the round trip from the serving region (servingRegion, shown in the header) to that row’s region. It is annotated two ways:
  • Unmarked — measured: true. An ICMP average taken from the serving-region engine on the date in measuredAt.
  • Marked with * — measured: false. A placement target or estimate derived from verified datacenter availability. No packet was sent for this number.
Colour on the RTT cell is a pure threshold on the value the procedure returned: green at or below 5 ms, red at or above 200 ms, default in between. It is not a health signal. The region matrix (regionMatrix) is derived, not stored: rows are deduplicated by region code, and when two rows share a region the measured one wins. It is sorted ascending by RTT and supplies the maximum used to scale the latency bars.

Where the RTT numbers come from and how to refresh them

inference.gpuFleet returns a hardcoded array. There is no probe loop behind this screen. The procedure builds fleet, derives regionMatrix, computes liveWorstRttMs as the maximum RTT among live rows, and returns a fixed inRegionRttMs and measuredAt. Consequences worth stating plainly:
  • measuredAt is the date the ICMP averages were taken, not the date the page was loaded. A row marked as measured is measured as of that date.
  • Adding or removing a rented pod does not change this screen. The list changes when the list is edited.
  • inRegionRttMs — the green “in-region” comparison bar and KPI — is a constant in the procedure, not a probe of an in-region GPU. The in-region row in the table is a target, so there is nothing in the fleet to probe.
The procedure is baked in for demonstration purposes and is designed to move behind a Redis key (clutchcall:inference:fleet) for live operations. Until that key is wired up, treat the table as a curated inventory document rendered as a table, and refresh it by editing the procedure. If the query fails or returns ok: false, the whole screen collapses to a single Fleet unavailable card. The compute fleet card is not rendered in that state either.

The 8 ms stack allowance and what it does not include

STACK_MS = 8 is a constant declared in the screen component. It represents the engine-side QUIC transport and framing cost per round trip. It is not read from any procedure, and no telemetry feeds it. It is used in three places:
  • “Our stack overhead” KPI renders it directly as ~8 ms.
  • “Attributable to ocean” KPI is round((liveWorstRttMs - 8) / liveWorstRttMs * 100).
  • Latency bars split each bar into a red ocean segment sized (rttMs - 8) / max and a green stack segment sized 8 / max. Rows at or below 5 ms RTT are drawn as an all-green bar with no ocean segment.
So the red/green split on every bar is arithmetic against a constant, not two separately measured quantities. If the real transport cost changes, nothing on this screen will notice. What the 8 ms figure does not cover: model execution time, queueing or admission at the GPU, ASR/TTS/LLM generation, or anything on the caller’s access network. Those show up, blended, in the per-call occupancy figures below.

The compute fleet beacon: TTL, draining, reporting failures

compute.fleetStatus reads a single Redis key, clutchcall:compute:fleet, and parses it as a flat string map. It returns one of two shapes:
  • { reporting: false } — the key is absent or the payload failed to parse.
  • { reporting: true, ... } — the numeric fields below, each coerced from the string map with a default of 0 for missing keys.
mod_compute republishes these totals on a TTL. That is why absence is meaningful: when nothing has published recently, the key expires and the card shows “No engine is reporting a fleet.” This is a reporting outage, not a fleet of zero nodes. A malformed payload is treated identically. The procedure catches the parse error and returns reporting: false rather than a record of zeroes, precisely so that a corrupt beacon cannot be mistaken for an empty fleet. If you see the empty state, check that an engine is publishing before you conclude that nodes are gone. Enrolled vs alive vs draining are three different counts and the card shows all three. nodes is enrollment. nodesAlive is what is currently reporting in. nodesDraining is a subset that is being taken out of service and should be expected to fall out of nodesAlive.

Community capacity and workload opt-in

nodesCommunity and cpuMillisCommunity are the contributed share of what is alive — capacity on hardware ClutchCall does not operate. They are deliberately broken out rather than folded into the totals. Capacity that a workload cannot reach unless it opted in is not interchangeable with capacity that it can, and a single combined number would assert that it is. The tile renders only when nodesCommunity > 0. On an all-secure fleet a permanent zero tile would be noise; the moment it is non-zero it is the most consequential number on the card, because only workloads that opted in can be scheduled onto it. If you are sizing against the CPU, memory, and GPU tiles, subtract the community share unless every workload you care about has opted in.

Reading per-call occupancy against leg RTT

The Per-call metrics card comes from inference.gatewayStats({ windowDays }), defaulting to a 7-day window. Unlike the GPU fleet, this is a real aggregation — over engine_billing_event, filtered to dimension LIKE 'inference.%'. mod_inference emits these names through the meter path, so the metric table holds none of them; querying it returns nothing. The critical joint reading: avg generation includes the ocean RTT of whichever leg served the turn. GPU time held per turn is wall time on a leg whose round trip you can read off the table above. A turn served from a live row is paying that row’s rttMs inside the occupancy figure; the same model co-located would not. So the comparison the screen invites is: take Avg generation, and subtract the RTT of the leg for the role in question to reason about how much is generation and how much is distance. The screen does not do that subtraction for you, and it cannot — occupancy is aggregated fleet-wide across all roles and legs, with no per-leg breakdown in this query. gatewayStats also computes cancellation counters and a cancel rate from the same window. Those are returned by the procedure but are not displayed on this screen. When the query throws, the procedure returns ok: false with a null aggregate and the card shows “No calls in window.” An empty result set produces the same card. The two are not distinguished here.
  • Telemetry — the metric, trace, and CDR streams the gateway emits
  • Telephony Metrics — RTT, one-way latency, and jitter definitions