# Relay network and edge cache

> How stream PoPs register, how the edge cache-hit ratio and fan-out ratio are computed, what PoP statuses mean for viewers, and what to do when the rollup is unavailable.

The **Relay network** screen is the operable view of the ClutchCall MoQ
relay mesh. Where the architecture pages describe the mesh as a design,
this screen shows it as a list of points of presence (PoPs) with live
numbers attached: how many viewers each edge is serving, and how much of
that traffic it served from its own cache instead of pulling from origin.

Everything on the screen comes from one procedure:

```ts
trpc.streams.analytics.popHealth.useQuery({ orgId })
```

It joins the **`stream_pop` catalog** in Postgres — the registry of PoPs
and where they are — with the **`stream_pop_cache_1m_q` rollup** in
ClickHouse, which supplies the last hour of cache-hit and viewer numbers.
The screen re-runs the query every 30 seconds, and the **Refresh** button
runs it on demand.

## How a PoP is registered

PoPs are not created from the console. There is no "add PoP" action on
this screen. A row appears in the `stream_pop` catalog when a relay
deployment registers itself; the console only reads that catalog.

That is why the empty state reads:

> No PoPs registered yet — PoPs appear when the relay deploy registers them.

An empty mesh is therefore a deployment fact, not a telemetry failure. If
you expect PoPs and see none, the question to ask is whether the relay
deploy in that environment came up and registered, not whether the
analytics pipeline is healthy.

The catalog also supplies each PoP's map coordinates (`map_x`, `map_y`),
which is what places the pins on the **Edge relay map**. Map pins are
catalog data; the colour of a pin comes from `status`.

## Fields on each PoP row

| Field | Source | Meaning |
| ----- | ------ | ------- |
| `code` | `stream_pop` (PG) | The PoP identifier, shown in mono in both panels. |
| `city` | `stream_pop` (PG) | Human-readable location, shown next to the code. |
| `map_x`, `map_y` | `stream_pop` (PG) | Pin position on the relay map. |
| `status` | `stream_pop` (PG) | `active`, `maintenance`, or any other value (treated as bad). |
| `hit_ratio` | `stream_pop_cache_1m_q` (CH) | Edge cache-hit percentage over the query window. |
| `viewers` | `stream_pop_cache_1m_q` (CH) | Concurrent viewers attributed to that PoP. |

The response also carries a top-level `telemetry_ok` flag. See
[When telemetry is unavailable](#when-telemetry-is-unavailable).

## What the cache-hit ratio counts

`hit_ratio` is a per-PoP number computed by the ClickHouse rollup and read
verbatim by the console — the screen does not recompute it. It is the
edge-cache hit share for that PoP, labelled **origin offload** on the KPI
card: the higher the number, the more of the fan-out that relay absorbed
locally instead of forwarding upstream to origin.

The window is the **last hour**, assembled from the one-minute buckets in
`stream_pop_cache_1m_q`. It is not an instantaneous reading. A PoP that
recovered two minutes ago still carries the previous fifty-eight minutes
of misses in its bar.

The **Edge cache hit** KPI at the top of the screen is the **plain
arithmetic mean of `hit_ratio` across every PoP in the response**. It is
not weighted by viewers, and it does not exclude PoPs in maintenance. A
small PoP with a bad ratio moves the headline number exactly as much as
your largest one does, so treat the KPI as a mesh-wide smoke alarm and the
per-PoP bars as the actual diagnosis.

The KPI card and the per-PoP bars use different colour cut-points, which
is worth knowing before you reconcile them:

| Surface | Green / accent | Warn | Bad |
| ------- | -------------- | ---- | --- |
| **Edge cache hit** KPI (mesh average) | ≥ 95% | ≥ 85% | below 85% |
| Per-PoP bar in **Cache-hit ratio · per PoP** | ≥ 96% | ≥ 90% | below 90% |

So a PoP can sit in warn on its bar while the mesh average it feeds is
still green. That is expected, not a bug.

## Reading the fan-out ratio

The **Fan-out ratio** KPI is displayed as `1:N`, where:

```
N = round(sum(viewers across all PoPs) / number of PoPs)
```

It answers one question: *on average, how many viewers is a single relay
serving from one upstream copy?* That is the number that justifies the
relay mesh — the origin publishes once, and each edge multiplies it. The
**Origin → edge fan-out** panel restates the same arithmetic in its two
ends: `ORIGIN / ingest.live` on the left, `<N> RELAYS` and the summed
viewer count on the right.

Two caveats when you quote this number:

- The denominator is **every PoP in the catalog**, including any in
  maintenance. Taking a region out for maintenance lowers the displayed
  fan-out ratio even though the surviving relays are each doing *more*
  work.
- It is a mesh-wide average, so it hides skew. A mesh where one PoP carries
  most of the audience and the rest carry a handful shows the same `1:N` as
  an evenly balanced one. Read the per-PoP viewer counts in the fan-out
  panel to see the distribution.

When there are no PoPs, the ratio renders as `—` rather than `1:0`.

## PoP statuses and what they mean for viewers

`status` comes from the Postgres catalog and drives the pin colour on the
map, the status pip in the fan-out list, and the two counts under the map.

| Status | Indicator | What it means for delivery |
| ------ | --------- | -------------------------- |
| `active` | Green pip, counted under **Healthy** | The PoP is in rotation. Anycast can land viewers on it and it serves them from its edge cache. |
| `maintenance` | Amber pip, counted under **Maintenance** | The PoP is flagged as deliberately out of normal service. It still appears in the catalog, still contributes its rollup numbers to the KPIs, and is still counted in the PoP total. |
| anything else | Red pip | Not counted under either legend total. Treat an unrecognised status as a PoP to investigate — it is in the catalog but is neither healthy nor deliberately drained. |

The **Active PoPs** KPI shows the *total* number of rows returned, with the
healthy (`active`) count in the hint beneath it. When those two numbers
diverge, the difference is PoPs in maintenance plus PoPs in some other
state. A red pip with no corresponding legend entry is the case that most
often goes unnoticed.

## Live relay counters versus the rollup

The screen deliberately stacks two sources of truth, in this order:

1. **Relay health** (the `RelayHealth` panel, scoped to the `stream`
   namespace) — a direct scrape of the relay's own counters, read now.
2. **The PoP KPIs and bars** — ClickHouse rollups, which lag behind by
   about a minute because they are assembled from one-minute buckets.

The live scrape sits above the rollup KPIs for a reason: **when the two
disagree, the live panel is the current truth.** During an incident, the
rollup is the shape of the last hour and the live panel is the shape of
right now. Use the live panel to decide whether something is still
happening, and the rollup to decide how long it has been happening and
which regions it touched.

## When telemetry is unavailable

The `popHealth` response carries `telemetry_ok`. When the ClickHouse side
of the join is unavailable, the catalog still returns PoPs but the
cache-hit numbers cannot be trusted. The screen degrades rather than
showing zeros:

- The **Edge cache hit** KPI renders `—` instead of a percentage, drops its
  severity colour entirely, and sets its hint to `telemetry unavailable`.
- Every per-PoP cache-hit value in the **Cache-hit ratio · per PoP** panel
  renders `—`.
- The **Telemetry health banner** at the top of the page explains the
  degraded state.

This is a display rule with a specific intent: a red `0.0%` would claim
that every edge is missing cache, which is a different and far more
alarming statement than "we cannot currently measure it". Alarm colour on
this screen is reserved for conditions that are genuinely wrong.

The same suppression applies when the mesh is simply **empty**. With zero
PoPs, the KPI shows `—` with the hint `no traffic yet` rather than a red
`0.0%`. No PoPs means no measurement, not a total cache miss.

Note that the per-PoP **bars** still draw at their last-read widths while
`telemetry_ok` is false; only the numeric labels blank out. Read the `—`
values, not the bar lengths, when the banner is up.

Viewer counts and the fan-out ratio are **not** suppressed by
`telemetry_ok`. If the banner is up, treat those two numbers with the same
suspicion as the cache-hit numbers — they come from the same rollup.

## Diagnosing one unhealthy region

When a single PoP's bar is amber or red while the rest of the mesh is
green, work through the panels in this order:

1. **Check the status pip first.** If the PoP is in `maintenance`, its
   cache-hit number is expected to sag and needs no action — it is being
   drained on purpose, and it is still dragging the unweighted mesh
   average down. If the pip is red with no legend entry, the catalog state
   itself is the problem.
2. **Check the telemetry banner.** A degraded rollup makes every number on
   the page provisional. Resolve that before you chase a region.
3. **Compare against the live relay health panel.** The bar covers the last
   hour. If the live counters look normal, you are most likely looking at
   the tail of an incident that has already ended, and the bar will recover
   as the older one-minute buckets age out of the window.
4. **Read that PoP's viewer count in the fan-out panel.** A low hit ratio
   on a PoP with very few viewers has little effect on origin load, however
   red the bar looks; a low ratio on the PoP carrying most of the audience
   is where origin egress is actually going.
5. **Re-check the mesh average.** If exactly one PoP moved and the **Edge
   cache hit** KPI changed noticeably, that is the unweighted mean at work
   — a reminder that the KPI is a signal to open the per-PoP panel, not a
   measure of how much origin traffic you are actually paying for.

## Related

- [Telemetry](/platform/telemetry) — the metric, trace, and rollup streams behind these numbers
- [Architecture](/concepts/architecture) — where the relay mesh sits in the data plane
- [Authentication](/concepts/authentication) — relay tokens and the `namespace_auth` hook that gates edge subscribe
