The Relay network screen is the operable view of the ClutchCall MoQ relay mesh. Where the architecture pages describe the mesh as a design, this screen shows it as a list of points of presence (PoPs) with live numbers attached: how many viewers each edge is serving, and how much of that traffic it served from its own cache instead of pulling from origin. Everything on the screen comes from one procedure:
It joins the stream_pop catalog in Postgres — the registry of PoPs and where they are — with the stream_pop_cache_1m_q rollup in ClickHouse, which supplies the last hour of cache-hit and viewer numbers. The screen re-runs the query every 30 seconds, and the Refresh button runs it on demand.

How a PoP is registered

PoPs are not created from the console. There is no “add PoP” action on this screen. A row appears in the stream_pop catalog when a relay deployment registers itself; the console only reads that catalog. That is why the empty state reads:
No PoPs registered yet — PoPs appear when the relay deploy registers them.
An empty mesh is therefore a deployment fact, not a telemetry failure. If you expect PoPs and see none, the question to ask is whether the relay deploy in that environment came up and registered, not whether the analytics pipeline is healthy. The catalog also supplies each PoP’s map coordinates (map_x, map_y), which is what places the pins on the Edge relay map. Map pins are catalog data; the colour of a pin comes from status.

Fields on each PoP row

The response also carries a top-level telemetry_ok flag. See When telemetry is unavailable.

What the cache-hit ratio counts

hit_ratio is a per-PoP number computed by the ClickHouse rollup and read verbatim by the console — the screen does not recompute it. It is the edge-cache hit share for that PoP, labelled origin offload on the KPI card: the higher the number, the more of the fan-out that relay absorbed locally instead of forwarding upstream to origin. The window is the last hour, assembled from the one-minute buckets in stream_pop_cache_1m_q. It is not an instantaneous reading. A PoP that recovered two minutes ago still carries the previous fifty-eight minutes of misses in its bar. The Edge cache hit KPI at the top of the screen is the plain arithmetic mean of hit_ratio across every PoP in the response. It is not weighted by viewers, and it does not exclude PoPs in maintenance. A small PoP with a bad ratio moves the headline number exactly as much as your largest one does, so treat the KPI as a mesh-wide smoke alarm and the per-PoP bars as the actual diagnosis. The KPI card and the per-PoP bars use different colour cut-points, which is worth knowing before you reconcile them: So a PoP can sit in warn on its bar while the mesh average it feeds is still green. That is expected, not a bug.

Reading the fan-out ratio

The Fan-out ratio KPI is displayed as 1:N, where:
It answers one question: on average, how many viewers is a single relay serving from one upstream copy? That is the number that justifies the relay mesh — the origin publishes once, and each edge multiplies it. The Origin → edge fan-out panel restates the same arithmetic in its two ends: ORIGIN / ingest.live on the left, <N> RELAYS and the summed viewer count on the right. Two caveats when you quote this number:
  • The denominator is every PoP in the catalog, including any in maintenance. Taking a region out for maintenance lowers the displayed fan-out ratio even though the surviving relays are each doing more work.
  • It is a mesh-wide average, so it hides skew. A mesh where one PoP carries most of the audience and the rest carry a handful shows the same 1:N as an evenly balanced one. Read the per-PoP viewer counts in the fan-out panel to see the distribution.
When there are no PoPs, the ratio renders as — rather than 1:0.

PoP statuses and what they mean for viewers

status comes from the Postgres catalog and drives the pin colour on the map, the status pip in the fan-out list, and the two counts under the map. The Active PoPs KPI shows the total number of rows returned, with the healthy (active) count in the hint beneath it. When those two numbers diverge, the difference is PoPs in maintenance plus PoPs in some other state. A red pip with no corresponding legend entry is the case that most often goes unnoticed.

Live relay counters versus the rollup

The screen deliberately stacks two sources of truth, in this order:
  1. Relay health (the RelayHealth panel, scoped to the stream namespace) — a direct scrape of the relay’s own counters, read now.
  2. The PoP KPIs and bars — ClickHouse rollups, which lag behind by about a minute because they are assembled from one-minute buckets.
The live scrape sits above the rollup KPIs for a reason: when the two disagree, the live panel is the current truth. During an incident, the rollup is the shape of the last hour and the live panel is the shape of right now. Use the live panel to decide whether something is still happening, and the rollup to decide how long it has been happening and which regions it touched.

When telemetry is unavailable

The popHealth response carries telemetry_ok. When the ClickHouse side of the join is unavailable, the catalog still returns PoPs but the cache-hit numbers cannot be trusted. The screen degrades rather than showing zeros:
  • The Edge cache hit KPI renders — instead of a percentage, drops its severity colour entirely, and sets its hint to telemetry unavailable.
  • Every per-PoP cache-hit value in the Cache-hit ratio · per PoP panel renders —.
  • The Telemetry health banner at the top of the page explains the degraded state.
This is a display rule with a specific intent: a red 0.0% would claim that every edge is missing cache, which is a different and far more alarming statement than “we cannot currently measure it”. Alarm colour on this screen is reserved for conditions that are genuinely wrong. The same suppression applies when the mesh is simply empty. With zero PoPs, the KPI shows — with the hint no traffic yet rather than a red 0.0%. No PoPs means no measurement, not a total cache miss. Note that the per-PoP bars still draw at their last-read widths while telemetry_ok is false; only the numeric labels blank out. Read the — values, not the bar lengths, when the banner is up. Viewer counts and the fan-out ratio are not suppressed by telemetry_ok. If the banner is up, treat those two numbers with the same suspicion as the cache-hit numbers — they come from the same rollup.

Diagnosing one unhealthy region

When a single PoP’s bar is amber or red while the rest of the mesh is green, work through the panels in this order:
  1. Check the status pip first. If the PoP is in maintenance, its cache-hit number is expected to sag and needs no action — it is being drained on purpose, and it is still dragging the unweighted mesh average down. If the pip is red with no legend entry, the catalog state itself is the problem.
  2. Check the telemetry banner. A degraded rollup makes every number on the page provisional. Resolve that before you chase a region.
  3. Compare against the live relay health panel. The bar covers the last hour. If the live counters look normal, you are most likely looking at the tail of an incident that has already ended, and the bar will recover as the older one-minute buckets age out of the window.
  4. Read that PoP’s viewer count in the fan-out panel. A low hit ratio on a PoP with very few viewers has little effect on origin load, however red the bar looks; a low ratio on the PoP carrying most of the audience is where origin egress is actually going.
  5. Re-check the mesh average. If exactly one PoP moved and the Edge cache hit KPI changed noticeably, that is the unweighted mean at work — a reminder that the KPI is a signal to open the per-PoP panel, not a measure of how much origin traffic you are actually paying for.
  • Telemetry — the metric, trace, and rollup streams behind these numbers
  • Architecture — where the relay mesh sits in the data plane
  • Authentication — relay tokens and the namespace_auth hook that gates edge subscribe