stream_pop catalog in Postgres — the registry of PoPs
and where they are — with the stream_pop_cache_1m_q rollup in
ClickHouse, which supplies the last hour of cache-hit and viewer numbers.
The screen re-runs the query every 30 seconds, and the Refresh button
runs it on demand.
How a PoP is registered
PoPs are not created from the console. There is no “add PoP” action on this screen. A row appears in thestream_pop catalog when a relay
deployment registers itself; the console only reads that catalog.
That is why the empty state reads:
No PoPs registered yet — PoPs appear when the relay deploy registers them.An empty mesh is therefore a deployment fact, not a telemetry failure. If you expect PoPs and see none, the question to ask is whether the relay deploy in that environment came up and registered, not whether the analytics pipeline is healthy. The catalog also supplies each PoP’s map coordinates (
map_x, map_y),
which is what places the pins on the Edge relay map. Map pins are
catalog data; the colour of a pin comes from status.
Fields on each PoP row
The response also carries a top-level
telemetry_ok flag. See
When telemetry is unavailable.
What the cache-hit ratio counts
hit_ratio is a per-PoP number computed by the ClickHouse rollup and read
verbatim by the console — the screen does not recompute it. It is the
edge-cache hit share for that PoP, labelled origin offload on the KPI
card: the higher the number, the more of the fan-out that relay absorbed
locally instead of forwarding upstream to origin.
The window is the last hour, assembled from the one-minute buckets in
stream_pop_cache_1m_q. It is not an instantaneous reading. A PoP that
recovered two minutes ago still carries the previous fifty-eight minutes
of misses in its bar.
The Edge cache hit KPI at the top of the screen is the plain
arithmetic mean of hit_ratio across every PoP in the response. It is
not weighted by viewers, and it does not exclude PoPs in maintenance. A
small PoP with a bad ratio moves the headline number exactly as much as
your largest one does, so treat the KPI as a mesh-wide smoke alarm and the
per-PoP bars as the actual diagnosis.
The KPI card and the per-PoP bars use different colour cut-points, which
is worth knowing before you reconcile them:
So a PoP can sit in warn on its bar while the mesh average it feeds is
still green. That is expected, not a bug.
Reading the fan-out ratio
The Fan-out ratio KPI is displayed as1:N, where:
ORIGIN / ingest.live on the left, <N> RELAYS and the summed
viewer count on the right.
Two caveats when you quote this number:
- The denominator is every PoP in the catalog, including any in maintenance. Taking a region out for maintenance lowers the displayed fan-out ratio even though the surviving relays are each doing more work.
- It is a mesh-wide average, so it hides skew. A mesh where one PoP carries
most of the audience and the rest carry a handful shows the same
1:Nas an evenly balanced one. Read the per-PoP viewer counts in the fan-out panel to see the distribution.
— rather than 1:0.
PoP statuses and what they mean for viewers
status comes from the Postgres catalog and drives the pin colour on the
map, the status pip in the fan-out list, and the two counts under the map.
The Active PoPs KPI shows the total number of rows returned, with the
healthy (
active) count in the hint beneath it. When those two numbers
diverge, the difference is PoPs in maintenance plus PoPs in some other
state. A red pip with no corresponding legend entry is the case that most
often goes unnoticed.
Live relay counters versus the rollup
The screen deliberately stacks two sources of truth, in this order:- Relay health (the
RelayHealthpanel, scoped to thestreamnamespace) — a direct scrape of the relay’s own counters, read now. - The PoP KPIs and bars — ClickHouse rollups, which lag behind by about a minute because they are assembled from one-minute buckets.
When telemetry is unavailable
ThepopHealth response carries telemetry_ok. When the ClickHouse side
of the join is unavailable, the catalog still returns PoPs but the
cache-hit numbers cannot be trusted. The screen degrades rather than
showing zeros:
- The Edge cache hit KPI renders
—instead of a percentage, drops its severity colour entirely, and sets its hint totelemetry unavailable. - Every per-PoP cache-hit value in the Cache-hit ratio · per PoP panel
renders
—. - The Telemetry health banner at the top of the page explains the degraded state.
0.0% would claim
that every edge is missing cache, which is a different and far more
alarming statement than “we cannot currently measure it”. Alarm colour on
this screen is reserved for conditions that are genuinely wrong.
The same suppression applies when the mesh is simply empty. With zero
PoPs, the KPI shows — with the hint no traffic yet rather than a red
0.0%. No PoPs means no measurement, not a total cache miss.
Note that the per-PoP bars still draw at their last-read widths while
telemetry_ok is false; only the numeric labels blank out. Read the —
values, not the bar lengths, when the banner is up.
Viewer counts and the fan-out ratio are not suppressed by
telemetry_ok. If the banner is up, treat those two numbers with the same
suspicion as the cache-hit numbers — they come from the same rollup.
Diagnosing one unhealthy region
When a single PoP’s bar is amber or red while the rest of the mesh is green, work through the panels in this order:- Check the status pip first. If the PoP is in
maintenance, its cache-hit number is expected to sag and needs no action — it is being drained on purpose, and it is still dragging the unweighted mesh average down. If the pip is red with no legend entry, the catalog state itself is the problem. - Check the telemetry banner. A degraded rollup makes every number on the page provisional. Resolve that before you chase a region.
- Compare against the live relay health panel. The bar covers the last hour. If the live counters look normal, you are most likely looking at the tail of an incident that has already ended, and the bar will recover as the older one-minute buckets age out of the window.
- Read that PoP’s viewer count in the fan-out panel. A low hit ratio on a PoP with very few viewers has little effect on origin load, however red the bar looks; a low ratio on the PoP carrying most of the audience is where origin egress is actually going.
- Re-check the mesh average. If exactly one PoP moved and the Edge cache hit KPI changed noticeably, that is the unweighted mean at work — a reminder that the KPI is a signal to open the per-PoP panel, not a measure of how much origin traffic you are actually paying for.
Related
- Telemetry — the metric, trace, and rollup streams behind these numbers
- Architecture — where the relay mesh sits in the data plane
- Authentication — relay tokens and the
namespace_authhook that gates edge subscribe

