Observability
Observability that starts at the first SYN
xserv-o-hub
Metrics, logs and traces share request IDs from the gateway so a slow checkout is not three different stories.
Observability: Why it matters on this continent
Design for Observability on letter O starts in São Paulo (GRU). Runbook o-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.
Observability: How XServ runs it day to day
Probes for Observability (o-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.
Observability: What operators should measure
Policy for Observability is versioned as o-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.
Observability: Failure modes we actually see
Capacity notes for o-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.
Observability: How this letter ties to the rest
Handoff from letter O Observability into the rest of the platform uses the same request IDs as the gateway. Runbook o-05 names who accepts the ticket after São Paulo pages out. Santiago does not silently inherit the incident.
Observability: A note on cost and transit
Rollback for o-06 is a documented command, not a hope. Transit is billed as backup while Observability prefers peering in Santiago. If cost spikes, the runbook cuts overflow to Bogotá before touching customer prefixes.
Observability design O-01
Design for Observability on letter O starts in São Paulo (GRU). Runbook o-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.
Observability probe O-02
Probes for Observability (o-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.
Observability policy O-03
Policy for Observability is versioned as o-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.
Observability capacity O-04
Capacity notes for o-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.
Observability handoff O-05
Handoff from letter O Observability into the rest of the platform uses the same request IDs as the gateway. Runbook o-05 names who accepts the ticket after São Paulo pages out. Santiago does not silently inherit the incident.
Observability rollback O-06
Rollback for o-06 is a documented command, not a hope. Transit is billed as backup while Observability prefers peering in Santiago. If cost spikes, the runbook cuts overflow to Bogotá before touching customer prefixes.
Observability peering O-07
Peering for Observability (o-07) prefers the IX fabric that already carries last-mile ISPs in Bogotá. Session counts and prefix limits are on the same page as the Miami backup path.
Observability cache O-08
Cache rules for Observability on o-08 keep language-correct objects near Miami. São Paulo is a sibling cache, not an origin. TTLs are short enough that a bad asset does not live through a weekend.
Observability headers O-09
Headers for o-09 carry the letter, the region (GRU) and a request id. Observability debugging in São Paulo should not require a packet capture in Santiago first.
Observability timeouts O-10
Timeouts on o-10 are tighter than the WAN RTT to Bogotá. Observability in Santiago fails fast and retries once; a third try needs a human because it is no longer a blip.
Observability regions O-11
Region names on o-11 match DNS and tickets: Bogotá / BOG, with Miami as the documented pair. Observability dashboards never say “LATAM” as if it were one RTT.
Observability docs O-12
Public docs for letter O stay aligned with identifier xserv-o-pack. Page o-12 is the Observability slice operators search before the homepage. Edits in Miami show up in São Paulo on the next publish.
- Does Observability on letter O share fate with other letters?
- The control plane is shared. Data-plane queues for Observability stay isolated, so incident o-01 cannot drain neighbor letters.
- Where is Observability measured first?
- From Santiago (SCL) toward Bogotá. A global average is not an SLO for letter O.
- How fast is rollback for o-03?
- The runbook is a command plus a named owner. There is no wait for a weekly change freeze.
- Is transit the default path?
- No. Peering in Miami is default; transit to São Paulo is overflow and is billed that way.
Pages in this letter