Uptime

Uptime as a calendar of incidents, not a round badge

xserv-u-hub

XServ records every regional interruption with start, end and customer impact so the next design review has facts.

Uptime: Why it matters on this continent

Design for Uptime on letter U starts in São Paulo (GRU). Runbook u-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.

Uptime: How XServ runs it day to day

Probes for Uptime (u-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.

Uptime: What operators should measure

Policy for Uptime is versioned as u-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.

Uptime: Failure modes we actually see

Capacity notes for u-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.

Uptime: How this letter ties to the rest

Handoff from letter U Uptime into the rest of the platform uses the same request IDs as the gateway. Runbook u-05 names who accepts the ticket after São Paulo pages out. Santiago does not silently inherit the incident.

Uptime: A note on cost and transit

Rollback for u-06 is a documented command, not a hope. Transit is billed as backup while Uptime prefers peering in Santiago. If cost spikes, the runbook cuts overflow to Bogotá before touching customer prefixes.

Uptime design U-01

Design for Uptime on letter U starts in São Paulo (GRU). Runbook u-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.

Uptime probe U-02

Probes for Uptime (u-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.

Uptime policy U-03

Policy for Uptime is versioned as u-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.

Uptime capacity U-04

Capacity notes for u-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.

Uptime handoff U-05

Handoff from letter U Uptime into the rest of the platform uses the same request IDs as the gateway. Runbook u-05 names who accepts the ticket after São Paulo pages out. Santiago does not silently inherit the incident.

Uptime rollback U-06

Rollback for u-06 is a documented command, not a hope. Transit is billed as backup while Uptime prefers peering in Santiago. If cost spikes, the runbook cuts overflow to Bogotá before touching customer prefixes.

Uptime peering U-07

Peering for Uptime (u-07) prefers the IX fabric that already carries last-mile ISPs in Bogotá. Session counts and prefix limits are on the same page as the Miami backup path.

Uptime cache U-08

Cache rules for Uptime on u-08 keep language-correct objects near Miami. São Paulo is a sibling cache, not an origin. TTLs are short enough that a bad asset does not live through a weekend.

Uptime headers U-09

Headers for u-09 carry the letter, the region (GRU) and a request id. Uptime debugging in São Paulo should not require a packet capture in Santiago first.

Uptime timeouts U-10

Timeouts on u-10 are tighter than the WAN RTT to Bogotá. Uptime in Santiago fails fast and retries once; a third try needs a human because it is no longer a blip.

Uptime regions U-11

Region names on u-11 match DNS and tickets: Bogotá / BOG, with Miami as the documented pair. Uptime dashboards never say “LATAM” as if it were one RTT.

Uptime docs U-12

Public docs for letter U stay aligned with identifier xserv-u-pack. Page u-12 is the Uptime slice operators search before the homepage. Edits in Miami show up in São Paulo on the next publish.

Does Uptime on letter U share fate with other letters?
The control plane is shared. Data-plane queues for Uptime stay isolated, so incident u-01 cannot drain neighbor letters.
Where is Uptime measured first?
From Santiago (SCL) toward Bogotá. A global average is not an SLO for letter U.
How fast is rollback for u-03?
The runbook is a command plus a named owner. There is no wait for a weekly change freeze.
Is transit the default path?
No. Peering in Miami is default; transit to São Paulo is overflow and is billed that way.

Pages in this letter