Uptime
Uptime as a calendar of incidents, not a round badge
xserv-u-hub
XServ records every regional interruption with start, end and customer impact so the next design review has facts.
Uptime: Why it matters on this continent
Design for Uptime on letter U starts in São Paulo (GRU). Runbook u-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.
Uptime: How XServ runs it day to day
Probes for Uptime (u-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.
Uptime: What operators should measure
Policy for Uptime is versioned as u-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.
Uptime: Failure modes we actually see
Capacity notes for u-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.
Uptime: How this letter ties to the rest
Handoff from letter U Uptime into the rest of the platform uses the same request IDs as the gateway. Runbook u-05 names who accepts the ticket after São Paulo pages out. Santiago does not silently inherit the incident.
Uptime: A note on cost and transit
Rollback for u-06 is a documented command, not a hope. Transit is billed as backup while Uptime prefers peering in Santiago. If cost spikes, the runbook cuts overflow to Bogotá before touching customer prefixes.
Uptime design U-01
Design for Uptime on letter U starts in São Paulo (GRU). Runbook u-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.
Uptime probe U-02
Probes for Uptime (u-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.
Uptime policy U-03
Policy for Uptime is versioned as u-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.
Uptime capacity U-04
Capacity notes for u-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.
Uptime handoff U-05
Handoff from letter U Uptime into the rest of the platform uses the same request IDs as the gateway. Runbook u-05 names who accepts the ticket after São Paulo pages out. Santiago does not silently inherit the incident.
Uptime rollback U-06
Rollback for u-06 is a documented command, not a hope. Transit is billed as backup while Uptime prefers peering in Santiago. If cost spikes, the runbook cuts overflow to Bogotá before touching customer prefixes.
Uptime peering U-07
Peering for Uptime (u-07) prefers the IX fabric that already carries last-mile ISPs in Bogotá. Session counts and prefix limits are on the same page as the Miami backup path.
Uptime cache U-08
Cache rules for Uptime on u-08 keep language-correct objects near Miami. São Paulo is a sibling cache, not an origin. TTLs are short enough that a bad asset does not live through a weekend.
Uptime headers U-09
Headers for u-09 carry the letter, the region (GRU) and a request id. Uptime debugging in São Paulo should not require a packet capture in Santiago first.
Uptime timeouts U-10
Timeouts on u-10 are tighter than the WAN RTT to Bogotá. Uptime in Santiago fails fast and retries once; a third try needs a human because it is no longer a blip.
Uptime regions U-11
Region names on u-11 match DNS and tickets: Bogotá / BOG, with Miami as the documented pair. Uptime dashboards never say “LATAM” as if it were one RTT.
Uptime docs U-12
Public docs for letter U stay aligned with identifier xserv-u-pack. Page u-12 is the Uptime slice operators search before the homepage. Edits in Miami show up in São Paulo on the next publish.
- Does Uptime on letter U share fate with other letters?
- The control plane is shared. Data-plane queues for Uptime stay isolated, so incident u-01 cannot drain neighbor letters.
- Where is Uptime measured first?
- From Santiago (SCL) toward Bogotá. A global average is not an SLO for letter U.
- How fast is rollback for u-03?
- The runbook is a command plus a named owner. There is no wait for a weekly change freeze.
- Is transit the default path?
- No. Peering in Miami is default; transit to São Paulo is overflow and is billed that way.
Pages in this letter