High availability

High availability as a budget, not a slogan

xserv-h-hub

HA is the number of nines you can pay for in power, people and spare capacity. XServ writes that number next to each region.

High availability: Why it matters on this continent

Design for High availability on letter H starts in São Paulo (GRU). Runbook h-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.

High availability: How XServ runs it day to day

Probes for High availability (h-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.

High availability: What operators should measure

Policy for High availability is versioned as h-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.

High availability: Failure modes we actually see

Capacity notes for h-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.

High availability: How this letter ties to the rest

Handoff from letter H High availability into the rest of the platform uses the same request IDs as the gateway. Runbook h-05 names who accepts the ticket after São Paulo pages out. Santiago does not silently inherit the incident.

High availability: A note on cost and transit

Rollback for h-06 is a documented command, not a hope. Transit is billed as backup while High availability prefers peering in Santiago. If cost spikes, the runbook cuts overflow to Bogotá before touching customer prefixes.

High availability design H-01

Design for High availability on letter H starts in São Paulo (GRU). Runbook h-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.

High availability probe H-02

Probes for High availability (h-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.

High availability policy H-03

Policy for High availability is versioned as h-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.

High availability capacity H-04

Capacity notes for h-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.

High availability handoff H-05

Handoff from letter H High availability into the rest of the platform uses the same request IDs as the gateway. Runbook h-05 names who accepts the ticket after São Paulo pages out. Santiago does not silently inherit the incident.

High availability rollback H-06

Rollback for h-06 is a documented command, not a hope. Transit is billed as backup while High availability prefers peering in Santiago. If cost spikes, the runbook cuts overflow to Bogotá before touching customer prefixes.

High availability peering H-07

Peering for High availability (h-07) prefers the IX fabric that already carries last-mile ISPs in Bogotá. Session counts and prefix limits are on the same page as the Miami backup path.

High availability cache H-08

Cache rules for High availability on h-08 keep language-correct objects near Miami. São Paulo is a sibling cache, not an origin. TTLs are short enough that a bad asset does not live through a weekend.

High availability headers H-09

Headers for h-09 carry the letter, the region (GRU) and a request id. High availability debugging in São Paulo should not require a packet capture in Santiago first.

High availability timeouts H-10

Timeouts on h-10 are tighter than the WAN RTT to Bogotá. High availability in Santiago fails fast and retries once; a third try needs a human because it is no longer a blip.

High availability regions H-11

Region names on h-11 match DNS and tickets: Bogotá / BOG, with Miami as the documented pair. High availability dashboards never say “LATAM” as if it were one RTT.

High availability docs H-12

Public docs for letter H stay aligned with identifier xserv-h-pack. Page h-12 is the High availability slice operators search before the homepage. Edits in Miami show up in São Paulo on the next publish.

Does High availability on letter H share fate with other letters?
The control plane is shared. Data-plane queues for High availability stay isolated, so incident h-01 cannot drain neighbor letters.
Where is High availability measured first?
From Santiago (SCL) toward Bogotá. A global average is not an SLO for letter H.
How fast is rollback for h-03?
The runbook is a command plus a named owner. There is no wait for a weekly change freeze.
Is transit the default path?
No. Peering in Miami is default; transit to São Paulo is overflow and is billed that way.

Pages in this letter