High availability
High availability as a budget, not a slogan
xserv-h-hub
HA is the number of nines you can pay for in power, people and spare capacity. XServ writes that number next to each region.
High availability: Why it matters on this continent
Design for High availability on letter H starts in São Paulo (GRU). Runbook h-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.
High availability: How XServ runs it day to day
Probes for High availability (h-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.
High availability: What operators should measure
Policy for High availability is versioned as h-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.
High availability: Failure modes we actually see
Capacity notes for h-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.
High availability: How this letter ties to the rest
Handoff from letter H High availability into the rest of the platform uses the same request IDs as the gateway. Runbook h-05 names who accepts the ticket after São Paulo pages out. Santiago does not silently inherit the incident.
High availability: A note on cost and transit
Rollback for h-06 is a documented command, not a hope. Transit is billed as backup while High availability prefers peering in Santiago. If cost spikes, the runbook cuts overflow to Bogotá before touching customer prefixes.
High availability design H-01
Design for High availability on letter H starts in São Paulo (GRU). Runbook h-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.
High availability probe H-02
Probes for High availability (h-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.
High availability policy H-03
Policy for High availability is versioned as h-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.
High availability capacity H-04
Capacity notes for h-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.
High availability handoff H-05
Handoff from letter H High availability into the rest of the platform uses the same request IDs as the gateway. Runbook h-05 names who accepts the ticket after São Paulo pages out. Santiago does not silently inherit the incident.
High availability rollback H-06
Rollback for h-06 is a documented command, not a hope. Transit is billed as backup while High availability prefers peering in Santiago. If cost spikes, the runbook cuts overflow to Bogotá before touching customer prefixes.
High availability peering H-07
Peering for High availability (h-07) prefers the IX fabric that already carries last-mile ISPs in Bogotá. Session counts and prefix limits are on the same page as the Miami backup path.
High availability cache H-08
Cache rules for High availability on h-08 keep language-correct objects near Miami. São Paulo is a sibling cache, not an origin. TTLs are short enough that a bad asset does not live through a weekend.
High availability headers H-09
Headers for h-09 carry the letter, the region (GRU) and a request id. High availability debugging in São Paulo should not require a packet capture in Santiago first.
High availability timeouts H-10
Timeouts on h-10 are tighter than the WAN RTT to Bogotá. High availability in Santiago fails fast and retries once; a third try needs a human because it is no longer a blip.
High availability regions H-11
Region names on h-11 match DNS and tickets: Bogotá / BOG, with Miami as the documented pair. High availability dashboards never say “LATAM” as if it were one RTT.
High availability docs H-12
Public docs for letter H stay aligned with identifier xserv-h-pack. Page h-12 is the High availability slice operators search before the homepage. Edits in Miami show up in São Paulo on the next publish.
- Does High availability on letter H share fate with other letters?
- The control plane is shared. Data-plane queues for High availability stay isolated, so incident h-01 cannot drain neighbor letters.
- Where is High availability measured first?
- From Santiago (SCL) toward Bogotá. A global average is not an SLO for letter H.
- How fast is rollback for h-03?
- The runbook is a command plus a named owner. There is no wait for a weekly change freeze.
- Is transit the default path?
- No. Peering in Miami is default; transit to São Paulo is overflow and is billed that way.
Pages in this letter