Failover
Failover that is boring on purpose
xserv-f-hub
When a PoP fails, traffic should move before a human opens a ticket. XServ rehearses that path, including DNS and BGP.
Failover: Why it matters on this continent
Design for Failover on letter F starts in São Paulo (GRU). Runbook f-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.
Failover: How XServ runs it day to day
Probes for Failover (f-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.
Failover: What operators should measure
Policy for Failover is versioned as f-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.
Failover: Failure modes we actually see
Capacity notes for f-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.
Failover: How this letter ties to the rest
Handoff from letter F Failover into the rest of the platform uses the same request IDs as the gateway. Runbook f-05 names who accepts the ticket after São Paulo pages out. Santiago does not silently inherit the incident.
Failover: A note on cost and transit
Rollback for f-06 is a documented command, not a hope. Transit is billed as backup while Failover prefers peering in Santiago. If cost spikes, the runbook cuts overflow to Bogotá before touching customer prefixes.
Failover design F-01
Design for Failover on letter F starts in São Paulo (GRU). Runbook f-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.
Failover probe F-02
Probes for Failover (f-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.
Failover policy F-03
Policy for Failover is versioned as f-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.
Failover capacity F-04
Capacity notes for f-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.
Failover handoff F-05
Handoff from letter F Failover into the rest of the platform uses the same request IDs as the gateway. Runbook f-05 names who accepts the ticket after São Paulo pages out. Santiago does not silently inherit the incident.
Failover rollback F-06
Rollback for f-06 is a documented command, not a hope. Transit is billed as backup while Failover prefers peering in Santiago. If cost spikes, the runbook cuts overflow to Bogotá before touching customer prefixes.
Failover peering F-07
Peering for Failover (f-07) prefers the IX fabric that already carries last-mile ISPs in Bogotá. Session counts and prefix limits are on the same page as the Miami backup path.
Failover cache F-08
Cache rules for Failover on f-08 keep language-correct objects near Miami. São Paulo is a sibling cache, not an origin. TTLs are short enough that a bad asset does not live through a weekend.
Failover headers F-09
Headers for f-09 carry the letter, the region (GRU) and a request id. Failover debugging in São Paulo should not require a packet capture in Santiago first.
Failover timeouts F-10
Timeouts on f-10 are tighter than the WAN RTT to Bogotá. Failover in Santiago fails fast and retries once; a third try needs a human because it is no longer a blip.
Failover regions F-11
Region names on f-11 match DNS and tickets: Bogotá / BOG, with Miami as the documented pair. Failover dashboards never say “LATAM” as if it were one RTT.
Failover docs F-12
Public docs for letter F stay aligned with identifier xserv-f-pack. Page f-12 is the Failover slice operators search before the homepage. Edits in Miami show up in São Paulo on the next publish.
- Does Failover on letter F share fate with other letters?
- The control plane is shared. Data-plane queues for Failover stay isolated, so incident f-01 cannot drain neighbor letters.
- Where is Failover measured first?
- From Santiago (SCL) toward Bogotá. A global average is not an SLO for letter F.
- How fast is rollback for f-03?
- The runbook is a command plus a named owner. There is no wait for a weekly change freeze.
- Is transit the default path?
- No. Peering in Miami is default; transit to São Paulo is overflow and is billed that way.
Pages in this letter