Failover

Failover that is boring on purpose

xserv-f-hub

When a PoP fails, traffic should move before a human opens a ticket. XServ rehearses that path, including DNS and BGP.

Failover: Why it matters on this continent

Design for Failover on letter F starts in São Paulo (GRU). Runbook f-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.

Failover: How XServ runs it day to day

Probes for Failover (f-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.

Failover: What operators should measure

Policy for Failover is versioned as f-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.

Failover: Failure modes we actually see

Capacity notes for f-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.

Failover: How this letter ties to the rest

Handoff from letter F Failover into the rest of the platform uses the same request IDs as the gateway. Runbook f-05 names who accepts the ticket after São Paulo pages out. Santiago does not silently inherit the incident.

Failover: A note on cost and transit

Rollback for f-06 is a documented command, not a hope. Transit is billed as backup while Failover prefers peering in Santiago. If cost spikes, the runbook cuts overflow to Bogotá before touching customer prefixes.

Failover design F-01

Design for Failover on letter F starts in São Paulo (GRU). Runbook f-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.

Failover probe F-02

Probes for Failover (f-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.

Failover policy F-03

Policy for Failover is versioned as f-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.

Failover capacity F-04

Capacity notes for f-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.

Failover handoff F-05

Handoff from letter F Failover into the rest of the platform uses the same request IDs as the gateway. Runbook f-05 names who accepts the ticket after São Paulo pages out. Santiago does not silently inherit the incident.

Failover rollback F-06

Rollback for f-06 is a documented command, not a hope. Transit is billed as backup while Failover prefers peering in Santiago. If cost spikes, the runbook cuts overflow to Bogotá before touching customer prefixes.

Failover peering F-07

Peering for Failover (f-07) prefers the IX fabric that already carries last-mile ISPs in Bogotá. Session counts and prefix limits are on the same page as the Miami backup path.

Failover cache F-08

Cache rules for Failover on f-08 keep language-correct objects near Miami. São Paulo is a sibling cache, not an origin. TTLs are short enough that a bad asset does not live through a weekend.

Failover headers F-09

Headers for f-09 carry the letter, the region (GRU) and a request id. Failover debugging in São Paulo should not require a packet capture in Santiago first.

Failover timeouts F-10

Timeouts on f-10 are tighter than the WAN RTT to Bogotá. Failover in Santiago fails fast and retries once; a third try needs a human because it is no longer a blip.

Failover regions F-11

Region names on f-11 match DNS and tickets: Bogotá / BOG, with Miami as the documented pair. Failover dashboards never say “LATAM” as if it were one RTT.

Failover docs F-12

Public docs for letter F stay aligned with identifier xserv-f-pack. Page f-12 is the Failover slice operators search before the homepage. Edits in Miami show up in São Paulo on the next publish.

Does Failover on letter F share fate with other letters?
The control plane is shared. Data-plane queues for Failover stay isolated, so incident f-01 cannot drain neighbor letters.
Where is Failover measured first?
From Santiago (SCL) toward Bogotá. A global average is not an SLO for letter F.
How fast is rollback for f-03?
The runbook is a command plus a named owner. There is no wait for a weekly change freeze.
Is transit the default path?
No. Peering in Miami is default; transit to São Paulo is overflow and is billed that way.

Pages in this letter