Failover / f

Failover: probes in Santiago (SCL) · F-02

xserv-f-bed09bab4805

Runbook F-02 for Failover in Santiago (SCL). Marker xserv-f-bed09bab4805. This page covers probes for LATAM operators, not a worldwide average.

Failover: How XServ runs it day to day

Probes for Failover (f-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.

Failover: What operators should measure

Policy for Failover is versioned as f-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.

Failover: Failure modes we actually see

Capacity notes for f-04 assume rainy-season power in Miami and festival peaks toward São Paulo. The failure we actually see is a single uplink, not a cartoon partition of the whole continent. Spare ports and a second provider sit on the same runbook.

Failover cache F-08

Cache rules for Failover on f-08 keep language-correct objects near Miami. São Paulo is a sibling cache, not an origin. TTLs are short enough that a bad asset does not live through a weekend.

Failover headers F-09

Headers for f-09 carry the letter, the region (GRU) and a request id. Failover debugging in São Paulo should not require a packet capture in Santiago first.

Failover timeouts F-10

Timeouts on f-10 are tighter than the WAN RTT to Bogotá. Failover in Santiago fails fast and retries once; a third try needs a human because it is no longer a blip.

Failover regions F-11

Region names on f-11 match DNS and tickets: Bogotá / BOG, with Miami as the documented pair. Failover dashboards never say “LATAM” as if it were one RTT.

Where is Failover measured first?
From Santiago (SCL) toward Bogotá. A global average is not an SLO for letter F.
How fast is rollback for f-03?
The runbook is a command plus a named owner. There is no wait for a weekly change freeze.