SLA / s
SLA: design in São Paulo (GRU) · S-61
xserv-s-dcdfbb499ffe
Runbook S-61 for SLA in São Paulo (GRU). Marker xserv-s-dcdfbb499ffe. This page covers design for LATAM operators, not a worldwide average.
SLA: Why it matters on this continent
Design for SLA on letter S starts in São Paulo (GRU). Runbook s-01 keeps a single owner, a written rollback, and a traffic split that can be reversed without a global freeze. Neighbors in Santiago only take overflow after the local pool fails a health window.
SLA: How XServ runs it day to day
Probes for SLA (s-02) leave Santiago every 15s toward Bogotá. XServ pages the named owner if loss or RTT crosses the letter budget. A probe never shares a queue with bulk transfers, so a saturated WAN does not hide a dead PoP.
SLA: What operators should measure
Policy for SLA is versioned as s-03. Operators measure error rate, p95 from Bogotá, and time-to-rollback — not a worldwide average. Changes land in Bogotá first, then Miami, with a hold if either region regresses.
SLA peering S-07
Peering for SLA (s-07) prefers the IX fabric that already carries last-mile ISPs in Bogotá. Session counts and prefix limits are on the same page as the Miami backup path.
SLA cache S-08
Cache rules for SLA on s-08 keep language-correct objects near Miami. São Paulo is a sibling cache, not an origin. TTLs are short enough that a bad asset does not live through a weekend.
SLA headers S-09
Headers for s-09 carry the letter, the region (GRU) and a request id. SLA debugging in São Paulo should not require a packet capture in Santiago first.
SLA timeouts S-10
Timeouts on s-10 are tighter than the WAN RTT to Bogotá. SLA in Santiago fails fast and retries once; a third try needs a human because it is no longer a blip.
- Does SLA on letter S share fate with other letters?
- The control plane is shared. Data-plane queues for SLA stay isolated, so incident s-01 cannot drain neighbor letters.
- Where is SLA measured first?
- From Santiago (SCL) toward Bogotá. A global average is not an SLO for letter S.