Lesson 2 / 25
SLIs, SLOs, SLAs and Error Budgets
Turn vague goals into measurable targets and use error budgets to balance speed and stability.
Measuring what users feel
A Service Level Indicator (SLI) is a measurement of user-visible behaviour, usually a ratio of good events to total events: the proportion of HTTP requests that succeed, or that complete in under 300 ms. A Service Level Objective (SLO) is a target for an SLI over a window: "99.9% of checkout requests succeed over 30 days". A Service Level Agreement (SLA) is a contract with customers that specifies consequences, such as service credits, if a promise is missed; SLAs are usually looser than internal SLOs so you get warned before you owe money. The error budget is what the SLO allows to go wrong: at 99.9% over 30 days, 0.1% of requests may fail. While budget remains, teams ship features; when it is spent, they prioritise reliability work. Each extra "nine" costs far more engineering and money, so choose targets users actually need.
What each availability target allows
Allowed downtime if the whole service were down continuously.
target per year per 30 days
---------- -------------- -----------
99% 3.65 days 7.2 hours
99.9% 8.76 hours 43.2 minutes
99.95% 4.38 hours 21.6 minutes
99.99% 52.6 minutes 4.3 minutes
99.999% 5.26 minutes 26 seconds
error budget = (1 - SLO) x total events
99.9% of 50,000,000 monthly requests -> 50,000 requests may failA monthly data pack
An error budget is like a monthly mobile data pack. You can spend it freely on experiments and risky releases, but once it is gone you slow down until the next cycle instead of paying for overages with angry users.
Quick check: With a 99.9% monthly availability SLO and 10 million requests, how many failed requests does the error budget allow?
- 100
- 1,000
- 100,000
- 10,000
Answer
10,000 — 0.1% of 10,000,000 is 10,000 requests.