SLOs and Error Budgets
Service Level Objectives (SLOs) translate reliability from a vague aspiration into a number you can measure, alert on, and make engineering decisions against. The error budget — the allowed amount of unreliability — is what makes the number actionable: it turns a compliance question (“are we meeting the SLO?”) into a spend question (“how fast are we spending the budget, and does that change what we ship next?”).
Core Definitions
Service Level Indicator (SLI) — a quantitative measurement of a service behaviour that reflects the user experience. Common SLIs: request success rate, request latency at a given percentile, data freshness lag, throughput.
Service Level Objective (SLO) — a target level of reliability expressed as a percentage over a rolling time window. Examples: “99.9% of requests complete in under 200ms over 30 days”; “99.95% of API calls return successfully over 28 days.”
Error Budget — the complement of the SLO: 100% minus the target. A 99.9% SLO means 0.1% of requests may fail — that 0.1% is the error budget. When it runs out, the team cannot absorb more risk until it recovers. On a rolling window that happens gradually, as old failures age out of the window; on a calendar window it resets on a fixed date.
Burn Rate — the rate at which the error budget is being consumed relative to the pace that would exhaust it exactly at the end of the window. For a 30-day window, a burn rate of 1× means you’ll exhaust the budget in exactly 30 days. A burn rate of 14.4× means you’ll exhaust it in about 2.1 days.
Error Budget Calculation
For a 99.9% SLO over a 30-day window:
- Error budget = 100% − 99.9% = 0.1%
- Allowed downtime = 0.001 × 30 days × 24 hr × 60 min = 43.2 minutes
For a 99.95% SLO over 28 days:
- Error budget = 0.05%
- Allowed downtime = 0.0005 × 28 × 24 × 60 = 20.2 minutes
The budget is not a ceiling for incidents — it is a risk allocation. Planned changes, deployments, experiments, and maintenance all draw from the same budget as unplanned failures. A team with budget remaining can move fast; a team at budget must stop and stabilise.
Burn Rate Alerts
A single threshold alert on the SLO catches incidents too slowly. By the time you’ve failed enough to violate a 30-day SLO target, you’ve already had a significant outage. Multi-window, multi-burn-rate alerting detects incidents at different severities while they can still be contained.
Standard thresholds for a 30-day SLO, from the Google SRE Workbook:
| Burn Rate | Long window | Short window | Budget Consumed | Action |
|---|---|---|---|---|
| 14.4× | 1 hour | 5 minutes | 2% | Page immediately |
| 6× | 6 hours | 30 minutes | 5% | Page |
| 1× | 3 days | 6 hours | 10% | Create ticket |
At 14.4× burn, the budget exhausts in 30 ÷ 14.4 ≈ 2.1 days. The 1-hour window is short enough to catch fast-moving incidents early; the 6-hour window catches sustained moderate burns that the 1-hour window misses. The ticket tier catches the slow leak: at 1× for three days you’ve spent a tenth of the month’s budget with no headroom left for anything else to go wrong. An alert fires only when both its windows are over the threshold. The short window is what lets it stop firing minutes after the burn does, instead of hours later.
Defining Good SLIs
An SLI that doesn’t reflect user experience produces an SLO that doesn’t protect users. Good SLIs share four properties:
Prefer SLIs measured at the edge of the system (from the user’s perspective) over internal measurements. A p99 latency measured at the load balancer is a better SLI than p99 measured at a single microservice — it captures the whole user experience, including infrastructure above and below your code.
Which SLIs fit depends on the service type. A batch job or a queue consumer needs different ones from an API. Choosing SLIs for Your Service has the matrix.
Time windows: 28–30 days for most services (aligns with billing cycles, long enough to absorb weekday/weekend variance). 7 days for services with very high change rates where a 30-day window obscures recent trends.
Setting SLO Targets
Set initial targets from observed performance, not aspirational performance:
Setting the initial SLO below current performance gives you a buffer while you learn the service’s baseline. An SLO set at exactly current p99 latency will fire alerts immediately on any regression. Start conservative; tighten as you understand the service’s normal variation.
Error Budget Policy
The error budget only works as a decision-making tool if the policy around it is written down and enforced. Three areas need explicit policy:
Consumption rules define what draws from the budget: unplanned outages, degraded performance events, and — critically — planned changes that cause errors. A deployment that causes a 10-minute error spike draws from the same budget as an unplanned incident.
Reset policy defines what happens when the budget runs out. The standard response is a feature freeze: no new deployments until the budget recovers or until the team has made targeted reliability improvements. Define this in advance, not during an incident.
Response thresholds sit alongside the burn rate alerts: burn rate says how fast the budget is going, consumption says how much is left. A common ladder is: at 50% budget consumed, increased review; at 75%, no non-critical deployments; at 100%, feature freeze. These thresholds need to be agreed across engineering and product before they’re ever invoked.
Integration with Business Processes
Change management — the error budget is the input to deploy decisions. A team with a full budget can move fast; a team at 10% remaining needs to treat each deployment as a risk event. Build error budget remaining into your deploy checklist.
Capacity planning — SLO burn patterns reveal growth pressure. A service whose burn rate trends upward across multiple windows without a corresponding incident is approaching a capacity ceiling. Error budget data feeds infrastructure investment decisions before the ceiling is hit.
Product decisions — when the budget is exhausted, the conversation about whether to ship or fix should already be resolved by policy. The budget makes that conversation objective: not “how important is this feature?” but “do we have budget to absorb this risk?”
See Also
- Choosing SLIs for Your Service — which SLIs fit APIs, queues, pipelines, batch jobs, storage, models and CDNs
- How to Set Up Your First SLO and Burn Rate Alerts — the Prometheus recording and alert rules for all three tiers, step by step
- Alert Severity Levels, Rebuilt for Burn Rate — mapping these burn rate tiers onto P0–P4 and who gets woken up
- Alert Design Principles — how burn rate alerts fit into a broader alerting strategy
One observability idea, every Tuesday
A short take, one thing to try that week, and a link to the full article. No vendor pitches.
Subscribe on Substack