scryops
menu

SLOs and Error Budgets

Service Level Objectives and error budgets give reliability a quantitative shape — a target, a budget for deviation, and burn rate signals that tell you when to stop shipping and start fixing.

Service Level Objectives (SLOs) translate reliability from a vague aspiration into a number you can measure, alert on, and make engineering decisions against. The error budget — the allowed amount of unreliability — is what makes the number actionable: it turns a compliance question (“are we meeting the SLO?”) into a spend question (“how fast are we spending the budget, and does that change what we ship next?”).

Core Definitions

Service Level Indicator (SLI) — a quantitative measurement of a service behaviour that reflects the user experience. Common SLIs: request success rate, request latency at a given percentile, data freshness lag, throughput.

Service Level Objective (SLO) — a target level of reliability expressed as a percentage over a rolling time window. Examples: “99.9% of requests complete in under 200ms over 30 days”; “99.95% of API calls return successfully over 28 days.”

Error Budget — the complement of the SLO: 100% minus the target. A 99.9% SLO means 0.1% of requests may fail — that 0.1% is the error budget. When it runs out, the team cannot absorb more risk until it recovers. On a rolling window that happens gradually, as old failures age out of the window; on a calendar window it resets on a fixed date.

Burn Rate — the rate at which the error budget is being consumed relative to the pace that would exhaust it exactly at the end of the window. For a 30-day window, a burn rate of 1× means you’ll exhaust the budget in exactly 30 days. A burn rate of 14.4× means you’ll exhaust it in about 2.1 days.

graph TD A[SLI] -->|measured against| B[SLO target] B -->|defines| C[Error Budget] C -->|monitored via| D[Burn Rate] D -->|triggers| E[Alerts and actions]
Fig. — An SLI feeds an SLO target, which defines the error budget; burn rate tracks how fast that budget is spent and triggers alerts and actions.

Error Budget Calculation

For a 99.9% SLO over a 30-day window:

  • Error budget = 100% − 99.9% = 0.1%
  • Allowed downtime = 0.001 × 30 days × 24 hr × 60 min = 43.2 minutes

For a 99.95% SLO over 28 days:

  • Error budget = 0.05%
  • Allowed downtime = 0.0005 × 28 × 24 × 60 = 20.2 minutes
graph TD A[30-day window] --> B[99.9% uptime target] A --> C[0.1% error budget] B --> D[43,157 minutes available] C --> E[43.2 minutes allowed downtime]
Fig. — A 99.9% target over a 30-day window allows just 43.2 of the window’s 43,200 minutes of downtime.

The budget is not a ceiling for incidents — it is a risk allocation. Planned changes, deployments, experiments, and maintenance all draw from the same budget as unplanned failures. A team with budget remaining can move fast; a team at budget must stop and stabilise.

Burn Rate Alerts

A single threshold alert on the SLO catches incidents too slowly. By the time you’ve failed enough to violate a 30-day SLO target, you’ve already had a significant outage. Multi-window, multi-burn-rate alerting detects incidents at different severities while they can still be contained.

Standard thresholds for a 30-day SLO, from the Google SRE Workbook:

Burn RateLong windowShort windowBudget ConsumedAction
14.4×1 hour5 minutes2%Page immediately
6×6 hours30 minutes5%Page
1×3 days6 hours10%Create ticket

At 14.4× burn, the budget exhausts in 30 ÷ 14.4 ≈ 2.1 days. The 1-hour window is short enough to catch fast-moving incidents early; the 6-hour window catches sustained moderate burns that the 1-hour window misses. The ticket tier catches the slow leak: at 1× for three days you’ve spent a tenth of the month’s budget with no headroom left for anything else to go wrong. An alert fires only when both its windows are over the threshold. The short window is what lets it stop firing minutes after the burn does, instead of hours later.

Burn-rate triageBurn-rate triage as two questions. First: is the SLO error budget burning? If not, monitor and log it. If yes, ask how fast it is burning, which sets a proportional response. Three outcomes, least to most severe, each marked by a distinct shape rather than a colour. Marked with a circle: a slow burn, 1 times the sustainable rate or more on the three-day window, weeks of budget left, becomes a ticket to investigate this week. Marked with a triangle: a moderate burn, 6 times on the six-hour window, notifies the team and escalates cross-functionally if on-call cannot resolve it. Marked with a square: a fast burn, 14 times on the one-hour window, pages on-call and is declared a major incident if users are severely hit. Severity is carried by shape, line weight and line style, so the figure reads the same in greyscale and with any colour vision.BURN-RATE TRIAGE// two questions decide the responseIncident detectedSLO budget burning?no → monitor & logyesHow fast is itburning?SLOW · 1× · 3-day window · weeks leftTicket / Slack — investigate this weekMODERATE · 6× · 6h windowNotify the team — escalate if on-call can’t resolveFAST · 14× · 1h windowPage on-call — major incident if users hit hard
Fig. — Two questions — is the budget burning, and how fast — and the response picks itself.
A 14.4× alert on the 1-hour window doesn’t mean the budget is gone in an hour. It means that if the last hour’s error rate held, a 30-day budget would last about 50 hours. The alert window is how long you measured, not how long you have left: days to exhaustion = SLO window days ÷ burn rate.

Defining Good SLIs

An SLI that doesn’t reflect user experience produces an SLO that doesn’t protect users. Good SLIs share four properties:

graph LR A[Good SLI] --> B[User-focused] A --> C[Controllable] A --> D[Measurable] A --> E[Meaningful] B --> F[Reflects user experience] C --> G[Your team can affect it] D --> H[Consistently collectable] E --> I[Correlated with business outcomes]
Fig. — A good SLI satisfies four properties at once: it reflects user experience, your team can affect it, it is consistently collectable, and it correlates with business outcomes.

Prefer SLIs measured at the edge of the system (from the user’s perspective) over internal measurements. A p99 latency measured at the load balancer is a better SLI than p99 measured at a single microservice — it captures the whole user experience, including infrastructure above and below your code.

Which SLIs fit depends on the service type. A batch job or a queue consumer needs different ones from an API. Choosing SLIs for Your Service has the matrix.

Time windows: 28–30 days for most services (aligns with billing cycles, long enough to absorb weekday/weekend variance). 7 days for services with very high change rates where a 30-day window obscures recent trends.

Setting SLO Targets

Set initial targets from observed performance, not aspirational performance:

graph TD A[Measure current performance] --> B[Set target below current level] B --> C[Monitor 2-3 months] C --> D[Gather user feedback] D --> E[Adjust target] E --> F[Review quarterly]
Fig. — Set the initial SLO below observed performance, then tighten it on a quarterly review cycle as the service’s real variance becomes clear.

Setting the initial SLO below current performance gives you a buffer while you learn the service’s baseline. An SLO set at exactly current p99 latency will fire alerts immediately on any regression. Start conservative; tighten as you understand the service’s normal variation.

Error Budget Policy

The error budget only works as a decision-making tool if the policy around it is written down and enforced. Three areas need explicit policy:

graph LR A[Error Budget Policy] --> B[Consumption rules] A --> C[Reset policy] A --> D[Response actions] B --> B1[What counts as a burn?] B --> B2[How is impact measured?] C --> C1[Reset interval] C --> C2[Overrun handling] D --> D1[Alert thresholds] D --> D2[Required actions per threshold]
Fig. — A usable error budget policy defines three things in advance: what counts as consumption, when the budget resets, and what action each threshold requires.

Consumption rules define what draws from the budget: unplanned outages, degraded performance events, and — critically — planned changes that cause errors. A deployment that causes a 10-minute error spike draws from the same budget as an unplanned incident.

Reset policy defines what happens when the budget runs out. The standard response is a feature freeze: no new deployments until the budget recovers or until the team has made targeted reliability improvements. Define this in advance, not during an incident.

Response thresholds sit alongside the burn rate alerts: burn rate says how fast the budget is going, consumption says how much is left. A common ladder is: at 50% budget consumed, increased review; at 75%, no non-critical deployments; at 100%, feature freeze. These thresholds need to be agreed across engineering and product before they’re ever invoked.

Integration with Business Processes

graph LR A[SLO / Error Budget] --> B[Change management] A --> C[Capacity planning] A --> D[Product decisions] B --> B1[Deploy go/no-go] B --> B2[Maintenance windows] C --> C1[Growth projections] C --> C2[Infrastructure investment] D --> D1[Feature vs. reliability priority] D --> D2[Technical debt timing]
Fig. — The error budget feeds three business processes directly: deploy go/no-go decisions, capacity investment timing, and the feature-versus-reliability tradeoff.

Change management — the error budget is the input to deploy decisions. A team with a full budget can move fast; a team at 10% remaining needs to treat each deployment as a risk event. Build error budget remaining into your deploy checklist.

Capacity planning — SLO burn patterns reveal growth pressure. A service whose burn rate trends upward across multiple windows without a corresponding incident is approaching a capacity ceiling. Error budget data feeds infrastructure investment decisions before the ceiling is hit.

Product decisions — when the budget is exhausted, the conversation about whether to ship or fix should already be resolved by policy. The budget makes that conversation objective: not “how important is this feature?” but “do we have budget to absorb this risk?”

See Also

One observability idea, every Tuesday

A short take, one thing to try that week, and a link to the full article. No vendor pitches.

Subscribe on Substack