scryops
menu
Guide 6 min read Alerting SLOs On-Call Reliability

Alert Severity Levels, Rebuilt for Burn Rate

The P0-P4 framework was built for a world of static thresholds. Here's how to reconnect it to SLO burn rates — so severity reflects actual user impact, not arbitrary lines.

Severity levels are a promise to the person on the other end of the page. A P0 says: this cannot wait. A P2 says: this matters, but morning is soon enough. A P4 says: someone should know about this, but nobody needs to move.

Most severity frameworks were designed for a world of static thresholds — a metric crosses a line, you assign it a P-level, you page someone or you don’t. That worked when systems were simpler and incidents were more discrete. It fails when you’re trying to catch degradations before they become full outages. Static thresholds tell you what happened. Burn rate tells you what’s happening. A modern severity framework needs to handle both.

Five Levels, One Question

The P0-P4 scale is still useful — not as a prescriptive checklist, but as a shared language for “how fast do we need to move?” The question that matters: how fast is user experience degrading, and how long do we have before it becomes unacceptable?

P0 — The Clock Is Running. Complete outage or severe degradation affecting a large percentage of users. Revenue impact is active and direct. No workaround exists. Fast burn rate: error budget exhausted within roughly two days at the current rate. Response is immediate, all hands, round-the-clock until mitigated.

P1 — Significant and Getting Worse. Partial outage or major feature failure affecting a significant portion of users. Some revenue impact. Limited workarounds. Moderate burn rate: budget will exhaust in roughly five days at the current rate. One on-call engineer starts; pull in a dedicated team if it isn’t contained quickly.

P2 — Real, But Not an Emergency. Non-critical functionality degraded, affecting a small subset of users. Minimal revenue impact. Usable workarounds exist. Slow but noticeable burn: budget gone in roughly two weeks. Worth investigating within the sprint, but nobody needs to be woken up for it. Addressed during business hours.

P3 — Low and Stable. Minor issues: cosmetic, edge-case, or affecting very few users with no meaningful business impact. Burning at 1–2×: the budget runs out in 15–30 days, slowly enough for a ticket. Handled in the normal sprint cycle.

P4 — Informational. No service impact. Purely observational — capacity thresholds, trend signals, things worth knowing about but requiring no action right now.

Severity levels as game threat tiersFive severity levels shown as a descending threat-tier ladder, each driven by SLO burn rate. P0, The Clock Is Running, 14x burn, budget gone in about 2 days, routed to PagerDuty plus a phone call, acknowledge under 5 minutes, 24/7, threat 5 of 5. P1, Significant and worsening, 6x, about 5 days, PagerDuty, under 30 minutes, 24/7, threat 4 of 5. A dashed line marks the boundary: P0 and P1 wake on-call 24/7, everything below is business hours. P2, Real but not urgent, 2x, about 15 days, incident channel, under 2 hours, business hours, threat 3 of 5. P3, Low and stable, 1 to 2x, budget gone in 15 to 30 days, ticket, under 8 hours, business hours, threat 2 of 5. P4, Informational, under 1x, no impact, documentation, next cycle, no page, threat 1 of 5. Each tier's left rail and badge border get fainter and thinner as severity drops.SEVERITY = THREAT LEVELP0–P4 · set by burn rateP0THE CLOCK IS RUNNINGburn 14× / ~2dPagerDuty + CALL · <5m · 24/7THREATP1SIGNIFICANT, WORSENINGburn 6× / ~5dPagerDuty · <30m · 24/7THREAT▲ wake on-call 24/7 / ▼ business hoursP2REAL, NOT URGENTburn 2× / ~15dIncident channel · <2h · business hrsTHREATP3LOW & STABLEburn 1–2× / 15–30dTicket · <8h · business hrsTHREATP4INFORMATIONALburn <1× / no impactDocs · next cycle · no pageTHREAT
Fig. — Severity as a threat ladder: the burn rate sets the tier, and the tier sets who gets woken. P0 and P1 cross the line into 24/7 paging; P2 and below wait for business hours. The mismatch to avoid is a P2 paging someone at 3am.
flowchart LR subgraph assess["Assess impact"] direction TB ui["User impact
and scope"] ~~~ br["SLO burn rate
and budget left"] end subgraph classify["Classify severity"] p0["P0 — immediate
14× burn / ~2 days"] p1["P1 — urgent
6× burn / ~5 days"] p2["P2 — business hours
2× burn / ~15 days"] p34["P3/P4 — low
under 2× burn"] end subgraph route["Route response"] wake["Wake on-call
24/7"] bh["Business hours
response"] ticket["Ticket /
documentation"] end assess --> classify p0 --> wake p1 --> wake p2 --> bh p34 --> ticket
Fig. — The burn rate sets the level, and the level sets who gets woken.

Let the Burn Rate Set the Level

The cleanest way to drive severity from observability data is to wire it to your SLO burn rate, not to individual metric thresholds. Burn rate is how fast you’re spending the error budget relative to the pace that would use it up exactly at the end of the SLO window. At 1×, a 30-day budget lasts 30 days. Days to exhaustion is simply the window divided by the burn rate.

All the numbers below assume a 30-day window. A burn rate of 14× sustained over one hour empties the budget in roughly two days — act now. That’s a P0. A 6× burn over six hours empties it in about five days — P1. A 2× burn sustained over three days is P2: two weeks of runway, real and worth fixing, but not worth anyone’s night. Between 1× and 2× is P3. Below 1× you’re living inside the budget — P4 at most.

The P0 and P1 lines aren’t arbitrary. They’re the page-worthy thresholds from the multiwindow, multi-burn-rate alerting in Google’s SRE Workbook (14.4× over an hour, 6× over six hours), which also pairs each long window with a short one (5 and 30 minutes) so the alert stops firing soon after the burn does. The Workbook tickets anything slower at 1× over three days; the P2/P3 split above just decides which of those tickets is worth this sprint.

The advantage is proportionality: the severity reflects how fast users are actually being hurt, not whatever threshold someone set three years ago and never reviewed.

Who Gets Woken Up, and When

The severity level should directly determine three things: who gets notified, through what channel, and by when. The channel matters as much as the time.

LevelNotify viaAcknowledgmentAfter-hours?
P0PagerDuty + call< 5 minYes, 24/7
P1PagerDuty< 30 minYes, 24/7
P2Incident channel< 2 hoursBusiness hours
P3Ticket< 8 hoursBusiness hours
P4DocumentationNext cycleNo

A P2 that pages someone at 3am is a mismatch between severity and routing — even if the acknowledgment time is fast. Lower-severity alerts should go to passive channels that a team reviews during working hours, not active interrupt channels. Protecting the sleep of your on-call team isn’t just kindness; it’s a reliability investment. Every page that didn’t need a human teaches the human to trust pages a little less — the loop behind alert fatigue.

When in Doubt, Go Higher

Not every incident arrives with a clear severity label. When the initial assessment is uncertain, default to higher severity and downgrade. A P1 that turns out to be a P2 costs the team some sleep. A P2 that should have been a P1 costs users.

Escalate when: the on-call responder can’t resolve or contain within the expected window, the impact is spreading to additional services or customers, or the error budget burn rate increases after the initial response begins.

Document the escalation decision when it happens. Post-incident reviews that can trace why severity was reassigned — and when — surface the systemic gaps that reviews limited to resolution timelines miss entirely.

The Framework That Never Reviews Itself Goes Stale

A severity framework written once and never revisited is a framework that slowly drifts out of alignment with how your system actually behaves. Business priorities change. Services get added. What counts as a P0 revenue impact for a company doing $1M/month is a different calculation from a company doing $10M/month.

Run a quarterly review: pull the last quarter’s incidents, compare assigned severity to actual impact, and adjust the thresholds where they drifted. That’s what keeps the framework calibrated before the gaps turn into outages.

A severity framework tied to real burn rate data will tell you when your P0 threshold is miscalibrated. One written to a whiteboard and never reviewed won’t.

The test worth running. Pull your last 20 P0 and P1 incidents. For each one, check when the SLO burn rate first exceeded 6x — and when the alert actually fired. The gap between those two timestamps is how far behind your alerting is — and every minute of it was spent budget nobody was watching.

One observability idea, every Tuesday

A short take, one thing to try that week, and a link to the full article. No vendor pitches.

Subscribe on Substack