scryops
menu

On-Call Procedures: From Page to Postmortem

A page is just the starting gun. What happens between the alert firing and the postmortem closing determines whether your team gets better or just gets tired.

On-call is a machine. Like any machine, it either gets designed or it gets improvised — and improvised on-call is the kind that burns people out, misses incidents, and produces postmortems nobody reads. The sections below lay out the components: roles, rotation, triage, handoffs, and postmortems. Each is a gear. None works without the others.

flowchart TD A[Monitor] --> B[Detect] B --> C[Respond] C --> D[Resolve] D --> E[Postmortem] E --> F[Improve] F --> A
Fig. — The postmortem feeds back into monitoring. Without that loop, on-call keeps meeting the same incidents.

The cycle is intentional. The postmortem feeds back into the monitoring layer — new thresholds, new runbook steps, new alerts for failure modes you hadn’t considered. Without that feedback loop, on-call is just reactive. With it, each incident makes the next one shorter.

Roles and Responsibilities

On-call works best with two roles in rotation simultaneously: a primary who owns detection and response, and a secondary who provides backup escalation and covers gaps.

flowchart LR A[Primary On-Call] --> B[Detect and Respond] A --> C[Escalate if Needed] D[Secondary On-Call] --> E[Assist Primary] D --> F[Take Over if
Primary Unavailable]
Fig. — Two roles, written down in advance: the primary owns the response, the secondary backs it up.

The boundary between the roles should be written down, not negotiated during an incident. Common conventions:

  • Primary acknowledges each page within its severity’s target (under 5 minutes for P0, under 30 for P1, per Alert Severity Levels); secondary covers if primary doesn’t acknowledge within a grace window
  • Primary drives communication on small incidents — status page updates, stakeholder pings, postmortem authorship. On a P0, split the jobs: an incident commander coordinates, the primary works the problem, and someone else owns updates. One person fixing and narrating at once does both badly
  • Secondary does not self-activate; primary explicitly hands off or requests backup
  • Both roles rotate on the same schedule so secondary experience is real, not theoretical

Document the expectations in the team runbook. Engineers who haven’t read the expectations before they’re on-call will invent their own, usually incorrectly.

Rotation Schedule

A fair rotation balances on-call load across the team and accounts for time zones, weekends, and holidays. The standard pattern is a weekly rotation with a one-week offset between the two roles: this week’s secondary becomes next week’s primary.

WeekPrimarySecondary
1Engineer AEngineer B
2Engineer BEngineer C
3Engineer CEngineer D
4Engineer DEngineer A

The offset is the point. Whoever takes over as primary has just spent a week as secondary, watching the same pages, so they start the week with context instead of a cold handover.

Practical notes:

  • Publish the schedule at least two rotation cycles in advance
  • Have a lightweight swap process — a Slack thread with ack from the team lead is enough; bureaucratic swap processes get ignored under deadline pressure
  • Track holiday coverage explicitly; do not assume “the calendar handles it”
  • If the team is globally distributed across more than two time zones, a follow-the-sun model reduces unsociable hours — but requires clean handoff documentation at each shift boundary

Alert Routing by Severity

Not every alert warrants a page. Route based on the urgency of required action, not on the technical severity of the condition.

flowchart LR A[Alert fires] --> B{Severity} B --> C[P0 / P1: page via PagerDuty
immediate response, 24/7] B --> D[P2: incident channel
business-hours response] B --> E[P3: ticket
handled in the sprint] B --> F[P4: documentation or dashboard
no notification]
Fig. — Route by how urgently someone has to act. Only P0 and P1 wake anyone.

The alert routing policy and severity definitions are covered in Alert Severity Levels and Alert Design Principles. The key constraint: if an alert fires and no action is required, it should not be in the paging channel. Every page trains the on-call engineer on what a page means. Page noise is learned helplessness.

Incident Response

When a page fires, the first step is triage — establishing severity before committing resources. The triage decision determines who gets engaged, how fast, and what communication channels open.

flowchart TD A[Incident Detected] --> B[Assess Severity] B --> C{Severity Level} C -->|P0| D[Activate Incident
Response: bridge,
incident commander,
status page update] C -->|P1| E[Investigate and
Mitigate: one
on-call engineer
starts, team pulled
in if not contained] C -->|P2| F[Resolve in
Business Hours
from the incident
channel, nobody
woken up] D --> G[Communicate to
stakeholders;
escalate if
unresolved in SLA] G --> I[Resolve] E --> I F --> I I --> J[Postmortem]
Fig. — Severity decides who gets engaged and how fast. Every path still ends in a postmortem.

Keep the triage step deliberate and short — under five minutes. The most common triage mistake is jumping to mitigation before establishing severity, which leads to P0-level urgency applied to a P2 problem (or the reverse).

Five questions get you to a severity fast:

  • Who’s affected? All users, a region, one customer tier, or nobody yet?
  • How fast is the error budget burning? The burn rate on the alert usually answers this before you’ve opened a dashboard. See SLOs and Error Budgets.
  • Is revenue or data at risk? Failed payments and data loss raise the severity regardless of scope.
  • Is there a workaround? A usable workaround is often the difference between P1 and P2.
  • Is it getting worse? A spreading failure gets the higher severity now, not after it has spread.

Once the severity is set, communication follows one rule: every update says what’s affected, what you’re doing, and when the next update comes, and then you keep that promise. A status page that goes quiet for an hour reads as an outage nobody is handling, even when the fix is ten minutes away.

Handoffs and Escalation

Incidents that span shift boundaries require explicit handoff. A handoff without documentation is a context wipe — the incoming engineer restarts diagnosis from scratch.

flowchart TD A[Incident in Progress] --> B[Handoff at Shift Change] B --> C[Update: current status,
actions taken, next steps,
open hypotheses] C --> D{Resolved?} D -->|Yes| E[Close Incident] D -->|No| F{Escalation Needed?} F -->|Yes| G[Escalate to Next Level
or Incident Commander] F -->|No| H[Continue Working] G --> H H --> D
Fig. — A handoff is a written update, then the same loop the next engineer keeps running until the incident closes.

A minimum handoff note contains:

  • Current status (is the incident actively degrading, stabilised, or in recovery?)
  • What has been tried and ruled out
  • The leading hypothesis
  • Immediate next action
  • Who else is engaged

This takes three minutes to write and saves thirty.

Escalation criteria should be pre-defined, not negotiated mid-incident. Common triggers: incident has been active for N minutes without a mitigation path identified; the failure domain has expanded; external dependencies (payment provider, cloud region) appear to be involved.

Postmortems

Every significant incident produces a postmortem. The purpose is not accountability — it is systemic learning. A postmortem that identifies a person as the root cause has found the wrong root cause.

flowchart TD A[Incident Resolved] --> B[Conduct Blameless Postmortem] B --> C[Reconstruct Timeline] C --> D[Identify Contributing Factors] D --> E[Determine Improvements] E --> F[Assign Action Items with Owners] F --> G[Implement Changes] G --> H[Track Progress] H --> I[Share Learnings with Team] I --> J[Feed Improvements Back
into Monitoring / Runbooks]
Fig. — A postmortem is done when its improvements reach monitoring and the runbooks, not when the meeting ends.

A postmortem action item without an owner and a due date is decorative. Assign each item at the postmortem meeting; review open items at the next team sync. The loop closes when the monitoring layer reflects what you learned — a new alert, a tighter threshold, a runbook step that would have halved the time to detect.

Decide what counts as “significant” before you need to. Google’s SRE book lists common triggers: user-visible downtime or degradation beyond a threshold, any data loss, an on-call intervention such as a rollback or traffic reroute, a resolution time over a threshold, and a monitoring failure, which usually means a human found the incident before the alerts did. Anything below those lines gets a short incident note instead. Anyone can still ask for a full postmortem.

A postmortem doesn’t need a long template. It needs five sections:

  • Summary — what happened and who it affected, in two or three sentences
  • Timeline — timestamped events from first signal to resolution, including when the alert fired relative to when users were hurt
  • Contributing factors — the conditions that let it happen, plural; there is rarely one root cause
  • What went well — the parts of detection and response worth keeping
  • Action items — each with an owner and a due date

See Also

One observability idea, every Tuesday

A short take, one thing to try that week, and a link to the full article. No vendor pitches.

Subscribe on Substack