scryops
menu
01
How-to 5 min read

The Alert Body Template

A copy-paste alert body that answers four questions before the on-call has to ask: what's broken, how fast it's getting worse, what's already been checked, and the first three steps.
02
Guide 6 min read

Alert Correlation: Finding the Signal in the Flood

A single failure in a distributed system can trigger dozens of alerts across every layer it touches. Correlation groups the symptoms back into one cause — so the on-call engineer sees a problem, not a storm.
03
How-to 9 min read

How to Set Up Your First SLO and Burn Rate Alerts

A step-by-step walkthrough: define an SLI, calculate your error budget, write Prometheus recording rules, and wire up multi-window burn rate alerts that page you before users notice.
04
Guide 6 min read

On-Call Procedures: From Page to Postmortem

A page is just the starting gun. What happens between the alert firing and the postmortem closing determines whether your team gets better or just gets tired.
05
Guide 11 min read

Writing Runbooks That Work at 3am

A runbook that's hard to follow under pressure isn't a runbook. It's a liability. Here's the anatomy of one that actually shortens incident response, and how to keep it true.
06
Guide 6 min read

Alert Severity Levels, Rebuilt for Burn Rate

The P0-P4 framework was built for a world of static thresholds. Here's how to reconnect it to SLO burn rates — so severity reflects actual user impact, not arbitrary lines.
07
Article 6 min read

An Alert Without a Next Step Is Just Noise

The alert fires. The on-call is up. Now what? If the answer is 'check the dashboard', the alert isn't finished. The alert body is where the fix starts, or where you lose an hour chasing context.
08
Guide 8 min read

SLOs and Error Budgets

Service Level Objectives and error budgets give reliability a quantitative shape — a target, a budget for deviation, and burn rate signals that tell you when to stop shipping and start fixing.
09
Guide 11 min read

Log-Based Monitoring: Alerting on the Evidence

Logs carry operational state at a resolution metrics can't match. Most teams only open them after something breaks. This guide covers how to query them continuously, turn them into metrics, and alert on what they surface without blowing up cardinality or cost.
10
How-to 8 min read

Set Up Log-Based Alerting with Loki and Grafana

Turn a LogQL query into a Grafana-managed alert rule that fires on error volume, a specific error type or a failing dependency. Covers the query, the rule settings that trip people up, routing to PagerDuty and Slack, and an end-to-end test.
11
Article 5 min read

Alert Fatigue Is an Observability Problem

Every alert that fires is doing its job. That is the problem. The model is wrong, not the thresholds.