scryops
menu
01
Guide 6 min read

Alert Correlation: Finding the Signal in the Flood

A single failure in a distributed system can trigger dozens of alerts across every layer it touches. Correlation groups the symptoms back into one cause — so the on-call engineer sees a problem, not a storm.
02
Guide 9 min read

Choosing SLIs for Your Service: A Practitioner's Matrix

Availability and latency are the obvious SLIs. But they don't fit every service type. This guide provides SLI selection frameworks for APIs, data pipelines, batch jobs, storage systems, event-driven services, and more.
03
Guide 6 min read

On-Call Procedures: From Page to Postmortem

A page is just the starting gun. What happens between the alert firing and the postmortem closing determines whether your team gets better or just gets tired.
04
Q&A 3 min read

What is an SLI and how do you choose one?

An SLI is a quantitative measure of service behaviour from the user's perspective. Choose the metric that, when it degrades, users notice — not server-side proxies that feel measurable but don't map to user experience.
05
Guide 11 min read

Writing Runbooks That Work at 3am

A runbook that's hard to follow under pressure isn't a runbook. It's a liability. Here's the anatomy of one that actually shortens incident response, and how to keep it true.
06
Guide 6 min read

Alert Severity Levels, Rebuilt for Burn Rate

The P0-P4 framework was built for a world of static thresholds. Here's how to reconnect it to SLO burn rates — so severity reflects actual user impact, not arbitrary lines.
07
Article 6 min read

An Alert Without a Next Step Is Just Noise

The alert fires. The on-call is up. Now what? If the answer is 'check the dashboard', the alert isn't finished. The alert body is where the fix starts, or where you lose an hour chasing context.
08
Guide 8 min read

SLOs and Error Budgets

Service Level Objectives and error budgets give reliability a quantitative shape — a target, a budget for deviation, and burn rate signals that tell you when to stop shipping and start fixing.
09
Q&A 3 min read

What is synthetic monitoring, and how does it differ from RUM?

Synthetic monitoring runs scripted tests on a schedule. RUM captures what real users actually experience. They answer different questions, and you need both.
10
Guide 11 min read

Async Logging: Keeping Your Application Threads Free

A synchronous log write makes the request thread wait on disk or network I/O. Async logging hands that work to a background thread, and quietly adds a queue that can fill up, drop records and lose them at shutdown. How to size it, watch it and flush it, with Serilog, the OpenTelemetry SDK and Python.
11
Guide 9 min read

High-Throughput Logging: Keeping the Hot Path Fast

At hundreds of thousands of requests per second, the logging call itself becomes the bottleneck. Async channels, pooling, batching, and circuit breakers keep log I/O off the request thread.