scryops
menu
01
Guide 8 min read

OpenTelemetry: What It Is and How It Fits Together

OpenTelemetry is a single instrumentation layer that produces traces, metrics, and logs in a vendor-neutral format. This guide explains what each signal is for, how the SDK and Collector relate, and where to go next.
02
Guide 6 min read

Alert Correlation: Finding the Signal in the Flood

A single failure in a distributed system can trigger dozens of alerts across every layer it touches. Correlation groups the symptoms back into one cause — so the on-call engineer sees a problem, not a storm.
03
Guide 9 min read

Choosing SLIs for Your Service: A Practitioner's Matrix

Availability and latency are the obvious SLIs. But they don't fit every service type. This guide provides SLI selection frameworks for APIs, data pipelines, batch jobs, storage systems, event-driven services, and more.
04
Guide 11 min read

Context Propagation: How Distributed Traces Stay Connected Across Services

A distributed trace is only as complete as its weakest propagation link. One hop that drops the context and the trace splits in two. W3C Trace Context and Baggage, the propagator settings that matter, and the places context gets lost — between services and inside them — in .NET, Java, Go, Python and Node.
05
How-to 8 min read

How to Instrument a .NET Service with OpenTelemetry

Add OpenTelemetry to an ASP.NET Core service: traces, metrics and logs in one setup block, manual spans and metrics for business logic, Serilog, and the zero-code agent for services you can't change. All verified against a local Collector.
06
How-to 7 min read

How to Instrument a Java Spring Boot Service with OpenTelemetry

Instrument a Spring Boot service with the OpenTelemetry Java agent or the Spring Boot starter: traces, metrics and logs with no code, then your own spans and metrics, the Micrometer bridge and log correlation. All verified against a local Collector.
07
Guide 12 min read

Log Levels: When to Whisper, Speak, or Shout

Log levels are the emotional register of your system's voice. How to use ERROR, WARN, INFO, DEBUG and TRACE consistently, how they map onto OpenTelemetry's severity numbers, and what each one costs you.
08
Guide 9 min read

Logging Foundations

What logging is for, how it fits next to metrics and traces, and what to log, where and how. The mental model to have before touching a logging framework, and the starting point for the rest of the logging guides.
09
Guide 6 min read

On-Call Procedures: From Page to Postmortem

A page is just the starting gun. What happens between the alert firing and the postmortem closing determines whether your team gets better or just gets tired.
10
Guide 8 min read

Structured Logging: Teaching Machines to Read

Logs were designed for humans grepping text files at 2am. Now they also have to feed query engines, correlation and anomaly detection. Here's what that changes about what you write, and which field names to use.
11
Guide 11 min read

Telemetry Data Sovereignty: Where Your Data Lives Matters

A system that spans continents produces telemetry that spans legal jurisdictions. Here's how to keep traces and logs where the law wants them, and still see your whole system.
12
Q&A 3 min read

What is an SLI and how do you choose one?

An SLI is a quantitative measure of service behaviour from the user's perspective. Choose the metric that, when it degrades, users notice — not server-side proxies that feel measurable but don't map to user experience.
13
Article 6 min read

An Alert Without a Next Step Is Just Noise

The alert fires. The on-call is up. Now what? If the answer is 'check the dashboard', the alert isn't finished. The alert body is where the fix starts, or where you lose an hour chasing context.
14
Guide 8 min read

SLOs and Error Budgets

Service Level Objectives and error budgets give reliability a quantitative shape — a target, a budget for deviation, and burn rate signals that tell you when to stop shipping and start fixing.
15
Article 7 min read

Your Sampling Strategy Is Lying to You

A flat 5% sampling rate sounds like a sensible trade between cost and coverage. It isn't. A random slice of your traffic is mostly the requests you'll never look at, and it throws away the rare ones you need at the same rate.
16
Article 7 min read

Your Tagging Standard Is a Wiki Page. That's Why It Doesn't Work.

Every team has a tagging standard. Most of them live in a wiki, enforced by nobody, remembered by almost nobody, and invisible to the CI pipeline. Open Policy Agent fixes the root cause.
17
Article 6 min read

The dashboard was green, but the request was broken.

Metrics tell you how the crowd is doing. Logs tell you what one service saw. A trace tells you what one request went through, and at 2 a.m. that is usually the question you are asking.
18
Q&A 3 min read

What is synthetic monitoring, and how does it differ from RUM?

Synthetic monitoring runs scripted tests on a schedule. RUM captures what real users actually experience. They answer different questions, and you need both.
19
Guide 11 min read

Async Logging: Keeping Your Application Threads Free

A synchronous log write makes the request thread wait on disk or network I/O. Async logging hands that work to a background thread, and quietly adds a queue that can fill up, drop records and lose them at shutdown. How to size it, watch it and flush it, with Serilog, the OpenTelemetry SDK and Python.
20
Guide 9 min read

Common Logging Pitfalls and How to Avoid Them

The same logging mistakes turn up in every team and every stack: inconsistent field names, values buried in message strings, missing trace context, personal data and secrets in error logs, and loops that log the same thing ten thousand times. Here is where to look and what to fix.
21
Guide 8 min read

Log Context Enrichment: Adding Meaning to Your Events

Enrichment turns isolated log records into connected business events. Here is the architecture that makes it work — static resource attributes, background-refreshed caches, and per-request scopes — without taxing the request path.
22
Guide 11 min read

Log-Based Monitoring: Alerting on the Evidence

Logs carry operational state at a resolution metrics can't match. Most teams only open them after something breaks. This guide covers how to query them continuously, turn them into metrics, and alert on what they surface without blowing up cardinality or cost.
23
Guide 10 min read

Observability Under Compliance: GDPR, HIPAA, SOC 2, and PCI DSS

Regulated industries need observability too. A guide to building telemetry pipelines that satisfy GDPR, HIPAA, SOC 2, and PCI DSS requirements — covering data minimisation, retention mandates, audit trails, and what each framework actually requires.
24
How-to 8 min read

Set Up Log-Based Alerting with Loki and Grafana

Turn a LogQL query into a Grafana-managed alert rule that fires on error volume, a specific error type or a failing dependency. Covers the query, the rule settings that trip people up, routing to PagerDuty and Slack, and an end-to-end test.
25
Guide 8 min read

Data Masking in Telemetry: The Art of Safe Transformation

Telemetry data is just as risky for PII as any database. Here's how to turn sensitive fields into safe, useful signals: hashing, tokenising, coarsening, and picking the right tool for the job.
26
Guide 10 min read

Distributed Logging: Ten Services, One Story

When a request crosses ten services, you get ten log streams that share nothing but a timestamp you can't trust. How to collect them on every node, carry the trace ID through, ship them through the OpenTelemetry Collector, and notice when lines go missing.
27
Guide 9 min read

High-Throughput Logging: Keeping the Hot Path Fast

At hundreds of thousands of requests per second, the logging call itself becomes the bottleneck. Async channels, pooling, batching, and circuit breakers keep log I/O off the request thread.
28
Article 6 min read

Observability vs. Monitoring: Why the Distinction Matters

Monitoring tells you when something you predicted goes wrong. Observability lets you work out what's happening when it's something you didn't. You need both, and the gap between them is where incidents drag on.
29
Article 5 min read

The Evolution of System Understanding

From grepping one log file to querying wide, trace-linked events: how the questions we can ask a running system changed when monoliths split apart, and why OpenTelemetry had to exist.
30
Article 5 min read

Alert Fatigue Is an Observability Problem

Every alert that fires is doing its job. That is the problem. The model is wrong, not the thresholds.
31
Guide 8 min read

Your Traces Are Leaking User Data

Every OTel span that includes a customer email, shipping address, or payment token is a GDPR audit waiting to happen. The fix isn't application code; it's a Collector pipeline.
32
Article 6 min read

Observability 1.0 meant forensics. Observability 2.0 means prevention.

Observability 1.0 taught us to look backward. Observability 2.0 asks us to look forward. Most teams haven’t made that shift yet. That’s why I named this site after a medieval divination practice.