in-depth technical references
guides
Deep technical references — heavy with multi-language code and diagrams. Start at the top; each one turns a system you can't see into one you can.
01
OpenTelemetry: What It Is and How It Fits Together
OpenTelemetry is a single instrumentation layer that produces traces, metrics, and logs in a vendor-neutral format. This guide explains what each signal is for, how the SDK and Collector relate, and where to go next.
02
Alert Correlation: Finding the Signal in the Flood
A single failure in a distributed system can trigger dozens of alerts across every layer it touches. Correlation groups the symptoms back into one cause — so the on-call engineer sees a problem, not a storm.
03
Choosing SLIs for Your Service: A Practitioner's Matrix
Availability and latency are the obvious SLIs. But they don't fit every service type. This guide provides SLI selection frameworks for APIs, data pipelines, batch jobs, storage systems, event-driven services, and more.
04
Context Propagation: How Distributed Traces Stay Connected Across Services
A distributed trace is only as complete as its weakest propagation link. One hop that drops the context and the trace splits in two. W3C Trace Context and Baggage, the propagator settings that matter, and the places context gets lost — between services and inside them — in .NET, Java, Go, Python and Node.
05
eBPF Continuous Profiling: A Practical Guide
Profile every process on a node with one DaemonSet and no code changes. How eBPF profilers work, which tool to pick, a tested Parca deployment, and the privileges, storage and trace-linking limits you need to plan around.
06
Log Levels: When to Whisper, Speak, or Shout
Log levels are the emotional register of your system's voice. How to use ERROR, WARN, INFO, DEBUG and TRACE consistently, how they map onto OpenTelemetry's severity numbers, and what each one costs you.
07
Logging Foundations
What logging is for, how it fits next to metrics and traces, and what to log, where and how. The mental model to have before touching a logging framework, and the starting point for the rest of the logging guides.
08
On-Call Procedures: From Page to Postmortem
A page is just the starting gun. What happens between the alert firing and the postmortem closing determines whether your team gets better or just gets tired.
09
Structured Logging: Teaching Machines to Read
Logs were designed for humans grepping text files at 2am. Now they also have to feed query engines, correlation and anomaly detection. Here's what that changes about what you write, and which field names to use.
10
Telemetry Data Sovereignty: Where Your Data Lives Matters
A system that spans continents produces telemetry that spans legal jurisdictions. Here's how to keep traces and logs where the law wants them, and still see your whole system.
11
Writing Runbooks That Work at 3am
A runbook that's hard to follow under pressure isn't a runbook. It's a liability. Here's the anatomy of one that actually shortens incident response, and how to keep it true.
12
Alert Severity Levels, Rebuilt for Burn Rate
The P0-P4 framework was built for a world of static thresholds. Here's how to reconnect it to SLO burn rates — so severity reflects actual user impact, not arbitrary lines.
13
SLOs and Error Budgets
Service Level Objectives and error budgets give reliability a quantitative shape — a target, a budget for deviation, and burn rate signals that tell you when to stop shipping and start fixing.
14
Async Logging: Keeping Your Application Threads Free
A synchronous log write makes the request thread wait on disk or network I/O. Async logging hands that work to a background thread, and quietly adds a queue that can fill up, drop records and lose them at shutdown. How to size it, watch it and flush it, with Serilog, the OpenTelemetry SDK and Python.
15
Common Logging Pitfalls and How to Avoid Them
The same logging mistakes turn up in every team and every stack: inconsistent field names, values buried in message strings, missing trace context, personal data and secrets in error logs, and loops that log the same thing ten thousand times. Here is where to look and what to fix.
16
Log Context Enrichment: Adding Meaning to Your Events
Enrichment turns isolated log records into connected business events. Here is the architecture that makes it work — static resource attributes, background-refreshed caches, and per-request scopes — without taxing the request path.
17
Implementing Audit Trails with OpenTelemetry
An audit trail is not a log. It's a tamper-evident, time-ordered record of who did what, when, and why. Most teams build this wrong. Here is how to do it correctly using OpenTelemetry and append-only storage.
18
Log-Based Monitoring: Alerting on the Evidence
Logs carry operational state at a resolution metrics can't match. Most teams only open them after something breaks. This guide covers how to query them continuously, turn them into metrics, and alert on what they surface without blowing up cardinality or cost.
19
Observability Under Compliance: GDPR, HIPAA, SOC 2, and PCI DSS
Regulated industries need observability too. A guide to building telemetry pipelines that satisfy GDPR, HIPAA, SOC 2, and PCI DSS requirements — covering data minimisation, retention mandates, audit trails, and what each framework actually requires.
20
Data Masking in Telemetry: The Art of Safe Transformation
Telemetry data is just as risky for PII as any database. Here's how to turn sensitive fields into safe, useful signals: hashing, tokenising, coarsening, and picking the right tool for the job.
21
Distributed Logging: Ten Services, One Story
When a request crosses ten services, you get ten log streams that share nothing but a timestamp you can't trust. How to collect them on every node, carry the trace ID through, ship them through the OpenTelemetry Collector, and notice when lines go missing.
22
High-Throughput Logging: Keeping the Hot Path Fast
At hundreds of thousands of requests per second, the logging call itself becomes the bottleneck. Async channels, pooling, batching, and circuit breakers keep log I/O off the request thread.
23
High-Throughput Logging: Sampling, Collectors, and the Wire
At 1.5 million log events per second you cannot keep, batch, or ship everything the way you did at moderate scale. Content-aware sampling, OTel exporter tuning, Collector-side batching, and cheaper bytes on the wire.
24
Your Traces Are Leaking User Data
Every OTel span that includes a customer email, shipping address, or payment token is a GDPR audit waiting to happen. The fix isn't application code; it's a Collector pipeline.