Tagged: Observability
01
OpenTelemetry: What It Is and How It Fits Together
OpenTelemetry is a single instrumentation layer that produces traces, metrics, and logs in a vendor-neutral format. This guide explains what each signal is for, how the SDK and Collector relate, and where to go next.
02
Alert Correlation: Finding the Signal in the Flood
A single failure in a distributed system can trigger dozens of alerts across every layer it touches. Correlation groups the symptoms back into one cause — so the on-call engineer sees a problem, not a storm.
03
Choosing SLIs for Your Service: A Practitioner's Matrix
Availability and latency are the obvious SLIs. But they don't fit every service type. This guide provides SLI selection frameworks for APIs, data pipelines, batch jobs, storage systems, event-driven services, and more.
04
Context Propagation: How Distributed Traces Stay Connected Across Services
A distributed trace is only as complete as its weakest propagation link. One hop that drops the context and the trace splits in two. W3C Trace Context and Baggage, the propagator settings that matter, and the places context gets lost — between services and inside them — in .NET, Java, Go, Python and Node.
05
How to Instrument a .NET Service with OpenTelemetry
Add OpenTelemetry to an ASP.NET Core service: traces, metrics and logs in one setup block, manual spans and metrics for business logic, Serilog, and the zero-code agent for services you can't change. All verified against a local Collector.
06
How to Instrument a Java Spring Boot Service with OpenTelemetry
Instrument a Spring Boot service with the OpenTelemetry Java agent or the Spring Boot starter: traces, metrics and logs with no code, then your own spans and metrics, the Micrometer bridge and log correlation. All verified against a local Collector.
07
Log Levels: When to Whisper, Speak, or Shout
Log levels are the emotional register of your system's voice. How to use ERROR, WARN, INFO, DEBUG and TRACE consistently, how they map onto OpenTelemetry's severity numbers, and what each one costs you.
08
Logging Foundations
What logging is for, how it fits next to metrics and traces, and what to log, where and how. The mental model to have before touching a logging framework, and the starting point for the rest of the logging guides.
09
On-Call Procedures: From Page to Postmortem
A page is just the starting gun. What happens between the alert firing and the postmortem closing determines whether your team gets better or just gets tired.
10
Structured Logging: Teaching Machines to Read
Logs were designed for humans grepping text files at 2am. Now they also have to feed query engines, correlation and anomaly detection. Here's what that changes about what you write, and which field names to use.
11
Telemetry Data Sovereignty: Where Your Data Lives Matters
A system that spans continents produces telemetry that spans legal jurisdictions. Here's how to keep traces and logs where the law wants them, and still see your whole system.
12
What is an SLI and how do you choose one?
An SLI is a quantitative measure of service behaviour from the user's perspective. Choose the metric that, when it degrades, users notice — not server-side proxies that feel measurable but don't map to user experience.
13
An Alert Without a Next Step Is Just Noise
The alert fires. The on-call is up. Now what? If the answer is 'check the dashboard', the alert isn't finished. The alert body is where the fix starts, or where you lose an hour chasing context.
14
SLOs and Error Budgets
Service Level Objectives and error budgets give reliability a quantitative shape — a target, a budget for deviation, and burn rate signals that tell you when to stop shipping and start fixing.
15
Your Sampling Strategy Is Lying to You
A flat 5% sampling rate sounds like a sensible trade between cost and coverage. It isn't. A random slice of your traffic is mostly the requests you'll never look at, and it throws away the rare ones you need at the same rate.
16
Your Tagging Standard Is a Wiki Page. That's Why It Doesn't Work.
Every team has a tagging standard. Most of them live in a wiki, enforced by nobody, remembered by almost nobody, and invisible to the CI pipeline. Open Policy Agent fixes the root cause.
17
The dashboard was green, but the request was broken.
Metrics tell you how the crowd is doing. Logs tell you what one service saw. A trace tells you what one request went through, and at 2 a.m. that is usually the question you are asking.
18
What is synthetic monitoring, and how does it differ from RUM?
Synthetic monitoring runs scripted tests on a schedule. RUM captures what real users actually experience. They answer different questions, and you need both.
19
Async Logging: Keeping Your Application Threads Free
A synchronous log write makes the request thread wait on disk or network I/O. Async logging hands that work to a background thread, and quietly adds a queue that can fill up, drop records and lose them at shutdown. How to size it, watch it and flush it, with Serilog, the OpenTelemetry SDK and Python.
20
Common Logging Pitfalls and How to Avoid Them
The same logging mistakes turn up in every team and every stack: inconsistent field names, values buried in message strings, missing trace context, personal data and secrets in error logs, and loops that log the same thing ten thousand times. Here is where to look and what to fix.
21
Log Context Enrichment: Adding Meaning to Your Events
Enrichment turns isolated log records into connected business events. Here is the architecture that makes it work — static resource attributes, background-refreshed caches, and per-request scopes — without taxing the request path.
22
Log-Based Monitoring: Alerting on the Evidence
Logs carry operational state at a resolution metrics can't match. Most teams only open them after something breaks. This guide covers how to query them continuously, turn them into metrics, and alert on what they surface without blowing up cardinality or cost.
23
Observability Under Compliance: GDPR, HIPAA, SOC 2, and PCI DSS
Regulated industries need observability too. A guide to building telemetry pipelines that satisfy GDPR, HIPAA, SOC 2, and PCI DSS requirements — covering data minimisation, retention mandates, audit trails, and what each framework actually requires.
24
Set Up Log-Based Alerting with Loki and Grafana
Turn a LogQL query into a Grafana-managed alert rule that fires on error volume, a specific error type or a failing dependency. Covers the query, the rule settings that trip people up, routing to PagerDuty and Slack, and an end-to-end test.
25
Data Masking in Telemetry: The Art of Safe Transformation
Telemetry data is just as risky for PII as any database. Here's how to turn sensitive fields into safe, useful signals: hashing, tokenising, coarsening, and picking the right tool for the job.
26
Distributed Logging: Ten Services, One Story
When a request crosses ten services, you get ten log streams that share nothing but a timestamp you can't trust. How to collect them on every node, carry the trace ID through, ship them through the OpenTelemetry Collector, and notice when lines go missing.
27
High-Throughput Logging: Keeping the Hot Path Fast
At hundreds of thousands of requests per second, the logging call itself becomes the bottleneck. Async channels, pooling, batching, and circuit breakers keep log I/O off the request thread.
28
Observability vs. Monitoring: Why the Distinction Matters
Monitoring tells you when something you predicted goes wrong. Observability lets you work out what's happening when it's something you didn't. You need both, and the gap between them is where incidents drag on.
29
The Evolution of System Understanding
From grepping one log file to querying wide, trace-linked events: how the questions we can ask a running system changed when monoliths split apart, and why OpenTelemetry had to exist.
30
Alert Fatigue Is an Observability Problem
Every alert that fires is doing its job. That is the problem. The model is wrong, not the thresholds.
31
Your Traces Are Leaking User Data
Every OTel span that includes a customer email, shipping address, or payment token is a GDPR audit waiting to happen. The fix isn't application code; it's a Collector pipeline.
32
Observability 1.0 meant forensics. Observability 2.0 means prevention.
Observability 1.0 taught us to look backward. Observability 2.0 asks us to look forward. Most teams haven’t made that shift yet. That’s why I named this site after a medieval divination practice.