Tagged: How-to
01
The Alert Body Template
A copy-paste alert body that answers four questions before the on-call has to ask: what's broken, how fast it's getting worse, what's already been checked, and the first three steps.
02
How to Instrument a .NET Service with OpenTelemetry
Add OpenTelemetry to an ASP.NET Core service: traces, metrics and logs in one setup block, manual spans and metrics for business logic, Serilog, and the zero-code agent for services you can't change. All verified against a local Collector.
03
How to Instrument a Java Spring Boot Service with OpenTelemetry
Instrument a Spring Boot service with the OpenTelemetry Java agent or the Spring Boot starter: traces, metrics and logs with no code, then your own spans and metrics, the Micrometer bridge and log correlation. All verified against a local Collector.
04
How to Set Up Your First SLO and Burn Rate Alerts
A step-by-step walkthrough: define an SLI, calculate your error budget, write Prometheus recording rules, and wire up multi-window burn rate alerts that page you before users notice.
05
How to Wire Trace IDs Into Your Logs
Logs and traces live in separate worlds until you connect them. Put the trace ID on every log line in .NET, Java, Go or Python, check where each runtime actually writes it, and make the field name a contract your backend can use.
06
Enrich Logs with Business Context in .NET
A log line that says 'payment failed' tells you something broke. One that says 'payment failed, enterprise customer, checkout-v2 experiment' tells you what to do about it. Here's how to add that context to every log event in a .NET service, safely, with Serilog.
07
How to Configure OTel Collector Tail Sampling
Move from flat probabilistic sampling to tail-based sampling in the OTel Collector. Keep every error and slow trace, cut health-check noise to 1%, and check that the Collector is doing what you think.
08
Scrub PII from Application Logs in .NET
Keep personal data out of your .NET logs before they leave the process: classify fields so the logger erases or pseudonymises them, scrub free text and exception messages in an OpenTelemetry processor, and prove nothing leaks.
09
Set Up Log-Based Alerting with Loki and Grafana
Turn a LogQL query into a Grafana-managed alert rule that fires on error volume, a specific error type or a failing dependency. Covers the query, the rule settings that trip people up, routing to PagerDuty and Slack, and an end-to-end test.
10
How to Benchmark Synchronous vs Channel Logging
Async logging is supposed to take I/O off the request thread. Measure it: a sync-vs-channel benchmark under concurrent producers, in .NET, Go, or Python, and how to read the result.