scryops
menu

The dashboard was green, but the request was broken.

Metrics tell you how the crowd is doing. Logs tell you what one service saw. A trace tells you what one request went through, and at 2 a.m. that is usually the question you are asking.

It’s 2 a.m., and your phone almost vibrates off the table. You stumble to your laptop, open the dashboard, and see a sea of green: error rate at 0.3%, latency p50 at 42ms, health checks all smiling. You start to think this is just another false alarm. Twenty minutes later, another engineer hops on and drops a user complaint into Slack. Turns out, checkout is broken—but only for folks using a promo code on a cart with one very specific product category.

Your metrics had no idea.

You burn another forty minutes playing log detective, piecing clues from four different services. If you’d had a distributed trace, you could have solved it in thirty seconds: follow the spans, spot the database call that handed back a null discount, and watch it travel downstream as a perfectly respectable $0.00 total. A mimic posing as a treasure chest: no error, no alarm, nothing until someone opens the lid.

This isn’t a rare edge case. I’ve seen this exact shape more than once: a value that is wrong but valid, travelling through healthy services and averaging away into a green dashboard. It’s the classic failure mode when you rely on metrics first, especially in systems that were never a tidy monolith to begin with.

The dashboard at 2 a.m.: every aggregate inside its alert lineThe dashboard shows an error rate of 0.3 percent against a 1 percent alert, latency p50 of 42 milliseconds against a 200 millisecond alert, and 4 of 4 health checks passing, so nothing pages.THE DASHBOARD, 02:04// every aggregate inside its lineerror ratealert at 1%0.3%✓oklatency p50alert at 200 ms42 ms✓okhealth checksall passing4 / 4✓ok→ nothing pages. nobody looks.One broken trace from the same minuteOne checkout trace, trace-id a1b2c3, drawn as a span waterfall. POST /checkout calls cart.validate, which calls promo.apply, which calls db.query. The db.query span takes 18 milliseconds and returns null without recording an error, and the whole request finishes in 112 milliseconds with 200 OK and a total of 0.00 dollars. Nothing is slow and nothing is red. After the fix, the same span carries app.discount.missing = true, which is what lets tail sampling keep it.ONE TRACE, a1b2c3// the request that actually broke060 ms120 msPOST /checkoutcart.validatepromo.applydb.querySILENT FAILdb.query 18 ms → nullspan status unset, no errorAFTER THE FIXapp.discount.missing = true→ 200 OK in 112 ms, total $0.00
Fig. 1 — The dashboard averages this request away. Nothing in the trace is slow or red either: the null discount is the only clue, and only the trace shows it.

The Three Witnesses

Picture your observability stack as three types of witnesses at the scene of an incident.

Logs are your classic eyewitnesses: detailed, specific, and a bit myopic. Each one saw exactly what happened in its own little corner, but none of them caught the whole show. Trying to rebuild a single request from five different log streams is slow, manual work, and nobody wants to do it at 2 a.m.

Metrics are the statisticians in the room. They’ll tell you how often things happen, how fast, and in what total. Perfect for capacity planning, SLOs, and trend-spotting. But if you want to know why this request failed, metrics shrug—they’re designed to toss out the individual stories and focus on the crowd.

Traces are your surveillance footage. They follow a single request from the front door to the exit, catching every service hop, every database query, and every millisecond spent twiddling its thumbs. Here the unit you study is the request itself: richer than any single log line, more specific than any metric. That’s a whole different way of asking questions.

WitnessAnswersMisses
LogsWhat happened in this serviceThe rest of the request
MetricsHow often, how fast, how manyWhich request, and why
TracesWhat this request went throughTrends, unless aggregated

The Dashboard Reflex

In 2018, Nicole Forsgren, Jez Humble and Gene Kim published Accelerate, one of the most influential engineering books of the decade, built on the research of the DevOps Research and Assessment (DORA) program. It gave the industry something it badly needed: a shared vocabulary for delivery performance. Its four metrics (deployment frequency, lead time for changes, change failure rate, and mean time to restore) became the benchmark you’ll measure every engineering team against. DORA has since moved on: in 2023 time to restore became failed deployment recovery time, and in 2024 a fifth metric, deployment rework rate, joined the set. The idea stuck. The research holds up, and the framework was exactly the right tool for the job it was built for.

But there was a sneaky side effect. Once software delivery got boiled down to four tidy numbers, a lot of teams started expecting the dashboard to have all the answers. The reflex kicked in: something’s wrong? Find the metric. Build the dashboard. Stare at the number.

Here’s where the wires get crossed. DORA metrics are all about your delivery process—they show how well your team ships software over time. They’re not built to diagnose a single failed request. Mean time to restore might tell you last quarter’s average was 47 minutes, but it won’t say a word about why tonight’s incident dragged on for an hour: which service tripped up, which team owned it, which code path went astray, or which upstream quietly ghosted you.

That mysterious hour? That’s distributed tracing’s home turf. Delivery metrics can tell you things are slowing down, but only a trace will show you exactly where this one began its journey to nowhere.

Why Tracing Didn’t Win Sooner

Distributed tracing has existed since Google published the Dapper paper in 2010. Twitter open-sourced Zipkin in 2012, and Uber open-sourced Jaeger in 2017. If the value was always there, why did mainstream adoption take another decade?

There were really two big roadblocks.

First up: instrumentation fragmentation. Tracing used to mean picking a tracer or a vendor (Zipkin, Jaeger, Lightstep, Datadog) and bolting its client library onto every service in your call graph. It was sticky, pricey, and migrating later was a headache. OpenTracing and OpenCensus tried to help, but having two standards just meant double the confusion. They finally merged into OpenTelemetry in 2019. If you had fifty services, that meant fifty SDK integrations, tied to one backend, before you saw a single complete trace, and a lot of teams just decided it wasn’t worth the hassle.

Second: the sampling trap. Capturing every single trace is expensive—store every span from every request and your storage bill balloons with every user click. The usual fix was head-based sampling: flip a coin at the front door, trace 5 or 10% of requests, and hope the sample tells the story.

But head-based sampling is flying blind. It tosses traces before it knows if they’re interesting. The slow requests, the errors, the weird edge cases—they get sampled just as often as the boring, happy-path traffic. You end up with a sample that’s statistically fair, but mostly full of traces you’ll never care about.

What Changed

OpenTelemetry came along and fixed the fragmentation mess. Now there’s one API and spec, with an SDK for every language, and the W3C Trace Context standard (that’s the traceparent header) keeps spans passing smoothly without vendor lock-in. Auto-instrumentation agents—bytecode for Java, monkey-patching for Python and Node.js—let you instrument existing services without touching the code. The sticky vendor lock-in that made tracing a pain? That’s mostly history.

Tail-based sampling solved the sampling trap. Instead of making a snap decision at the start, the OpenTelemetry Collector’s tail_sampling processor buffers spans and waits. It doesn’t know exactly when a trace is done, so it waits a set decision_wait after the first span shows up, then makes the call. Spans that arrive after the decision can end up judged separately, which leaves you with half a trace, so set decision_wait longer than your slowest normal request (and turn on the decision cache). Within that window, you can keep every trace with an error, every slowpoke over your latency threshold, and just a pinch of the rest:

processors:
  tail_sampling:
    decision_wait: 10s
    num_traces: 100000   # ≈ traces/sec × decision_wait, plus headroom
    policies:
      - name: errors
        type: status_code
        status_code: {status_codes: [ERROR]}
      - name: slow-traces
        type: latency
        latency: {threshold_ms: 1000}
      - name: silent-business-failures
        type: boolean_attribute
        boolean_attribute: {key: app.discount.missing, value: true}
      - name: baseline-sample
        type: probabilistic
        probabilistic: {sampling_percentage: 5}

One catch: the checkout bug from the top of this article would have slipped past the first two policies. It returned 200 OK, and it was fast. Tail sampling can only keep what the span says is interesting, so teach your code to say it. When the discount lookup comes back null, set app.discount.missing=true on the span, and the silent-business-failures policy keeps every one of those traces.

Head-based vs tail-based sampling on the same budgetAn illustrative window of 400 requests: 378 ok, 12 slow, 8 errors and 2 silent failures that return 200 OK quickly but are flagged on the span. Head-based sampling at 10 percent keeps about 40 traces chosen blind: roughly 1 error, 1 slow trace, no silent failures and 38 ok traces. Tail-based sampling keeps 41 traces: all 8 errors, both silent failures because of the flag, all 12 slow traces, and 19 ok traces from a 5 percent baseline. The storage cost is about the same; the contents are not.SAME BUDGET, DIFFERENT TRACES// 400 requests: 378 ok, 12 slow, 8 errors, 2 silent fails. illustrative.HEAD-BASED, DECIDED AT THE FRONT DOORkeeps 10% of everything, before it knows which trace matters~1 of 8 errors0 of 2 silent~1 of 12 slowTAIL-BASED, DECIDED AFTER A SHORT WAITkeeps every error, slow trace and flagged silent fail, 5% of the rest8 of 8 errors2 of 2 silent12 of 12 slowerrorsilent fail, 200 OKslow, over 1 sokone cell = one kept trace
Fig. 2 — Both strategies store about forty traces. Head sampling spends them on the happy path; tail sampling spends them on the requests you'd page about, including the silent ones, but only because the span says so.

Now the traces you actually care about, the weird, here-be-dragons ones, don’t get tossed out by accident. The routine, successful traffic gets a light sampling for stats.

The trade-off is memory: the Collector holds every span for decision_wait, so size num_traces to roughly your traces per second times decision_wait, with headroom.

Tail sampling uses OR semantics. A trace is kept if any policy samples it, unless a drop policy matches it, and policy order doesn’t set priority. At scale, with several Collector replicas, run a first tier of Collectors with the loadbalancingexporter in front of the sampling tier, so every span of a trace reaches the same replica. Tail sampling needs the whole trace in one place to decide correctly. Without that routing, a trace split across replicas gets judged on half its spans.
flowchart TD A[checkout] -->|"a1b2c3, d4e5f6"| LB B[cart] -->|"a1b2c3"| LB C[promo] -->|"d4e5f6"| LB LB["Collector tier 1
loadbalancingexporter
routing_key: traceID"] LB -->|"all of trace a1b2c3"| R1["tier 2, replica A
tail_sampling"] LB -->|"all of trace d4e5f6"| R2["tier 2, replica B
tail_sampling"] R1 --> BE[("Tracing backend")] R2 --> BE
Fig. 3 — Route by trace ID first, so each sampling replica sees whole traces.

Where to Start

You don’t need fifty services instrumented to get value. You need one request path, traced end to end.

  1. Drop the OTel SDK into one service: the entry point for your most important user flow.
  2. Stick with the default W3C propagator and set up an OTLP exporter.
  3. Instrument the next service downstream, or let a framework-level auto-instrumentation agent do it. The spans link up automatically, and a two-node trace is already a win.
  4. Keep extending the instrumented path, one service at a time, until the whole critical chain is covered.

These four signals pay for themselves right away: http.request.method, http.response.status_code, db.query.text (sanitized, of course), and a span status of Error with error.type set on failures. With just those, you can zero in on every faulty trace, every slow database call, and every error—no more log file scavenger hunts across services.

Then add one attribute your business cares about on the critical path, such as whether a promo code was applied and what discount came back (the app.discount.missing flag from the sampling policy above). That is the attribute that would have found the 2 a.m. checkout bug.

Next time the dashboard is green and checkout is broken, you won’t be piecing it together from four log streams. You’ll open one trace and watch the $0.00 discount happen.

One observability idea, every Tuesday

A short take, one thing to try that week, and a link to the full article. No vendor pitches.

Subscribe on Substack