Three Tabs and a Timestamp
Most investigations still follow the same choreography. A latency panel spikes at 14:07. The engineer notes the time, opens the tracing tool, filters by service and a duration threshold, scrolls through candidates that might be the right request, picks one, copies its trace ID, and pastes it into the log search. Ten minutes have passed and the incident has not been touched yet.
The three signals each answer a different question. Metrics say when and how much. Traces say where in the request path. Logs say why. What is usually missing is the join key that lets a person move between them without re-typing anything. Two pieces of plumbing supply that key: trace context and exemplars.
One ID to Bind Them
The join key is the trace ID, and the first job is to make sure every component agrees on it. The W3C Trace Context recommendation standardises the traceparent and tracestate headers so that a trace ID survives the hop from an API gateway written in Go to a queue consumer written in Java. OpenTelemetry SDKs propagate it by default, and, just as importantly, context propagation makes the current trace and span available to anything running inside the request, including your logger.
That is the second step: stamp trace_id and span_id onto every log record. The OpenTelemetry logging specification reserves fields for both, and most SDKs inject them automatically once logging is bridged. If your logs are already structured, this is two extra keys; if they are not, this is the reason to fix that.
Exemplars: Metrics That Remember a Request
Metrics are aggregates, which is what makes them cheap and also what makes them forget every individual request that went into them. An exemplar is the fix: a single representative sample, tagged with a trace ID, attached to a histogram bucket or counter. The OpenMetrics specification defines them and caps the label payload at 128 characters, deliberately, so that nobody is tempted to smuggle whole events through the metrics pipeline.
OpenTelemetry SDKs generate exemplars automatically: the default exemplar filter in the metrics SDK is trace-based, meaning a measurement recorded inside a sampled span becomes a candidate exemplar. On the storage side, Prometheus keeps them in an in-memory ring buffer once the exemplar-storage feature flag is enabled. Grafana then draws each exemplar as a star on the graph, and clicking one opens the trace in Tempo. The spike at 14:07 becomes a request you can read.
Closing the Loop in the Other Direction
The link also runs backwards. Tempo's metrics-generator derives request, error and duration metrics from the spans it ingests and attaches exemplars to them as it goes, so services with tracing but no metrics instrumentation still get SLI-grade series with the trace links built in. Commercial platforms ship equivalents: Datadog's log and trace correlation injects trace identifiers into logs and surfaces them in both views. The mechanism is the same everywhere; only the buttons differ.
A Rollout That Fits in a Sprint
- Standardise propagation on W3C Trace Context at every boundary, including message queues and scheduled jobs, which are the places traces usually break.
- Inject trace and span IDs into logs, then confirm with one query that a trace ID in the log store returns exactly the lines from that request.
- Enable exemplars on the latency histograms and error counters that feed your SLIs. Those are the metrics you will be staring at during an incident.
- Configure the data source links in your dashboarding tool so the click-through actually works, and test it from a burn rate alert.
One caveat matters. An exemplar is only useful if the trace it points to was kept, so sampling policy and exemplars have to be designed together. Tail-based sampling that always retains slow and failed requests, covered in distributed tracing on a budget, is the natural partner, because those are exactly the requests an exemplar will point to.
Why This Belongs in an SLO Programme
The payoff shows up in time to diagnosis. When a burn rate alert fires, the responder goes from the SLO panel to an exemplar, to a trace, to the logs of the failing span in three clicks, and the ten minutes of copy-and-paste become thirty seconds. Every minute saved there is error budget the service keeps. If you are still deciding which indicators deserve this treatment, start by mapping user journeys to SLIs; the metrics that define your SLOs are the ones that most need a path back to a single request.