Hover a dotted term for 5 seconds to lock its explanation. It closes after 5 seconds away; nested tooltips and keyboard focus keep it open. Click, tap or Enter locks immediately. Technical glossary.

SRE and observability

OpenTelemetry for Microservices: Tracing, Incidents and Cost Control

A real client engagement. The engineering and the results are described below.

A marketplace we worked with has a slow checkout, healthy individual service dashboards and an increasingly expensive telemetry bill. The team follows one customer journey instead of collecting more events without a question.

Request a time through the inquiry form. A meeting is confirmed separately by email.

The business problem behind the technology

Can the evidence shorten diagnosis without becoming another cost problem?

The team wants faster diagnosis and sustainable evidence retention without collecting sensitive customer data or discarding the traces needed for rare failures.

Read the client engagement ↓

Client engagement / Delivered results

The missing connection was between the dashboards

A marketplace we worked with investigates twelve comparable service incidents each month. It also paid a $40,000 monthly telemetry charge, with high-cardinality and low-value data crowding the signal.

The constraint

The team wants faster diagnosis and sustainable evidence retention without collecting sensitive customer data or discarding the traces needed for rare failures.

The engineering decision

It instruments the checkout journey, preserves correlation across queues, controls attributes at collection and chooses sampling from investigation questions. Diagnosis effort fell from 60 to 20 minutes per incident and retained telemetry charges fell to $28,000.

The delivered outcome

The engagement yielded 40 minutes less diagnosis per incident and $12,000 less telemetry charge per month. These are separate benefits and were not added into one ROI figure.

Measured inputs and delivered differences
Measure / unitBeforeAfterDifference
Comparable incident diagnosis time
minutes/incident
602040
Telemetry platform charges
USD/month
40,00028,00012,000

Comparable incident diagnosis time. The diagnostic interval for comparable incidents.

Telemetry platform charges. The billable reduction kept the investigative evidence the team needed.

The conditions behind the results

  • The twelve incidents have comparable diagnostic complexity; frequency was not expected to fall.
  • The same security/privacy and investigation requirements apply after collection changes.
  • The monetary estimate excludes implementation and staff costs, and is not added to time capacity.

What this does not prove. Diagnosis time, outage duration and platform charges are different quantities; none proves prevented revenue loss.

Evidence to collect for your own decision

  • Trace a real user journey through synchronous and asynchronous boundaries.
  • Test sampling against rare-error questions and verify cardinality/privacy controls.
  • Compare diagnostic intervals and actual retained-data charges over equivalent periods.

Key decisions

OpenTelemetry becomes useful when correlation survives the customer journey. Choose collection, sampling and retention policies that answer operational questions without uncontrolled cost or sensitive-data exposure.

  • Trace a real user journey through synchronous and asynchronous boundaries.
  • Test sampling against rare-error questions and verify cardinality/privacy controls.
  • Diagnosis time, outage duration and platform charges are different quantities; none proves prevented revenue loss.

Follow the decision

Can the evidence shorten diagnosis without becoming another cost problem?

Select a step to follow its reasoning, then continue into the technical chapters.

Problem → boundary → decision → evidence

Choose the journey

Instrument a useful request path rather than every function indiscriminately.

Read every component and connection
Choose the journey · Problem
Instrument a useful request path rather than every function indiscriminately.
Keep correlation · Boundary
Carry identity across synchronous and asynchronous boundaries.
Bound collection · Decision
Control cardinality, sensitive data and sampling.
Guide action · Evidence
Alert on symptoms and retain a path into the evidence.
  • Choose the journey → Keep correlation: identify the constraint
  • Keep correlation → Bound collection: choose a bounded change
  • Bound collection → Guide action: check the outcome

A conceptual decision map for this article, not a measured timeline, physical topology or a depiction of a specific client system.

Hover, focus or tap a component to inspect it. Motion adapts automatically to connection, device and accessibility signals; the component key remains readable without JavaScript.

Instrument the customer journey, not every function

The client’s marketplace starts with checkout because its diagnosis question must survive the service boundaries.

Begin with the outcome customers experience: a valid checkout receives an accurate confirmation within the agreed time. Break that journey into meaningful operations such as validation, inventory reservation, queue publication, and fulfillment processing. Automatic HTTP and database instrumentation provides a foundation, but custom spans should describe important domain boundaries rather than wrapping every helper function.

Assign consistent resource identity, including service name, environment, and release version. Keep span names stable, using a route template instead of a concrete customer URL. Record operation outcomes so a successful transport response that contains a business rejection can be distinguished from an infrastructure failure. Document which outcome counts as bad for the service-level objective.

Asynchronous processing requires special care. Queue wait is not the same as consumer execution time. Record publication and processing boundaries, and propagate context through message metadata using the messaging instrumentation conventions for the selected library. Batch consumers may need span links to several upstream contexts rather than pretending that unrelated messages have one parent.

Make correlation work across real boundaries

The request crosses a real service boundary. Correlation must survive that crossing instead of ending at the first dashboard.

OpenTelemetry context propagation carries trace identity across process boundaries. For supported HTTP instrumentation, this commonly uses W3C trace context. Extraction on the receiver and injection on the sender must both work; one missing hop can fragment a journey into apparently unrelated traces. Test the actual gateway, HTTP client, queue wrapper, and background worker rather than assuming instrumentation coverage.

Include trace and span identifiers in structured logs through the active logging integration. The identifiers let an investigation move from an interesting span into the logs emitted for that execution. They should not become ordinary metric labels. Metrics aggregate population behavior; traces preserve individual request paths; logs explain specific events. These signals complement one another rather than requiring identical retention or indexing.

Where the SDK and backend support them, metric exemplars can link an aggregate observation to a sampled trace without creating a time series for every request. Verify the link in the real telemetry path. A trace identifier in a log can still refer to a trace that was not retained, so correlation is useful navigation, not a guarantee of complete evidence.

Control cardinality and sensitive data at collection

Collection policy now controls both the risk of sensitive data and the cost of retaining evidence.

A bounded label such as operation or status class supports aggregation. Customer identifiers, raw URLs, arbitrary error messages, and request IDs create unbounded combinations and can overwhelm a metric backend. Histograms multiply this cost across buckets and label sets. Choose dimensions from actual diagnostic questions, estimate their product, and watch active series as instrumentation changes.

Telemetry is another data system with access and retention obligations. Do not record authorization headers, payment fields, full prompts, or raw request bodies by default. Prefer an allowlist of safe attributes and redact before export. Collector filtering is useful defense in depth, but sensitive values already written to local logs or SDK buffers may have crossed a boundary before the collector sees them.

Baggage propagates values downstream and must not carry secrets or unnecessary personal data. Treat inbound context from untrusted callers carefully and control propagation to external services. A hash of a customer identifier is not automatically anonymous: stable identifiers can remain linkable. Decide whether a diagnostic need justifies restricted trace attributes instead of assuming every useful field belongs everywhere.

Choose sampling that preserves the questions

Sampling decides which questions the retained evidence can answer. You choose it from the investigation need rather than a convenient percentage.

Head sampling decides early and is comparatively simple, but it cannot know that a later span will fail. Tail sampling can retain traces based on errors or latency after observing enough spans. It requires buffering, suitable trace routing, and a decision window that accounts for late spans. Collector capacity and dropped-span metrics are therefore part of application observability, not an unrelated concern.

Tail sampling cannot recover spans discarded by an upstream head sampler. If preserving rare failures is a requirement, design the entire pipeline around that requirement and its cost. Error-biased samples are excellent for debugging but are not an unbiased estimate of population error rate. Use unsampled request metrics for SLO accounting, and document the sampling policy next to trace-based dashboards.

For the checkout system, keep sufficient successful traces to compare healthy and slow paths, not only failures. Otherwise, an expensive-looking inventory span lacks a baseline. Exercise a controlled queue delay and verify that its trace remains useful after it traverses every collector and retention rule.

Alert on symptoms, then guide an investigation

The team compares diagnostic intervals and telemetry charges separately, rather than claiming one inflated return.

An alert should identify affected behavior, urgency, and the next useful view. Prefer sustained SLO budget burn or growing queue age over a CPU threshold with no demonstrated customer impact. Use short and long windows to distinguish a sharp incident from slow degradation. Route capacity trends to planned work unless immediate action is necessary.

The query below assumes application-defined counters that count the same eligible requests and increment bad once for an error or an SLO latency breach. It shows a five-minute bad-request fraction, not a complete paging rule. Handle missing telemetry separately and set traffic-volume conditions appropriate to the service; an empty vector is not proof of health.

Validate the design by introducing an authorized fault, following a metric to a trace and correlated logs, locating the delayed boundary, and confirming recovery. Measure telemetry coverage, alert usefulness, collector loss, storage cost, and investigation steps. A smaller set of connected signals is more valuable than a large dashboard collection that cannot explain one failed journey.

promql / example
sum(rate(checkout_bad_requests_total[5m]))
/
sum(rate(checkout_eligible_requests_total[5m]))

Questions behind the decision

How can OpenTelemetry help diagnose microservices failures?

It can carry correlated context across the path a request actually takes, including asynchronous boundaries. Instrumentation and retention must preserve the question being investigated; installing a collector alone does not create useful traces.

How do you lower observability costs without losing useful evidence?

Control attributes and cardinality before ingestion, then choose sampling and retention from concrete investigation needs. Recheck rare failures and privacy requirements before treating a smaller data bill as a safe improvement.

References & further reading

Engineering notes

Keep following the thread.

Real client engagements and the engineering behind them.

Technical glossary: definitions, connected ideas and further reading.

Optional analytics off. Contact works either way.

How measurement works