OpenTelemetry Observability: Follow One Request Before Collecting Everything

A customer reports that checkout is slow, but every server appears healthy. Infrastructure metrics cannot tell you which dependency delayed that particular request. OpenTelemetry helps applications describe their work consistently, so you can follow execution across boundaries and connect it with other operational evidence.

This guide is for developers and platform teams introducing observability to a service they already maintain. Read it to build one useful telemetry path, understand the Collector’s role, and avoid turning a troubleshooting project into an uncontrolled data collection exercise.

Choose a question before a dashboard

Start with a question such as: does checkout spend most of its time waiting for inventory, payment, or database access? Define what evidence would answer it. A trace records related operations as spans. Metrics summarize measurements across events or time. Logs capture individual records that can provide additional context. The official signals documentation explains these distinct roles.

OpenTelemetry provides instrumentation and telemetry transport components. It is not itself your long-term storage, search interface, or alerting service. You still select and operate a backend, or use a managed one.

The core approach is established, while stability varies across language SDKs, instrumentation packages, Collector components, and semantic conventions. Treat newer integrations as individual dependencies to evaluate. A project-level adoption claim does not establish the maturity of every component you install.

Prerequisites for a focused first run

  • A development copy of one service and a repeatable request that exercises a downstream dependency.
  • The official OpenTelemetry SDK or automatic instrumentation appropriate to that service’s language and framework.
  • A Collector distribution containing the OTLP receiver and debug exporter, pinned to a tested release.
  • A backend for the later pilot, plus an agreed retention and access policy.
  • A short attribute allowlist that excludes secrets and unnecessary personal data.

Pick a stable service name, such as checkout-api. Record deployment and version metadata consistently so two services do not accidentally appear as one. Initialize instrumentation before the framework modules it needs to instrument, following the language-specific instructions.

A local Collector you can understand

The Collector receives telemetry, optionally processes it, and exports it. Components become active when referenced by a service pipeline. The configuration guide also documents configuration validation.

For a local native Collector and an application running on the same machine, this minimal YAML accepts OTLP over HTTP and prints trace details:

receivers:
  otlp:
    protocols:
      http:
        endpoint: 127.0.0.1:4318
exporters:
  debug:
    verbosity: detailed
service:
  pipelines:
    traces:
      receivers: [otlp]
      exporters: [debug]

Save it as a local configuration, then run these commands with a distribution whose executable is named otelcol:

otelcol validate --config=otel-local.yaml
otelcol --config=otel-local.yaml

Configure your application’s installed OTLP exporter for HTTP/protobuf and the base endpoint http://127.0.0.1:4318. A trace-specific endpoint instead includes /v1/traces. Confirm the configuration mechanism supported by your SDK. This receiver is intentionally local; a container or remote application needs an explicitly designed network path.

Turn the first trace into an acceptance test

  1. Send the selected request and confirm the debug output contains the expected service identity.
  2. Check that downstream spans share the trace context. A new unrelated trace at each service often indicates missing propagation.
  3. Add a small manual span around meaningful application work that automatic instrumentation cannot describe.
  4. Introduce a controlled delay in a development dependency and verify that its span explains the increased request time.
  5. Replace the debug exporter with a supported backend exporter, configure authentication and transport security, and repeat the test.

This example enables traces only and has no durable queue or production capacity controls. Detailed debug output is useful for synthetic requests but can expose telemetry contents. Remove it from the production path after verification.

Spend telemetry where it answers questions

For the checkout scenario, a route template is usually more useful as a metric dimension than a unique order identifier. Unbounded attribute values can multiply time series and drive storage and query costs. Keep business identifiers out of metric labels; use carefully controlled trace context when investigation genuinely requires correlation.

Sampling trades trace completeness for lower volume. Head sampling decides early and can miss rare errors. Tail sampling can consider completed trace information but needs buffering and routing that brings the relevant spans together. Start with a policy you can explain and measure, then revisit it using actual traffic.

Budget for Collector resources, network transfer, backend ingestion, retention, and engineering time. Monitor exporter failures and queue pressure. A telemetry outage should be visible without silently exhausting the application’s resources.

Make the pipeline safe to operate

Follow the project’s security guidance: restrict receiver exposure, protect transport, and control access to telemetry. Avoid collecting authorization headers, credentials, and request bodies by default. Review instrumentation changes for new fields before rollout.

Your next milestone is one reproducible incident investigation: detect a slow request, locate its cause, and explain the result using retained evidence. Expand to additional services only after that path works reliably.