What Is OpenTelemetry? Traces, Metrics and Logs Explained
OpenTelemetry is the open standard for collecting traces, metrics and logs from your app and sending them to any observability backend. The three signals, spans and context propagation, auto-instrumenting a Node.js app, the Collector, backends, LLM and agent tracing, and what a small team actually needs.
OpenTelemetry (OTel) is an open-source standard — APIs, SDKs and a wire protocol — for collecting telemetry from your software: traces, metrics and logs. Instrument your app once with OpenTelemetry, and you can send the data to almost any monitoring tool, open source or commercial, without rewriting anything.
It's a CNCF project, and it's become the industry default for observability.
The three signals
Traces
A trace follows one request through your system. It's made of spans — timed operations with names and attributes:
POST /api/checkout 820 ms
├─ auth.verifySession 12 ms
├─ db SELECT cart_items 35 ms
├─ stripe.paymentIntents.create 610 ms ← slow
└─ db INSERT orders 28 ms
Traces answer "why was this request slow?" and "where did it fail?" — across services, queues and databases.
Metrics
Numbers over time: request rate, error rate, latency percentiles, queue length, memory. Cheap to store, good for dashboards and alerts.
Logs
Event records. With OpenTelemetry, logs can carry the trace ID, so you jump from a log line to the full trace of that request. (Structured logging)
Context propagation
The trick that makes distributed tracing work: when service A calls service B (or puts a job on a queue), it passes the trace context along — in HTTP, a traceparent header (W3C Trace Context). B's spans join the same trace. (HTTP headers explained)
Instrumenting a Node.js app
Auto-instrumentation patches common libraries (HTTP, Express, Postgres, Redis, fetch…) so you get useful traces without changing your code:
npm install @opentelemetry/sdk-node @opentelemetry/auto-instrumentations-node \
@opentelemetry/exporter-trace-otlp-http
// instrumentation.ts — load before your app
import { NodeSDK } from '@opentelemetry/sdk-node'
import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node'
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http'
new NodeSDK({
serviceName: 'invoices-api',
traceExporter: new OTLPTraceExporter({ url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT + '/v1/traces' }),
instrumentations: [getNodeAutoInstrumentations()],
}).start()
node --import ./instrumentation.js server.js
Add custom spans for business steps:
import { trace } from '@opentelemetry/api'
const tracer = trace.getTracer('invoices')
await tracer.startActiveSpan('generate-pdf', async span => {
span.setAttribute('invoice.id', invoice.id)
try { await renderPdf(invoice) } finally { span.end() }
})
Next.js has built-in OpenTelemetry support via an instrumentation.ts file. Python, Go, Java and others have equivalent SDKs.
The Collector
The OpenTelemetry Collector is an optional service that receives telemetry, processes it (sampling, filtering out sensitive attributes, batching) and forwards it to one or more backends. Apps send to the Collector; you change backends by changing its config, not your code.
Backends
OpenTelemetry doesn't store or display anything itself. Send data to:
- Open source: Jaeger or Grafana Tempo (traces), Prometheus (metrics), Loki (logs), SigNoz, Uptrace.
- Commercial: Grafana Cloud, Honeycomb, Datadog, New Relic and most others accept OTLP.
- Error-focused: Sentry also ingests OpenTelemetry data. (Error monitoring for beginners)
LLM apps and agents
There are OpenTelemetry semantic conventions for generative AI: spans for model calls with token counts, model names and latency. Tracing an agent shows each model call and tool call as spans — invaluable for debugging why an agent took 40 steps or cost $3. Claude Code itself can export OpenTelemetry metrics and events for usage monitoring, and the Claude Agent SDK supports OpenTelemetry. (How to build an AI agent, Evaluating LLM outputs)
Be careful with prompt and completion content in spans — it may contain personal data. (GDPR basics)
What a small team actually needs
You don't need a full observability platform on day one:
- Error monitoring and uptime checks first. (Know when your app is down)
- Structured logs with request IDs.
- Add OpenTelemetry auto-instrumentation when you have performance questions you can't answer from logs, or more than one service.
- Sample traces in production (say 10%) to control cost.
Starting with OpenTelemetry rather than a vendor's proprietary agent keeps you free to switch later.
EasySpawn runs your app on a server where you can add an OpenTelemetry Collector or self-hosted backend next to it — and Claude Code can add the instrumentation and custom spans for you. See how it works or join the waitlist.
Related: Structured Logging · Error Monitoring for Beginners · How to Know When Your App Is Down · pg_stat_statements
Keep reading
"No Space Left on Device": Finding and Freeing Disk Space on a Server
When a server's disk fills up, databases stop writing, deploys fail and apps crash with ENOSPC. How to find what's using space with df and du, the usual culprits (logs, Docker, journald, old releases, deleted-but-open files), inode exhaustion, and how to stop it happening again.
Structured Logging: Logs You Can Actually Search
console.log('user saved') is useless at 3am when you need every request user 4812 made in the last hour. How structured logs work, what fields to include, request IDs that tie a request together, log levels that mean something, what never to log, and how logs make AI agents better debuggers.