Tracing answers the one question metrics and logs cannot: where, along a request's path through many services, did the time go or the failure start. The vocabulary is small and the exam expects it exact.

needsmake up obs

Orientation

competency 4.1 · tracing solutions

Two things make tracing tasks solvable: knowing the exact OTLP address telemetry should go to, and knowing that a trace is only a trace if context propagated. Everything else is UI.

Lab topology

The OTel operator lives in opentelemetry-operator-system, the collector in tracing, Jaeger alongside it. So the in-cluster OTLP endpoints are otel-collector.tracing.svc:4317 (gRPC) and :4318 (HTTP). Memorize the form <svc>.<ns>.svc:4317; every tracing task starts by knowing that address.

Vocabulary and plumbing

spans · context · pipeline

A trace is one request's journey. A span is one named, timed operation within it, carrying a trace ID, its own span ID, a parent span ID, a status, attributes and events. A trace is therefore a tree, and the root span is the entry point.

Context propagation is what links spans across service boundaries: the caller injects the trace context into headers, the callee extracts it and continues the trace instead of starting a new one. The standard header is W3C traceparent (plus tracestate, and baggage for user-defined key/values that travel with the request). Older stacks use B3 headers; a mixed fleet needs a propagator configured for both.

The classic failure, and how it looks

Broken propagation: every service reports spans, and every span is its own orphan trace. Symptom in Jaeger: many single-span traces, one per service, none joined. Causes: a proxy or gateway stripping unknown headers, an SDK without the right propagator, a manual HTTP client that never injects, or an async hop (queue, cron) where nobody thought to carry the context. When you see one-span traces from an instrumented system, name that cause immediately.

The OpenTelemetry pipeline

app (SDK or auto-instrumentation agent)
   │ OTLP  gRPC :4317 / HTTP :4318
   ▼
Collector   receivers ──▶ processors ──▶ exporters
             otlp, jaeger,     batch, memory_limiter,   otlp/jaeger,
             prometheus,       attributes, k8sattributes,  debug,
             filelog           tail_sampling            prometheus
   ▼
backend  Jaeger (traces) · Prometheus (metrics) · Loki (logs)

The collector is the platform team's control point: sampling, batching, redaction and fan-out happen there, not in application code. That is the platform-engineering answer to "how do we control telemetry cost": you change one deployment, not forty services.

SamplingDecidesTrade
Head-basedat the start of the trace, usually a fixed probabilitycheap and simple; you may drop the one slow request you needed
Tail-basedafter the trace completes, on its properties (errors, latency)keeps the interesting traces; the collector must buffer whole traces, so it costs memory and needs all spans of a trace at one collector

Kubernetes-side: two CRDs

  • OpenTelemetryCollector: the operator deploys and configures a collector (as a Deployment, DaemonSet, StatefulSet or Sidecar). Read its config to see the receiver/exporter chain.
  • Instrumentation: the zero-code path. The lab ships one named auto in default; annotating a pod spec with instrumentation.opentelemetry.io/inject-java: "default/auto" (or -python, -nodejs, -dotnet) makes a webhook inject the language agent as an init container at admission.
Injection happens at pod creation

Annotate an existing Deployment and nothing happens to the running pods; they must restart to pick it up. And the annotation belongs on the pod template, not on the Deployment's own metadata. Those two mistakes account for most "the annotation does nothing" reports, and both are one-line fixes once you know to look.

The collector, field by field

pipelines, connectors, the processors that decide cost

A collector config has four component sections (receivers, processors, exporters, connectors), an extensions section for things that are not in the data path (health_check, pprof, zpages), and a service block that wires them into named pipelines per signal. Defining a component does nothing until a pipeline lists it. Components can be declared more than once with a suffix (otlp/2, otlphttp/loki), and a connector is both an exporter of one pipeline and a receiver of another, which is how spanmetrics turns traces into RED metrics without touching the application.

receivers:
  otlp: { protocols: { grpc: { endpoint: 0.0.0.0:4317 }, http: { endpoint: 0.0.0.0:4318 } } }
processors:
  memory_limiter: { check_interval: 1s, limit_mib: 800, spike_limit_mib: 160 }
  k8sattributes: {}
  batch: {}
exporters:
  otlp/jaeger: { endpoint: jaeger-collector.tracing.svc:4317, tls: { insecure: true } }
  prometheus: { endpoint: 0.0.0.0:8889 }
connectors:
  spanmetrics: {}
service:
  pipelines:
    traces:  { receivers: [otlp], processors: [memory_limiter, k8sattributes, batch], exporters: [otlp/jaeger, spanmetrics] }
    metrics: { receivers: [spanmetrics], processors: [batch], exporters: [prometheus] }
ProcessorFields that matterRule
memory_limitercheck_interval, limit_mib (hard), spike_limit_mib (soft = hard minus spike; start at 20% of hard) or limit_percentagefirst in every pipeline so back-pressure reaches the receiver; set GOMEMLIMIT as well
batchsend_batch_size (8192), timeout (200ms), send_batch_max_sizelast before the exporters; reduces request count and exporter cost
k8sattributespod_association (match by k8s.pod.ip or k8s.pod.name+k8s.namespace.name), extract.metadata, extract.labelsneeds RBAC to list pods, replicasets, namespaces; without it spans carry no k8s.* resource attributes and Loki/Tempo correlation breaks
resource / attributesactions: [{key, action: insert|upsert|delete|hash}]redaction (delete on http.request.header.authorization) belongs here, not in forty SDKs
filterOTTL conditions, e.g. drop spans where attributes["http.route"] == "/healthz"dropping individual spans breaks trace trees; prefer sampling whole traces
tail_samplingdecision_wait (30s), num_traces (50000), policies[] of type always_sample, latency, status_code, probabilistic, rate_limiting, string_attribute, ottl_condition, and, compositeevery span of a trace must reach the same instance: put a loadbalancing exporter with routing_key: traceID in a first tier, or run one replica
probabilistic_samplersampling_percentagehead sampling in the collector; cheap, blind to errors

Two deployment roles: an agent (DaemonSet or sidecar) close to the workload that enriches and batches, and a gateway (Deployment behind a Service) that samples, redacts and fans out. Tail sampling and rate limiting belong in the gateway. The debug exporter (verbosity: detailed) is the fastest way to see whether spans arrive at all; the old logging exporter is gone from current builds, so a config that still names it fails to start with an unknown-exporter error.

Trap

An exporter's endpoint is a host:port for otlp (gRPC) and a URL for otlphttp. Sending gRPC to 4318 or HTTP to 4317 produces connection resets or 415/unimplemented errors rather than a clear message. The Python auto-instrumentation talks http/protobuf only, so its Instrumentation endpoint must be the 4318 one.

Operator CRDs, Jaeger v2 and how traces go wrong

modes, annotations, backends, exemplars, failure modes

OpenTelemetryCollector and Instrumentation

  • spec.mode: deployment (default), daemonset, statefulset (required for the target allocator and for sharded tail sampling), sidecar. A sidecar collector is injected into pods annotated sidecar.opentelemetry.io/inject: "true" (or the collector's name), a second, separate annotation from the language injection ones. spec.config is the collector YAML above; the operator renders it into a ConfigMap and a Service whose port names follow the receivers (otlp-grpc, otlp-http). spec.targetAllocator lets a StatefulSet of collectors share Prometheus scrape targets discovered from ServiceMonitors and PodMonitors.
  • Instrumentation fields: exporter.endpoint, propagators (tracecontext, baggage, b3), sampler.type (parentbased_traceidratio) and sampler.argument ("0.25"), plus per-language blocks (java.image, python.env, nodejs, dotnet, go) that can pin an agent image or add OTEL_* environment variables.
  • Annotation values for instrumentation.opentelemetry.io/inject-<lang>: "true" (the single Instrumentation in the pod's namespace; ambiguous if there are two), "name" (same namespace), "namespace/name", "false". Multi-container pods need instrumentation.opentelemetry.io/container-names: "app,worker" or the operator instruments only the first container. inject-sdk injects environment variables only, for a language with no agent. Go injection is eBPF: it needs otel-go-auto-target-exe pointing at the binary and a privileged sidecar, which a restricted namespace refuses.

Jaeger v2

Jaeger v2 is a distribution of the OpenTelemetry Collector: one jaeger binary configured with collector-style YAML, playing the roles collector, query, ingester (reads Kafka) or all-in-one. Its own components are the jaeger_storage extension (backends: in-memory, Badger for single-node persistence, Cassandra, Elasticsearch, OpenSearch, and ClickHouse in newer releases), the jaeger_query extension serving the UI and API on 16686, the remote_sampling extension (serves per-service sampling strategies to SDKs, static or adaptive) and the adaptivesampling processor. Ingest is the standard otlp receiver on 4317/4318; the v1 jaeger-agent and its Thrift ports are gone, and the docs tell you to run an OpenTelemetry Collector where you used to run the agent. The HTTP query API the lab used (/api/traces?service=) is unchanged, so scripted verification still works.

Grafana Tempo is the other trace store you will see named: object storage, no index beyond trace ID plus a search over recent blocks, TraceQL for queries, and a metrics-generator that emits span metrics and service graphs into Prometheus, the same job the spanmetrics connector does in the collector.

Exemplars and semantic conventions

An exemplar is a trace ID attached to a histogram bucket sample; the OpenMetrics text form is # {trace_id="abc..."} after the sample. Prometheus stores them only with --enable-feature=exemplar-storage (operator: spec.enableFeatures), and Grafana links them to Tempo or Jaeger through the datasource's exemplarTraceIdDestinations. The spanmetrics connector produces exemplars by default. Semantic conventions are the attribute names that make correlation possible: service.name (the one every SDK must set, or you get unknown_service), service.namespace, k8s.namespace.name, k8s.pod.name, k8s.deployment.name, http.request.method, http.response.status_code, url.path. A dashboard that groups by k8s_namespace_name is reading these attributes after the Prometheus exporter replaced dots with underscores.

Failure modes and their tells

TellCauseWhere to look
hundreds of single-span tracespropagation broken (proxy strips traceparent, propagator mismatch W3C vs B3, hand-rolled client)the caller's outbound headers; propagators in the Instrumentation
a child span starts before its parentclock skew between nodesnode NTP; the trace is real
a gap in the waterfall with no spanan uninstrumented hop (queue consumer, cron, a library without instrumentation) or the span was dropped by a filter processorthe service at the gap; the collector config
the slow request is never in the storehead sampling dropped itswitch to tail sampling on latency/status_code, or raise the ratio
service shows as unknown_serviceOTEL_SERVICE_NAME / service.name not setthe Instrumentation resource block or the pod env
spans arrive, no k8s.* attributesk8sattributes missing, or its RBAC, or the pod IP is hidden behind a gateway so pod_association cannot matchcollector logs; associate on k8s.pod.name instead of IP
collector logs data refused due to high memory usagememory_limiter hard limit hitraise limits or sample earlier; this is working as designed
injected init container present, no spansendpoint or protocol wrong in the Instrumentation (gRPC vs HTTP port), or a NetworkPolicy blocking egress to the collector namespacekubectl exec and read the OTEL_EXPORTER_OTLP_ENDPOINT env; test the port
How this gets tested

"Deploy a collector that receives OTLP and forwards traces to Jaeger" is an OpenTelemetryCollector with a receiver, a pipeline and an otlp exporter pointed at the Jaeger Service on 4317. "Auto-instrument this Deployment" is an Instrumentation plus one pod-template annotation and a rollout restart. "Keep only traces with errors or over one second" is tail_sampling with status_code and latency policies, and the reminder that it needs all spans on one collector.

Exercises

tick the dot when its check passes

kubectl -n tracing get cm -o yaml | grep -A30 'receivers:' (or the OpenTelemetryCollector CR if present). Draw the chain on paper: which receivers listen, which exporter points at Jaeger, what processors sit between.

verify: you can name the exact address an application in this cluster should send OTLP to (otel-collector.tracing.svc:4317 or the HTTP twin). Every tracing task starts by knowing that address.

telemetrygen is the collector project's own traffic generator, and it makes tracing exercises deterministic:

kubectl run telemetrygen --restart=Never \
  --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest \
  -- traces --otlp-endpoint otel-collector.tracing.svc:4317 --otlp-insecure \
     --service curriculum-drill --traces 20 --child-spans 3
outputcaptured 2026-08-26
$ kubectl run telemetrygen --restart=Never \
  --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest \
  -- traces --otlp-endpoint otel-collector.tracing.svc:4317 --otlp-insecure \
     --service curriculum-drill --traces 20 --child-spans 3
pod/telemetrygen created
verify: in Jaeger (address from make urls): service curriculum-drill appears in the dropdown, 20 traces, each of 4 spans (one root plus three children, two levels). Then verify the API way, because exams grade with curl: curl -s 'http://<jaeger>/api/traces?service=curriculum-drill&limit=1' | jq '.data[0].spans | length' returns 4.

Open one trace, read the waterfall: parent/child structure, per-span duration, attributes on each span. Answer for that trace: which span is the critical path, and what would you look at next if the leaf span were slow (that span's service's logs, at that timestamp; this is why traces carry IDs you can grep logs for).

verify: that narration is what "trace analysis and root cause" means as a testable skill.

Run a small Java service (ghcr.io/open-telemetry/opentelemetry-java-examples images work, or any Spring Boot sample), annotate its pod template with instrumentation.opentelemetry.io/inject-java: "default/auto", restart, and hit its endpoint. If it does not appear, the diagnostic ladder is: annotation on the pod template (not the Deployment metadata), pod restarted since annotating, Instrumentation CR namespace/name correct in the annotation value, operator webhook alive.

verify: kubectl describe pod shows the injected init container, and the service appears in Jaeger with HTTP spans you never wrote.

No cluster needed: service A calls B through a proxy that strips unknown headers. Describe what Jaeger shows and which header must survive.

verify: if your answer names traceparent and predicts orphaned single-service traces, the concept is yours.

A sidecar collector buffers locally so the app never blocks on a network hop, and the operator injects it from an annotation. Seeing the injected container is the whole point; the annotation is two words.

kubectl apply -f - <<'EOF'
apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata: { name: side, namespace: default }
spec:
  mode: sidecar
  config:
    receivers:
      otlp:
        protocols:
          grpc: { endpoint: 0.0.0.0:4317 }
          http: { endpoint: 0.0.0.0:4318 }
    processors:
      batch: { timeout: 5s }
    exporters:
      otlp:
        endpoint: otel-collector.tracing.svc:4317
        tls: { insecure: true }
    service:
      pipelines:
        traces: { receivers: [otlp], processors: [batch], exporters: [otlp] }
EOF
sleep 15
kubectl -n default run sidecar-demo --image=ghcr.io/stefanprodan/podinfo:6.7.1 --annotations="sidecar.opentelemetry.io/inject=true"
kubectl -n default wait --for=condition=Ready pod/sidecar-demo --timeout=120s
kubectl -n default get pod sidecar-demo -o json | jq -r '[.spec.initContainers[]?, .spec.containers[]?] | .[] | "\(.name) \(.image)"'
kubectl -n default get opentelemetrycollector -o custom-columns=NAME:.metadata.name,MODE:.spec.mode
kubectl -n default delete pod sidecar-demo
kubectl -n default delete opentelemetrycollector side
outputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata: { name: side, namespace: default }
spec:
  mode: sidecar
  config:
    receivers:
      otlp:
        protocols:
          grpc: { endpoint: 0.0.0.0:4317 }
          http: { endpoint: 0.0.0.0:4318 }
    processors:
      batch: { timeout: 5s }
    exporters:
      otlp:
        endpoint: otel-collector.tracing.svc:4317
        tls: { insecure: true }
    service:
      pipelines:
        traces: { receivers: [otlp], processors: [batch], exporters: [otlp] }
EOF
opentelemetrycollector.opentelemetry.io/side created
$ sleep 15
$ kubectl -n default run sidecar-demo --image=ghcr.io/stefanprodan/podinfo:6.7.1 --annotations="sidecar.opentelemetry.io/inject=true"
pod/sidecar-demo created
$ kubectl -n default wait --for=condition=Ready pod/sidecar-demo --timeout=120s
pod/sidecar-demo condition met
$ kubectl -n default get pod sidecar-demo -o json | jq -r '[.spec.initContainers[]?, .spec.containers[]?] | .[] | "\(.name) \(.image)"'
otc-container otel/opentelemetry-collector-contrib:0.158.0
sidecar-demo ghcr.io/stefanprodan/podinfo:6.7.1
$ kubectl -n default get opentelemetrycollector -o custom-columns=NAME:.metadata.name,MODE:.spec.mode
NAME   MODE
side   sidecar
$ kubectl -n default delete pod sidecar-demo
pod "sidecar-demo" deleted from default namespace
$ kubectl -n default delete opentelemetrycollector side
opentelemetrycollector.opentelemetry.io "side" deleted from default namespace
verify: the listing names two containers, your podinfo container and otc-container, the collector the operator injected from the annotation. It may appear under initContainers rather than containers, which is the operator's native-sidecar form and still the injected collector. If only your container is listed, the operator found no sidecar-mode collector in the namespace, and the collector listing says so.

Head sampling throws away traces before anything is known about them. Tail sampling waits for the whole trace and then keeps the errors and the slow ones, paid for by buffering every span until the decision is made.

kubectl -n tracing get opentelemetrycollector otel -o jsonpath='{.spec.config.service.pipelines.traces}' | jq
kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"processors":{"tail_sampling":{"decision_wait":"10s","policies":[{"name":"errors","type":"status_code","status_code":{"status_codes":["ERROR"]}},{"name":"slow","type":"latency","latency":{"threshold_ms":500}}]}},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","tail_sampling","batch"],"exporters":["otlp/jaeger","spanmetrics"]}}}}}}'
kubectl -n tracing rollout status deploy otel-collector --timeout=180s
kubectl -n tracing run gen-ok --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service ok-service --traces 5
kubectl -n tracing run gen-err --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service err-service --status-code Error --traces 5
sleep 30
kubectl -n tracing port-forward svc/jaeger 16686:16686 & PF1=$!
sleep 5
curl -s 'http://localhost:16686/api/services' | jq
curl -s 'http://localhost:16686/api/traces?service=err-service&limit=10' | jq '.data | length'
curl -s 'http://localhost:16686/api/traces?service=ok-service&limit=10' | jq '.data | length'
kill $PF1
# put the shared pipeline back: left in place, this drops every ordinary trace for the rest of the page
kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"processors":{"tail_sampling":null},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","batch"],"exporters":["otlp/jaeger","spanmetrics"]}}}}}}'
kubectl -n tracing rollout status deploy otel-collector --timeout=180s
outputcaptured 2026-09-12
$ kubectl -n tracing get opentelemetrycollector otel -o jsonpath='{.spec.config.service.pipelines.traces}' | jq
{
  "exporters": [
    "otlp/jaeger",
    "spanmetrics"
  ],
  "processors": [
    "k8sattributes",
    "batch"
  ],
  "receivers": [
    "otlp"
  ]
}
$ kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"processors":{"tail_sampling":{"decision_wait":"10s","policies":[{"name":"errors","type":"status_code","status_code":{"status_codes":["ERROR"]}},{"name":"slow","type":"latency","latency":{"threshold_ms":500}}]}},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","tail_sampling","batch"],"exporters":["otlp/jaeger","spanmetrics"]}}}}}}'
opentelemetrycollector.opentelemetry.io/otel patched
$ kubectl -n tracing rollout status deploy otel-collector --timeout=180s
Waiting for deployment spec update to be observed...
Waiting for deployment spec update to be observed...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
deployment "otel-collector" successfully rolled out
$ kubectl -n tracing run gen-ok --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service ok-service --traces 5
2026-09-13T14:12:19.711Z	INFO	traces/traces.go:53	starting gRPC exporter
2026-09-13T14:12:19.711Z	INFO	grpclog/component.go:69	[core] original dial target is: "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T14:12:19.711Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel created for target "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T14:12:19.711Z	INFO	channelz/trace.go:200	[core] [Channel #1] parsed dial target is: resolver.Target{URL:url.URL{Scheme:"dns", Opaque:"", User:(*url.Userinfo)(nil), Host:"", Path:"/otel-collector.tracing.svc:4317", Fragment:"", RawQuery:"", RawPath:"", RawFragment:"", ForceQuery:false, OmitHost:false}}	{"grpc_log": true}
2026-09-13T14:12:19.711Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel authority set to "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T14:12:19.711Z	INFO	traces/traces.go:125	generation of traces is limited	{"per-second": 1}
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
2026-09-13T14:12:21.712Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel Connectivity change to CONNECTING	{"grpc_log": true}
2026-09-13T14:12:21.712Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel exiting idle mode	{"grpc_log": true}
2026-09-13T14:12:21.715Z	INFO	channelz/trace.go:200	[core] [Channel #1] Resolver state updated: {
  "Addresses": [
    {
      "Addr": "10.96.126.64:4317",
      "ServerName": "",
      "Attributes": null,
      "BalancerAttributes": null,
      "Metadata": null
    }
  ],
  "Endpoints": [
    {
      "Addresses": [
        {
          "Addr": "10.96.126.64:4317",
          "ServerName": "",
          "Attributes": null,
          "BalancerAttributes": null,
          "Metadata": null
        }
      ],
      "Attributes": null
    }
  ],
  "ServiceConfig": null,
  "Attributes": null
} (resolver returned new addresses)	{"grpc_log": true}
2026-09-13T14:12:21.715Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel switches to new LB policy "pick_first"	{"grpc_log": true}
2026-09-13T14:12:21.715Z	INFO	grpclog/prefix_logger.go:42	[pick-first-leaf-lb] [pick-first-leaf-lb 0x6f66b663710] Received new config {
  "shuffleAddressList": false
... 45 more lines
$ kubectl -n tracing run gen-err --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service err-service --status-code Error --traces 5
2026-09-13T14:12:35.956Z	INFO	traces/traces.go:53	starting gRPC exporter
2026-09-13T14:12:35.956Z	INFO	grpclog/component.go:69	[core] original dial target is: "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T14:12:35.956Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel created for target "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T14:12:35.956Z	INFO	channelz/trace.go:200	[core] [Channel #1] parsed dial target is: resolver.Target{URL:url.URL{Scheme:"dns", Opaque:"", User:(*url.Userinfo)(nil), Host:"", Path:"/otel-collector.tracing.svc:4317", Fragment:"", RawQuery:"", RawPath:"", RawFragment:"", ForceQuery:false, OmitHost:false}}	{"grpc_log": true}
2026-09-13T14:12:35.957Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel authority set to "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T14:12:35.957Z	INFO	traces/traces.go:125	generation of traces is limited	{"per-second": 1}
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
2026-09-13T14:12:37.958Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel Connectivity change to CONNECTING	{"grpc_log": true}
2026-09-13T14:12:37.958Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel exiting idle mode	{"grpc_log": true}
2026-09-13T14:12:37.962Z	INFO	channelz/trace.go:200	[core] [Channel #1] Resolver state updated: {
  "Addresses": [
    {
      "Addr": "10.96.126.64:4317",
      "ServerName": "",
      "Attributes": null,
      "BalancerAttributes": null,
      "Metadata": null
    }
  ],
  "Endpoints": [
    {
      "Addresses": [
        {
          "Addr": "10.96.126.64:4317",
          "ServerName": "",
          "Attributes": null,
          "BalancerAttributes": null,
          "Metadata": null
        }
      ],
      "Attributes": null
    }
  ],
  "ServiceConfig": null,
  "Attributes": null
} (resolver returned new addresses)	{"grpc_log": true}
2026-09-13T14:12:37.962Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel switches to new LB policy "pick_first"	{"grpc_log": true}
2026-09-13T14:12:37.962Z	INFO	grpclog/prefix_logger.go:42	[pick-first-leaf-lb] [pick-first-leaf-lb 0xc8833d00000] Received new config {
  "shuffleAddressList": false
... 45 more lines
$ sleep 30
$ kubectl -n tracing port-forward svc/jaeger 16686:16686 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:16686 -> 16686
Forwarding from [::1]:16686 -> 16686
$ curl -s 'http://localhost:16686/api/services' | jq
Handling connection for 16686
{
  "data": [
    "err-service",
    "jaeger",
    "metered"
  ],
  "total": 3,
  "limit": 0,
  "offset": 0,
  "errors": null
}
$ curl -s 'http://localhost:16686/api/traces?service=err-service&limit=10' | jq '.data | length'
Handling connection for 16686
10
$ curl -s 'http://localhost:16686/api/traces?service=ok-service&limit=10' | jq '.data | length'
Handling connection for 16686
0
$ kill $PF1
$ # put the shared pipeline back: left in place, this drops every ordinary trace for the rest of the page
$ kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"processors":{"tail_sampling":null},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","batch"],"exporters":["otlp/jaeger","spanmetrics"]}}}}}}'
opentelemetrycollector.opentelemetry.io/otel patched
$ kubectl -n tracing rollout status deploy otel-collector --timeout=180s
Waiting for deployment spec update to be observed...
Waiting for deployment spec update to be observed...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
deployment "otel-collector" successfully rolled out
verify: the error service's traces are in Jaeger and the clean ones are thinner or absent, depending on whether they tripped the latency policy. Say what decision_wait costs you and why a load balancer in front of two collectors breaks tail sampling.

The k8sattributes processor is what turns a span into something you can join against a pod. It needs RBAC to do that, and without it the spans still arrive, just anonymous.

# the tail-sampling exercise above leaves the shared pipeline dropping ordinary traces
kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"processors":{"tail_sampling":null},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","batch"],"exporters":["otlp/jaeger","spanmetrics"]}}}}}}'
# there is no ClusterRoleBinding to back up here: create the RBAC so there is something to take away
kubectl apply -f - <<'EOF'
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata: { name: otel-k8sattributes }
rules:
  - { apiGroups: [""], resources: [pods, namespaces, nodes], verbs: [get, list, watch] }
  - { apiGroups: ["apps"], resources: [replicasets], verbs: [get, list, watch] }
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata: { name: otel-k8sattributes }
roleRef: { apiGroup: rbac.authorization.k8s.io, kind: ClusterRole, name: otel-k8sattributes }
subjects:
  - { kind: ServiceAccount, name: otel-collector, namespace: tracing }
EOF
kubectl -n tracing rollout restart deploy otel-collector
kubectl -n tracing rollout status deploy otel-collector --timeout=180s
kubectl -n tracing run gen-before --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service enriched --traces 3
sleep 20
kubectl -n tracing port-forward svc/jaeger 16686:16686 & PF1=$!
sleep 5
# the collector batches for 5s and Jaeger indexes after that
sleep 30
curl -s 'http://localhost:16686/api/traces?service=enriched&limit=1' | jq -r '[.data[]?.processes[]?.tags[]? | select(.key|startswith("k8s")) | .key] | unique'
kubectl delete clusterrolebinding otel-k8sattributes
kubectl -n tracing rollout restart deploy otel-collector
kubectl -n tracing rollout status deploy otel-collector --timeout=120s
kubectl -n tracing run gen-after --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service unenriched --traces 3
sleep 20
# the collector batches for 5s and Jaeger indexes after that
sleep 30
curl -s 'http://localhost:16686/api/traces?service=unenriched&limit=1' | jq -r '[.data[]?.processes[]?.tags[]? | select(.key|startswith("k8s")) | .key] | unique'
kubectl -n tracing logs deploy/otel-collector --tail=30 | grep -i -m3 forbidden
kubectl apply -f - <<'EOF'
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata: { name: otel-k8sattributes }
roleRef: { apiGroup: rbac.authorization.k8s.io, kind: ClusterRole, name: otel-k8sattributes }
subjects:
  - { kind: ServiceAccount, name: otel-collector, namespace: tracing }
EOF
kubectl -n tracing rollout restart deploy otel-collector
kubectl -n tracing rollout status deploy otel-collector --timeout=180s
kill $PF1
outputcaptured 2026-09-13
$ # the tail-sampling exercise above leaves the shared pipeline dropping ordinary traces
$ kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"processors":{"tail_sampling":null},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","batch"],"exporters":["otlp/jaeger","spanmetrics"]}}}}}}'
opentelemetrycollector.opentelemetry.io/otel patched (no change)
$ # there is no ClusterRoleBinding to back up here: create the RBAC so there is something to take away
$ kubectl apply -f - <<'EOF'
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata: { name: otel-k8sattributes }
rules:
  - { apiGroups: [""], resources: [pods, namespaces, nodes], verbs: [get, list, watch] }
  - { apiGroups: ["apps"], resources: [replicasets], verbs: [get, list, watch] }
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata: { name: otel-k8sattributes }
roleRef: { apiGroup: rbac.authorization.k8s.io, kind: ClusterRole, name: otel-k8sattributes }
subjects:
  - { kind: ServiceAccount, name: otel-collector, namespace: tracing }
EOF
clusterrole.rbac.authorization.k8s.io/otel-k8sattributes unchanged
clusterrolebinding.rbac.authorization.k8s.io/otel-k8sattributes unchanged
$ kubectl -n tracing rollout restart deploy otel-collector
deployment.apps/otel-collector restarted
$ kubectl -n tracing rollout status deploy otel-collector --timeout=180s
Waiting for deployment spec update to be observed...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
deployment "otel-collector" successfully rolled out
$ kubectl -n tracing run gen-before --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service enriched --traces 3
2026-09-13T14:31:47.568Z	INFO	traces/traces.go:53	starting gRPC exporter
2026-09-13T14:31:47.568Z	INFO	grpclog/component.go:69	[core] original dial target is: "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T14:31:47.568Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel created for target "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T14:31:47.568Z	INFO	channelz/trace.go:200	[core] [Channel #1] parsed dial target is: resolver.Target{URL:url.URL{Scheme:"dns", Opaque:"", User:(*url.Userinfo)(nil), Host:"", Path:"/otel-collector.tracing.svc:4317", Fragment:"", RawQuery:"", RawPath:"", RawFragment:"", ForceQuery:false, OmitHost:false}}	{"grpc_log": true}
2026-09-13T14:31:47.568Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel authority set to "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T14:31:47.568Z	INFO	traces/traces.go:125	generation of traces is limited	{"per-second": 1}
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
2026-09-13T14:31:49.570Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel Connectivity change to CONNECTING	{"grpc_log": true}
2026-09-13T14:31:49.570Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel exiting idle mode	{"grpc_log": true}
2026-09-13T14:31:49.573Z	INFO	channelz/trace.go:200	[core] [Channel #1] Resolver state updated: {
  "Addresses": [
    {
      "Addr": "10.96.126.64:4317",
      "ServerName": "",
      "Attributes": null,
      "BalancerAttributes": null,
      "Metadata": null
    }
  ],
  "Endpoints": [
    {
      "Addresses": [
        {
          "Addr": "10.96.126.64:4317",
          "ServerName": "",
          "Attributes": null,
          "BalancerAttributes": null,
          "Metadata": null
        }
      ],
      "Attributes": null
    }
  ],
  "ServiceConfig": null,
  "Attributes": null
} (resolver returned new addresses)	{"grpc_log": true}
2026-09-13T14:31:49.573Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel switches to new LB policy "pick_first"	{"grpc_log": true}
2026-09-13T14:31:49.573Z	INFO	grpclog/prefix_logger.go:42	[pick-first-leaf-lb] [pick-first-leaf-lb 0x351f223e8000] Received new config {
  "shuffleAddressList": false
... 45 more lines
$ sleep 20
$ kubectl -n tracing port-forward svc/jaeger 16686:16686 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:16686 -> 16686
Forwarding from [::1]:16686 -> 16686
$ # the collector batches for 5s and Jaeger indexes after that
$ sleep 30
$ curl -s 'http://localhost:16686/api/traces?service=enriched&limit=1' | jq -r '[.data[]?.processes[]?.tags[]? | select(.key|startswith("k8s")) | .key] | unique'
Handling connection for 16686
[
  "k8s.namespace.name",
  "k8s.node.name",
  "k8s.pod.name",
  "k8s.pod.start_time",
  "k8s.pod.uid"
]
$ kubectl delete clusterrolebinding otel-k8sattributes
clusterrolebinding.rbac.authorization.k8s.io "otel-k8sattributes" deleted
$ kubectl -n tracing rollout restart deploy otel-collector
deployment.apps/otel-collector restarted
$ kubectl -n tracing rollout status deploy otel-collector --timeout=120s
Waiting for deployment spec update to be observed...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
deployment "otel-collector" successfully rolled out
$ kubectl -n tracing run gen-after --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service unenriched --traces 3
2026-09-13T14:32:59.131Z	INFO	traces/traces.go:53	starting gRPC exporter
2026-09-13T14:32:59.131Z	INFO	grpclog/component.go:69	[core] original dial target is: "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T14:32:59.131Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel created for target "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T14:32:59.131Z	INFO	channelz/trace.go:200	[core] [Channel #1] parsed dial target is: resolver.Target{URL:url.URL{Scheme:"dns", Opaque:"", User:(*url.Userinfo)(nil), Host:"", Path:"/otel-collector.tracing.svc:4317", Fragment:"", RawQuery:"", RawPath:"", RawFragment:"", ForceQuery:false, OmitHost:false}}	{"grpc_log": true}
2026-09-13T14:32:59.131Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel authority set to "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T14:32:59.131Z	INFO	traces/traces.go:125	generation of traces is limited	{"per-second": 1}
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
2026-09-13T14:33:01.133Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel Connectivity change to CONNECTING	{"grpc_log": true}
2026-09-13T14:33:01.133Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel exiting idle mode	{"grpc_log": true}
2026-09-13T14:33:01.136Z	INFO	channelz/trace.go:200	[core] [Channel #1] Resolver state updated: {
  "Addresses": [
    {
      "Addr": "10.96.126.64:4317",
      "ServerName": "",
      "Attributes": null,
      "BalancerAttributes": null,
      "Metadata": null
    }
  ],
  "Endpoints": [
    {
      "Addresses": [
        {
          "Addr": "10.96.126.64:4317",
          "ServerName": "",
          "Attributes": null,
          "BalancerAttributes": null,
          "Metadata": null
        }
      ],
      "Attributes": null
    }
  ],
  "ServiceConfig": null,
  "Attributes": null
} (resolver returned new addresses)	{"grpc_log": true}
2026-09-13T14:33:01.136Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel switches to new LB policy "pick_first"	{"grpc_log": true}
2026-09-13T14:33:01.136Z	INFO	grpclog/prefix_logger.go:42	[pick-first-leaf-lb] [pick-first-leaf-lb 0x27988ddd6000] Received new config {
  "shuffleAddressList": false
... 45 more lines
$ sleep 20
$ # the collector batches for 5s and Jaeger indexes after that
$ sleep 30
$ curl -s 'http://localhost:16686/api/traces?service=unenriched&limit=1' | jq -r '[.data[]?.processes[]?.tags[]? | select(.key|startswith("k8s")) | .key] | unique'
Handling connection for 16686
[]
$ kubectl -n tracing logs deploy/otel-collector --tail=30 | grep -i -m3 forbidden
E0913 14:32:53.207760       1 reflector.go:204] "Failed to watch" err="failed to list *v1.Pod: pods is forbidden: User \"system:serviceaccount:tracing:otel-collector\" cannot list resource \"pods\" in API group \"\" at the cluster scope" logger="UnhandledError" reflector="k8s.io/client-go@v0.35.4/tools/cache/reflector.go:289" type="*v1.Pod"
E0913 14:32:54.618610       1 reflector.go:204] "Failed to watch" err="failed to list *v1.Pod: pods is forbidden: User \"system:serviceaccount:tracing:otel-collector\" cannot list resource \"pods\" in API group \"\" at the cluster scope" logger="UnhandledError" reflector="k8s.io/client-go@v0.35.4/tools/cache/reflector.go:289" type="*v1.Pod"
E0913 14:32:57.318475       1 reflector.go:204] "Failed to watch" err="failed to list *v1.Pod: pods is forbidden: User \"system:serviceaccount:tracing:otel-collector\" cannot list resource \"pods\" in API group \"\" at the cluster scope" logger="UnhandledError" reflector="k8s.io/client-go@v0.35.4/tools/cache/reflector.go:289" type="*v1.Pod"
$ kubectl apply -f - <<'EOF'
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata: { name: otel-k8sattributes }
roleRef: { apiGroup: rbac.authorization.k8s.io, kind: ClusterRole, name: otel-k8sattributes }
subjects:
  - { kind: ServiceAccount, name: otel-collector, namespace: tracing }
EOF
clusterrolebinding.rbac.authorization.k8s.io/otel-k8sattributes created
$ kubectl -n tracing rollout restart deploy otel-collector
deployment.apps/otel-collector restarted
$ kubectl -n tracing rollout status deploy otel-collector --timeout=180s
Waiting for deployment spec update to be observed...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
deployment "otel-collector" successfully rolled out
$ kill $PF1
verify: the enriched spans carry k8s.pod.name and friends, the unenriched ones do not, and the collector log has a Forbidden line explaining why. Nothing fails; you just lose the ability to ask "which pod".

The spanmetrics connector derives RED metrics from traces you are already collecting, which gives you rate, errors and duration for a service that exports no metrics at all. It is the cheapest observability win on this page.

kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"connectors":{"spanmetrics":{"histogram":{"explicit":{"buckets":["100ms","500ms","1s"]}}}},"exporters":{"prometheus":{"endpoint":"0.0.0.0:8889"}},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","batch"],"exporters":["otlp/jaeger","spanmetrics"]},"metrics/spanmetrics":{"receivers":["spanmetrics"],"exporters":["prometheus"]}}}}}}'
kubectl -n tracing rollout status deploy otel-collector --timeout=180s
kubectl -n tracing run gen-metrics --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service metered --traces 20
sleep 30
kubectl -n tracing port-forward deploy/otel-collector 8889:8889 & PF1=$!
sleep 5
curl -s localhost:8889/metrics | grep -E '^(traces_span_metrics|calls)' | head -10
kill $PF1
outputcaptured 2026-09-12
$ kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"connectors":{"spanmetrics":{"histogram":{"explicit":{"buckets":["100ms","500ms","1s"]}}}},"exporters":{"prometheus":{"endpoint":"0.0.0.0:8889"}},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","batch"],"exporters":["otlp/jaeger","spanmetrics"]},"metrics/spanmetrics":{"receivers":["spanmetrics"],"exporters":["prometheus"]}}}}}}'
opentelemetrycollector.opentelemetry.io/otel patched
$ kubectl -n tracing rollout status deploy otel-collector --timeout=180s
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment spec update to be observed...
Waiting for deployment spec update to be observed...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
deployment "otel-collector" successfully rolled out
$ kubectl -n tracing run gen-metrics --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service metered --traces 20
2026-09-13T03:23:08.558Z	INFO	traces/traces.go:53	starting gRPC exporter
2026-09-13T03:23:08.558Z	INFO	grpclog/component.go:69	[core] original dial target is: "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T03:23:08.558Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel created for target "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T03:23:08.558Z	INFO	channelz/trace.go:200	[core] [Channel #1] parsed dial target is: resolver.Target{URL:url.URL{Scheme:"dns", Opaque:"", User:(*url.Userinfo)(nil), Host:"", Path:"/otel-collector.tracing.svc:4317", Fragment:"", RawQuery:"", RawPath:"", RawFragment:"", ForceQuery:false, OmitHost:false}}	{"grpc_log": true}
2026-09-13T03:23:08.558Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel authority set to "otel-collector.tracing.svc:4317"	{"grpc_log": true}
2026-09-13T03:23:08.558Z	INFO	traces/traces.go:125	generation of traces is limited	{"per-second": 1}
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
2026-09-13T03:23:10.559Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel Connectivity change to CONNECTING	{"grpc_log": true}
2026-09-13T03:23:10.559Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel exiting idle mode	{"grpc_log": true}
2026-09-13T03:23:10.562Z	INFO	channelz/trace.go:200	[core] [Channel #1] Resolver state updated: {
  "Addresses": [
    {
      "Addr": "10.96.126.64:4317",
      "ServerName": "",
      "Attributes": null,
      "BalancerAttributes": null,
      "Metadata": null
    }
  ],
  "Endpoints": [
    {
      "Addresses": [
        {
          "Addr": "10.96.126.64:4317",
          "ServerName": "",
          "Attributes": null,
          "BalancerAttributes": null,
          "Metadata": null
        }
      ],
      "Attributes": null
    }
  ],
  "ServiceConfig": null,
  "Attributes": null
} (resolver returned new addresses)	{"grpc_log": true}
2026-09-13T03:23:10.562Z	INFO	channelz/trace.go:200	[core] [Channel #1] Channel switches to new LB policy "pick_first"	{"grpc_log": true}
2026-09-13T03:23:10.562Z	INFO	grpclog/prefix_logger.go:42	[pick-first-leaf-lb] [pick-first-leaf-lb 0x6af722e0ab0] Received new config {
  "shuffleAddressList": false
... 45 more lines
$ sleep 30
$ kubectl -n tracing port-forward deploy/otel-collector 8889:8889 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:8889 -> 8889
Forwarding from [::1]:8889 -> 8889
$ curl -s localhost:8889/metrics | grep -E '^(traces_span_metrics|calls)' | head -10
Handling connection for 8889
traces_span_metrics_calls_total{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_CLIENT",span_name="lets-go",status_code="STATUS_CODE_UNSET"} 0
traces_span_metrics_calls_total{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_SERVER",span_name="okey-dokey-0",status_code="STATUS_CODE_UNSET"} 0
traces_span_metrics_duration_milliseconds_bucket{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_CLIENT",span_name="lets-go",status_code="STATUS_CODE_UNSET",le="100"} 20
traces_span_metrics_duration_milliseconds_bucket{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_CLIENT",span_name="lets-go",status_code="STATUS_CODE_UNSET",le="500"} 20
traces_span_metrics_duration_milliseconds_bucket{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_CLIENT",span_name="lets-go",status_code="STATUS_CODE_UNSET",le="1000"} 20
traces_span_metrics_duration_milliseconds_bucket{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_CLIENT",span_name="lets-go",status_code="STATUS_CODE_UNSET",le="+Inf"} 20
traces_span_metrics_duration_milliseconds_sum{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_CLIENT",span_name="lets-go",status_code="STATUS_CODE_UNSET"} 2.460000000000001
traces_span_metrics_duration_milliseconds_count{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_CLIENT",span_name="lets-go",status_code="STATUS_CODE_UNSET"} 20
traces_span_metrics_duration_milliseconds_bucket{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_SERVER",span_name="okey-dokey-0",status_code="STATUS_CODE_UNSET",le="100"} 20
traces_span_metrics_duration_milliseconds_bucket{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_SERVER",span_name="okey-dokey-0",status_code="STATUS_CODE_UNSET",le="500"} 20
$ kill $PF1
verify: the collector's own metrics endpoint serves call counts and duration buckets labeled by service and span name.

Jaeger v2 is an OpenTelemetry Collector with storage and query extensions, which means its configuration reads like a collector's. Knowing that changes how you debug it: the pipeline is right there.

kubectl -n tracing port-forward svc/jaeger 16686:16686 & PF1=$!
sleep 5
curl -s http://localhost:16686/api/services | jq
curl -s 'http://localhost:16686/api/operations?service=err-service' | jq '.data[:5]'
kill $PF1
kubectl -n tracing get cm -o name
kubectl -n tracing get deploy jaeger -o jsonpath='{.spec.template.spec.containers[0].args}{"\n"}'
kubectl -n tracing get deploy jaeger -o jsonpath='{.spec.template.spec.containers[0].env}' | jq
outputcaptured 2026-09-12
$ kubectl -n tracing port-forward svc/jaeger 16686:16686 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:16686 -> 16686
Forwarding from [::1]:16686 -> 16686
$ curl -s http://localhost:16686/api/services | jq
Handling connection for 16686
{
  "data": [
    "err-service",
    "jaeger",
    "metered"
  ],
  "total": 3,
  "limit": 0,
  "offset": 0,
  "errors": null
}
$ curl -s 'http://localhost:16686/api/operations?service=err-service' | jq '.data[:5]'
Handling connection for 16686
[
  {
    "name": "okey-dokey-0",
    "spanKind": "server"
  },
  {
    "name": "lets-go",
    "spanKind": "client"
  }
]
$ kill $PF1
$ kubectl -n tracing get cm -o name
configmap/kube-root-ca.crt
configmap/otel-collector-3a402bbd
configmap/otel-collector-a60f73e1
configmap/otel-collector-fa2c2caf
$ kubectl -n tracing get deploy jaeger -o jsonpath='{.spec.template.spec.containers[0].args}{"\n"}'
$ kubectl -n tracing get deploy jaeger -o jsonpath='{.spec.template.spec.containers[0].env}' | jq
[
  {
    "name": "COLLECTOR_OTLP_ENABLED",
    "value": "true"
  }
]
verify: the service list matches what you have sent, and you can say where this Jaeger keeps its spans. On the lab that is memory, which is why a restart loses everything; say what you would change for a cluster anyone depends on.

Self-check

answer before opening
Jaeger shows hundreds of one-span traces. Diagnosis?

Context propagation is broken: each service starts a new trace instead of continuing the caller's. Look for a proxy stripping traceparent, a missing or mismatched propagator (W3C vs B3), a hand-rolled HTTP client that never injects, or an async boundary nobody instrumented.

Head-based versus tail-based sampling: when does the difference bite?

When the interesting traces are rare. Head sampling at 1% will usually miss the one 5-second request that mattered; tail sampling keeps traces that contain errors or exceed a latency threshold, at the cost of buffering complete traces in the collector (and needing all spans of a trace to reach the same collector instance).

Where should sampling and redaction be configured, and why there?

In the collector. It is the platform's control point: one config change applies to every service, whereas SDK-side changes need forty teams to redeploy. This is the same "interface owned by the platform" argument as section 3.1, applied to telemetry.

You annotated a Deployment for auto-instrumentation and nothing changed. Two most likely reasons?

The annotation is on the Deployment's own metadata rather than spec.template.metadata, or the pods have not restarted since (injection happens at admission, at pod creation). Third candidate: the annotation value's <namespace>/<name> does not match an existing Instrumentation CR.

How do traces, metrics and logs tie together in practice?

Exemplars link metric buckets to sample trace IDs; logs carry the trace ID as a field so you can jump from a slow span to its exact log lines; and the collector can emit all three from the same pipeline with the same resource attributes (namespace, pod, service). "Correlated by trace ID and resource attributes" is the sentence.

Your tail-sampling collector runs three replicas behind a Service, and the sampling decisions look random. Why, and what fixes it?

Spans of one trace are spread across replicas, so no single instance sees the whole trace when decision_wait expires. Put a first-tier collector with a loadbalancing exporter using routing_key: traceID in front (or run the sampling tier as one replica, or a StatefulSet the load balancer can address per pod) so every span of a trace lands on the same instance.

What is Jaeger v2 architecturally, and where does the old jaeger-agent go?

A distribution of the OpenTelemetry Collector: one binary, collector-style YAML, roles collector/query/ingester/all-in-one, with Jaeger-specific extensions for storage, the query UI on 16686 and remote sampling. Ingest is plain OTLP on 4317/4318. The agent role is deprecated; run a standard OpenTelemetry Collector as DaemonSet or sidecar where the agent used to be.

A pod has the injected init container and the right annotation, yet no spans appear. Two collector-side reasons and one network reason?

The Instrumentation's exporter.endpoint uses the wrong port for the protocol the agent speaks (Python is HTTP-only, so it must be 4318; Java defaults to gRPC 4317), or the collector pipeline has no traces pipeline wired to that receiver. Network: a default-deny NetworkPolicy in the app namespace blocks egress to the collector's namespace. Read OTEL_EXPORTER_OTLP_ENDPOINT in the pod and test the port from inside it.

How does a metric panel jump to a trace, and what has to be enabled at each layer?

Exemplars: the histogram sample carries a trace_id. The instrumented service or the spanmetrics connector emits them, Prometheus keeps them only with --enable-feature=exemplar-storage, and the Grafana datasource maps trace_id to a Tempo or Jaeger datasource via exemplarTraceIdDestinations. Logs join the same picture when the log line carries the trace ID and Loki's derivedFields links it.

Docs to know your way around

study time, not exam time
  • opentelemetry.io: the Concepts pages (signals, context propagation), the collector's receivers/processors/exporters reference, and the Kubernetes operator's Instrumentation injection docs.
  • jaegertracing.io: mostly the UI is self-explanatory; know that the HTTP API exists for scripted verification.
  • w3.org/TR/trace-context: one page, and it is the header everything agrees on.
  • opentelemetry.io/docs/collector/configuration: sections, connectors and pipelines; the processor READMEs in github.com/open-telemetry/opentelemetry-collector-contrib (tailsamplingprocessor, k8sattributesprocessor) hold the field defaults.
  • opentelemetry.io/docs/platforms/kubernetes/operator/automatic: annotation values, container-names, the Go and Python caveats; the troubleshooting page beside it is the injection diagnostic ladder.
  • jaegertracing.io/docs/latest/architecture: roles, storage backends and the "we recommend the OpenTelemetry Collector instead of the agent" statement.