Tracing answers the one question metrics and logs cannot: where, along a request's path through many services, did the time go or the failure start. The vocabulary is small and the exam expects it exact.
make up obsOrientation
Two things make tracing tasks solvable: knowing the exact OTLP address telemetry should go to, and knowing that a trace is only a trace if context propagated. Everything else is UI.
The OTel operator lives in opentelemetry-operator-system, the collector in tracing, Jaeger alongside it. So the in-cluster OTLP endpoints are otel-collector.tracing.svc:4317 (gRPC) and :4318 (HTTP). Memorize the form <svc>.<ns>.svc:4317; every tracing task starts by knowing that address.
Vocabulary and plumbing
A trace is one request's journey. A span is one named, timed operation within it, carrying a trace ID, its own span ID, a parent span ID, a status, attributes and events. A trace is therefore a tree, and the root span is the entry point.
Context propagation is what links spans across service boundaries: the caller injects the trace context into headers, the callee extracts it and continues the trace instead of starting a new one. The standard header is W3C traceparent (plus tracestate, and baggage for user-defined key/values that travel with the request). Older stacks use B3 headers; a mixed fleet needs a propagator configured for both.
Broken propagation: every service reports spans, and every span is its own orphan trace. Symptom in Jaeger: many single-span traces, one per service, none joined. Causes: a proxy or gateway stripping unknown headers, an SDK without the right propagator, a manual HTTP client that never injects, or an async hop (queue, cron) where nobody thought to carry the context. When you see one-span traces from an instrumented system, name that cause immediately.
The OpenTelemetry pipeline
app (SDK or auto-instrumentation agent)
│ OTLP gRPC :4317 / HTTP :4318
▼
Collector receivers ──▶ processors ──▶ exporters
otlp, jaeger, batch, memory_limiter, otlp/jaeger,
prometheus, attributes, k8sattributes, debug,
filelog tail_sampling prometheus
▼
backend Jaeger (traces) · Prometheus (metrics) · Loki (logs)
The collector is the platform team's control point: sampling, batching, redaction and fan-out happen there, not in application code. That is the platform-engineering answer to "how do we control telemetry cost": you change one deployment, not forty services.
| Sampling | Decides | Trade |
|---|---|---|
| Head-based | at the start of the trace, usually a fixed probability | cheap and simple; you may drop the one slow request you needed |
| Tail-based | after the trace completes, on its properties (errors, latency) | keeps the interesting traces; the collector must buffer whole traces, so it costs memory and needs all spans of a trace at one collector |
Kubernetes-side: two CRDs
OpenTelemetryCollector: the operator deploys and configures a collector (as a Deployment, DaemonSet, StatefulSet or Sidecar). Read its config to see the receiver/exporter chain.Instrumentation: the zero-code path. The lab ships one namedautoindefault; annotating a pod spec withinstrumentation.opentelemetry.io/inject-java: "default/auto"(or-python,-nodejs,-dotnet) makes a webhook inject the language agent as an init container at admission.
Annotate an existing Deployment and nothing happens to the running pods; they must restart to pick it up. And the annotation belongs on the pod template, not on the Deployment's own metadata. Those two mistakes account for most "the annotation does nothing" reports, and both are one-line fixes once you know to look.
The collector, field by field
A collector config has four component sections (receivers, processors, exporters, connectors), an extensions section for things that are not in the data path (health_check, pprof, zpages), and a service block that wires them into named pipelines per signal. Defining a component does nothing until a pipeline lists it. Components can be declared more than once with a suffix (otlp/2, otlphttp/loki), and a connector is both an exporter of one pipeline and a receiver of another, which is how spanmetrics turns traces into RED metrics without touching the application.
receivers:
otlp: { protocols: { grpc: { endpoint: 0.0.0.0:4317 }, http: { endpoint: 0.0.0.0:4318 } } }
processors:
memory_limiter: { check_interval: 1s, limit_mib: 800, spike_limit_mib: 160 }
k8sattributes: {}
batch: {}
exporters:
otlp/jaeger: { endpoint: jaeger-collector.tracing.svc:4317, tls: { insecure: true } }
prometheus: { endpoint: 0.0.0.0:8889 }
connectors:
spanmetrics: {}
service:
pipelines:
traces: { receivers: [otlp], processors: [memory_limiter, k8sattributes, batch], exporters: [otlp/jaeger, spanmetrics] }
metrics: { receivers: [spanmetrics], processors: [batch], exporters: [prometheus] }| Processor | Fields that matter | Rule |
|---|---|---|
| memory_limiter | check_interval, limit_mib (hard), spike_limit_mib (soft = hard minus spike; start at 20% of hard) or limit_percentage | first in every pipeline so back-pressure reaches the receiver; set GOMEMLIMIT as well |
| batch | send_batch_size (8192), timeout (200ms), send_batch_max_size | last before the exporters; reduces request count and exporter cost |
| k8sattributes | pod_association (match by k8s.pod.ip or k8s.pod.name+k8s.namespace.name), extract.metadata, extract.labels | needs RBAC to list pods, replicasets, namespaces; without it spans carry no k8s.* resource attributes and Loki/Tempo correlation breaks |
| resource / attributes | actions: [{key, action: insert|upsert|delete|hash}] | redaction (delete on http.request.header.authorization) belongs here, not in forty SDKs |
| filter | OTTL conditions, e.g. drop spans where attributes["http.route"] == "/healthz" | dropping individual spans breaks trace trees; prefer sampling whole traces |
| tail_sampling | decision_wait (30s), num_traces (50000), policies[] of type always_sample, latency, status_code, probabilistic, rate_limiting, string_attribute, ottl_condition, and, composite | every span of a trace must reach the same instance: put a loadbalancing exporter with routing_key: traceID in a first tier, or run one replica |
| probabilistic_sampler | sampling_percentage | head sampling in the collector; cheap, blind to errors |
Two deployment roles: an agent (DaemonSet or sidecar) close to the workload that enriches and batches, and a gateway (Deployment behind a Service) that samples, redacts and fans out. Tail sampling and rate limiting belong in the gateway. The debug exporter (verbosity: detailed) is the fastest way to see whether spans arrive at all; the old logging exporter is gone from current builds, so a config that still names it fails to start with an unknown-exporter error.
An exporter's endpoint is a host:port for otlp (gRPC) and a URL for otlphttp. Sending gRPC to 4318 or HTTP to 4317 produces connection resets or 415/unimplemented errors rather than a clear message. The Python auto-instrumentation talks http/protobuf only, so its Instrumentation endpoint must be the 4318 one.
Operator CRDs, Jaeger v2 and how traces go wrong
OpenTelemetryCollector and Instrumentation
spec.mode:deployment(default),daemonset,statefulset(required for the target allocator and for sharded tail sampling),sidecar. A sidecar collector is injected into pods annotatedsidecar.opentelemetry.io/inject: "true"(or the collector's name), a second, separate annotation from the language injection ones.spec.configis the collector YAML above; the operator renders it into a ConfigMap and a Service whose port names follow the receivers (otlp-grpc,otlp-http).spec.targetAllocatorlets a StatefulSet of collectors share Prometheus scrape targets discovered from ServiceMonitors and PodMonitors.Instrumentationfields:exporter.endpoint,propagators(tracecontext,baggage,b3),sampler.type(parentbased_traceidratio) andsampler.argument("0.25"), plus per-language blocks (java.image,python.env,nodejs,dotnet,go) that can pin an agent image or addOTEL_*environment variables.- Annotation values for
instrumentation.opentelemetry.io/inject-<lang>:"true"(the single Instrumentation in the pod's namespace; ambiguous if there are two),"name"(same namespace),"namespace/name","false". Multi-container pods needinstrumentation.opentelemetry.io/container-names: "app,worker"or the operator instruments only the first container.inject-sdkinjects environment variables only, for a language with no agent. Go injection is eBPF: it needsotel-go-auto-target-exepointing at the binary and a privileged sidecar, which arestrictednamespace refuses.
Jaeger v2
Jaeger v2 is a distribution of the OpenTelemetry Collector: one jaeger binary configured with collector-style YAML, playing the roles collector, query, ingester (reads Kafka) or all-in-one. Its own components are the jaeger_storage extension (backends: in-memory, Badger for single-node persistence, Cassandra, Elasticsearch, OpenSearch, and ClickHouse in newer releases), the jaeger_query extension serving the UI and API on 16686, the remote_sampling extension (serves per-service sampling strategies to SDKs, static or adaptive) and the adaptivesampling processor. Ingest is the standard otlp receiver on 4317/4318; the v1 jaeger-agent and its Thrift ports are gone, and the docs tell you to run an OpenTelemetry Collector where you used to run the agent. The HTTP query API the lab used (/api/traces?service=) is unchanged, so scripted verification still works.
Grafana Tempo is the other trace store you will see named: object storage, no index beyond trace ID plus a search over recent blocks, TraceQL for queries, and a metrics-generator that emits span metrics and service graphs into Prometheus, the same job the spanmetrics connector does in the collector.
Exemplars and semantic conventions
An exemplar is a trace ID attached to a histogram bucket sample; the OpenMetrics text form is # {trace_id="abc..."} after the sample. Prometheus stores them only with --enable-feature=exemplar-storage (operator: spec.enableFeatures), and Grafana links them to Tempo or Jaeger through the datasource's exemplarTraceIdDestinations. The spanmetrics connector produces exemplars by default. Semantic conventions are the attribute names that make correlation possible: service.name (the one every SDK must set, or you get unknown_service), service.namespace, k8s.namespace.name, k8s.pod.name, k8s.deployment.name, http.request.method, http.response.status_code, url.path. A dashboard that groups by k8s_namespace_name is reading these attributes after the Prometheus exporter replaced dots with underscores.
Failure modes and their tells
| Tell | Cause | Where to look |
|---|---|---|
| hundreds of single-span traces | propagation broken (proxy strips traceparent, propagator mismatch W3C vs B3, hand-rolled client) | the caller's outbound headers; propagators in the Instrumentation |
| a child span starts before its parent | clock skew between nodes | node NTP; the trace is real |
| a gap in the waterfall with no span | an uninstrumented hop (queue consumer, cron, a library without instrumentation) or the span was dropped by a filter processor | the service at the gap; the collector config |
| the slow request is never in the store | head sampling dropped it | switch to tail sampling on latency/status_code, or raise the ratio |
service shows as unknown_service | OTEL_SERVICE_NAME / service.name not set | the Instrumentation resource block or the pod env |
spans arrive, no k8s.* attributes | k8sattributes missing, or its RBAC, or the pod IP is hidden behind a gateway so pod_association cannot match | collector logs; associate on k8s.pod.name instead of IP |
collector logs data refused due to high memory usage | memory_limiter hard limit hit | raise limits or sample earlier; this is working as designed |
| injected init container present, no spans | endpoint or protocol wrong in the Instrumentation (gRPC vs HTTP port), or a NetworkPolicy blocking egress to the collector namespace | kubectl exec and read the OTEL_EXPORTER_OTLP_ENDPOINT env; test the port |
"Deploy a collector that receives OTLP and forwards traces to Jaeger" is an OpenTelemetryCollector with a receiver, a pipeline and an otlp exporter pointed at the Jaeger Service on 4317. "Auto-instrument this Deployment" is an Instrumentation plus one pod-template annotation and a rollout restart. "Keep only traces with errors or over one second" is tail_sampling with status_code and latency policies, and the reminder that it needs all spans on one collector.
Exercises
kubectl -n tracing get cm -o yaml | grep -A30 'receivers:' (or the OpenTelemetryCollector CR if present). Draw the chain on paper: which receivers listen, which exporter points at Jaeger, what processors sit between.
otel-collector.tracing.svc:4317 or the HTTP twin). Every tracing task starts by knowing that address.telemetrygen is the collector project's own traffic generator, and it makes tracing exercises deterministic:
kubectl run telemetrygen --restart=Never \
--image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest \
-- traces --otlp-endpoint otel-collector.tracing.svc:4317 --otlp-insecure \
--service curriculum-drill --traces 20 --child-spans 3outputcaptured 2026-08-26
$ kubectl run telemetrygen --restart=Never \
--image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest \
-- traces --otlp-endpoint otel-collector.tracing.svc:4317 --otlp-insecure \
--service curriculum-drill --traces 20 --child-spans 3
pod/telemetrygen createdmake urls): service curriculum-drill appears in the dropdown, 20 traces, each of 4 spans (one root plus three children, two levels). Then verify the API way, because exams grade with curl: curl -s 'http://<jaeger>/api/traces?service=curriculum-drill&limit=1' | jq '.data[0].spans | length' returns 4.Open one trace, read the waterfall: parent/child structure, per-span duration, attributes on each span. Answer for that trace: which span is the critical path, and what would you look at next if the leaf span were slow (that span's service's logs, at that timestamp; this is why traces carry IDs you can grep logs for).
Run a small Java service (ghcr.io/open-telemetry/opentelemetry-java-examples images work, or any Spring Boot sample), annotate its pod template with instrumentation.opentelemetry.io/inject-java: "default/auto", restart, and hit its endpoint. If it does not appear, the diagnostic ladder is: annotation on the pod template (not the Deployment metadata), pod restarted since annotating, Instrumentation CR namespace/name correct in the annotation value, operator webhook alive.
kubectl describe pod shows the injected init container, and the service appears in Jaeger with HTTP spans you never wrote.No cluster needed: service A calls B through a proxy that strips unknown headers. Describe what Jaeger shows and which header must survive.
traceparent and predicts orphaned single-service traces, the concept is yours.A sidecar collector buffers locally so the app never blocks on a network hop, and the operator injects it from an annotation. Seeing the injected container is the whole point; the annotation is two words.
kubectl apply -f - <<'EOF'
apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata: { name: side, namespace: default }
spec:
mode: sidecar
config:
receivers:
otlp:
protocols:
grpc: { endpoint: 0.0.0.0:4317 }
http: { endpoint: 0.0.0.0:4318 }
processors:
batch: { timeout: 5s }
exporters:
otlp:
endpoint: otel-collector.tracing.svc:4317
tls: { insecure: true }
service:
pipelines:
traces: { receivers: [otlp], processors: [batch], exporters: [otlp] }
EOF
sleep 15
kubectl -n default run sidecar-demo --image=ghcr.io/stefanprodan/podinfo:6.7.1 --annotations="sidecar.opentelemetry.io/inject=true"
kubectl -n default wait --for=condition=Ready pod/sidecar-demo --timeout=120s
kubectl -n default get pod sidecar-demo -o json | jq -r '[.spec.initContainers[]?, .spec.containers[]?] | .[] | "\(.name) \(.image)"'
kubectl -n default get opentelemetrycollector -o custom-columns=NAME:.metadata.name,MODE:.spec.mode
kubectl -n default delete pod sidecar-demo
kubectl -n default delete opentelemetrycollector sideoutputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: opentelemetry.io/v1beta1
kind: OpenTelemetryCollector
metadata: { name: side, namespace: default }
spec:
mode: sidecar
config:
receivers:
otlp:
protocols:
grpc: { endpoint: 0.0.0.0:4317 }
http: { endpoint: 0.0.0.0:4318 }
processors:
batch: { timeout: 5s }
exporters:
otlp:
endpoint: otel-collector.tracing.svc:4317
tls: { insecure: true }
service:
pipelines:
traces: { receivers: [otlp], processors: [batch], exporters: [otlp] }
EOF
opentelemetrycollector.opentelemetry.io/side created
$ sleep 15
$ kubectl -n default run sidecar-demo --image=ghcr.io/stefanprodan/podinfo:6.7.1 --annotations="sidecar.opentelemetry.io/inject=true"
pod/sidecar-demo created
$ kubectl -n default wait --for=condition=Ready pod/sidecar-demo --timeout=120s
pod/sidecar-demo condition met
$ kubectl -n default get pod sidecar-demo -o json | jq -r '[.spec.initContainers[]?, .spec.containers[]?] | .[] | "\(.name) \(.image)"'
otc-container otel/opentelemetry-collector-contrib:0.158.0
sidecar-demo ghcr.io/stefanprodan/podinfo:6.7.1
$ kubectl -n default get opentelemetrycollector -o custom-columns=NAME:.metadata.name,MODE:.spec.mode
NAME MODE
side sidecar
$ kubectl -n default delete pod sidecar-demo
pod "sidecar-demo" deleted from default namespace
$ kubectl -n default delete opentelemetrycollector side
opentelemetrycollector.opentelemetry.io "side" deleted from default namespaceotc-container, the collector the operator injected from the annotation. It may appear under initContainers rather than containers, which is the operator's native-sidecar form and still the injected collector. If only your container is listed, the operator found no sidecar-mode collector in the namespace, and the collector listing says so.Head sampling throws away traces before anything is known about them. Tail sampling waits for the whole trace and then keeps the errors and the slow ones, paid for by buffering every span until the decision is made.
kubectl -n tracing get opentelemetrycollector otel -o jsonpath='{.spec.config.service.pipelines.traces}' | jq
kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"processors":{"tail_sampling":{"decision_wait":"10s","policies":[{"name":"errors","type":"status_code","status_code":{"status_codes":["ERROR"]}},{"name":"slow","type":"latency","latency":{"threshold_ms":500}}]}},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","tail_sampling","batch"],"exporters":["otlp/jaeger","spanmetrics"]}}}}}}'
kubectl -n tracing rollout status deploy otel-collector --timeout=180s
kubectl -n tracing run gen-ok --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service ok-service --traces 5
kubectl -n tracing run gen-err --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service err-service --status-code Error --traces 5
sleep 30
kubectl -n tracing port-forward svc/jaeger 16686:16686 & PF1=$!
sleep 5
curl -s 'http://localhost:16686/api/services' | jq
curl -s 'http://localhost:16686/api/traces?service=err-service&limit=10' | jq '.data | length'
curl -s 'http://localhost:16686/api/traces?service=ok-service&limit=10' | jq '.data | length'
kill $PF1
# put the shared pipeline back: left in place, this drops every ordinary trace for the rest of the page
kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"processors":{"tail_sampling":null},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","batch"],"exporters":["otlp/jaeger","spanmetrics"]}}}}}}'
kubectl -n tracing rollout status deploy otel-collector --timeout=180soutputcaptured 2026-09-12
$ kubectl -n tracing get opentelemetrycollector otel -o jsonpath='{.spec.config.service.pipelines.traces}' | jq
{
"exporters": [
"otlp/jaeger",
"spanmetrics"
],
"processors": [
"k8sattributes",
"batch"
],
"receivers": [
"otlp"
]
}
$ kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"processors":{"tail_sampling":{"decision_wait":"10s","policies":[{"name":"errors","type":"status_code","status_code":{"status_codes":["ERROR"]}},{"name":"slow","type":"latency","latency":{"threshold_ms":500}}]}},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","tail_sampling","batch"],"exporters":["otlp/jaeger","spanmetrics"]}}}}}}'
opentelemetrycollector.opentelemetry.io/otel patched
$ kubectl -n tracing rollout status deploy otel-collector --timeout=180s
Waiting for deployment spec update to be observed...
Waiting for deployment spec update to be observed...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
deployment "otel-collector" successfully rolled out
$ kubectl -n tracing run gen-ok --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service ok-service --traces 5
2026-09-13T14:12:19.711Z INFO traces/traces.go:53 starting gRPC exporter
2026-09-13T14:12:19.711Z INFO grpclog/component.go:69 [core] original dial target is: "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T14:12:19.711Z INFO channelz/trace.go:200 [core] [Channel #1] Channel created for target "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T14:12:19.711Z INFO channelz/trace.go:200 [core] [Channel #1] parsed dial target is: resolver.Target{URL:url.URL{Scheme:"dns", Opaque:"", User:(*url.Userinfo)(nil), Host:"", Path:"/otel-collector.tracing.svc:4317", Fragment:"", RawQuery:"", RawPath:"", RawFragment:"", ForceQuery:false, OmitHost:false}} {"grpc_log": true}
2026-09-13T14:12:19.711Z INFO channelz/trace.go:200 [core] [Channel #1] Channel authority set to "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T14:12:19.711Z INFO traces/traces.go:125 generation of traces is limited {"per-second": 1}
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
2026-09-13T14:12:21.712Z INFO channelz/trace.go:200 [core] [Channel #1] Channel Connectivity change to CONNECTING {"grpc_log": true}
2026-09-13T14:12:21.712Z INFO channelz/trace.go:200 [core] [Channel #1] Channel exiting idle mode {"grpc_log": true}
2026-09-13T14:12:21.715Z INFO channelz/trace.go:200 [core] [Channel #1] Resolver state updated: {
"Addresses": [
{
"Addr": "10.96.126.64:4317",
"ServerName": "",
"Attributes": null,
"BalancerAttributes": null,
"Metadata": null
}
],
"Endpoints": [
{
"Addresses": [
{
"Addr": "10.96.126.64:4317",
"ServerName": "",
"Attributes": null,
"BalancerAttributes": null,
"Metadata": null
}
],
"Attributes": null
}
],
"ServiceConfig": null,
"Attributes": null
} (resolver returned new addresses) {"grpc_log": true}
2026-09-13T14:12:21.715Z INFO channelz/trace.go:200 [core] [Channel #1] Channel switches to new LB policy "pick_first" {"grpc_log": true}
2026-09-13T14:12:21.715Z INFO grpclog/prefix_logger.go:42 [pick-first-leaf-lb] [pick-first-leaf-lb 0x6f66b663710] Received new config {
"shuffleAddressList": false
... 45 more lines
$ kubectl -n tracing run gen-err --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service err-service --status-code Error --traces 5
2026-09-13T14:12:35.956Z INFO traces/traces.go:53 starting gRPC exporter
2026-09-13T14:12:35.956Z INFO grpclog/component.go:69 [core] original dial target is: "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T14:12:35.956Z INFO channelz/trace.go:200 [core] [Channel #1] Channel created for target "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T14:12:35.956Z INFO channelz/trace.go:200 [core] [Channel #1] parsed dial target is: resolver.Target{URL:url.URL{Scheme:"dns", Opaque:"", User:(*url.Userinfo)(nil), Host:"", Path:"/otel-collector.tracing.svc:4317", Fragment:"", RawQuery:"", RawPath:"", RawFragment:"", ForceQuery:false, OmitHost:false}} {"grpc_log": true}
2026-09-13T14:12:35.957Z INFO channelz/trace.go:200 [core] [Channel #1] Channel authority set to "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T14:12:35.957Z INFO traces/traces.go:125 generation of traces is limited {"per-second": 1}
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
2026-09-13T14:12:37.958Z INFO channelz/trace.go:200 [core] [Channel #1] Channel Connectivity change to CONNECTING {"grpc_log": true}
2026-09-13T14:12:37.958Z INFO channelz/trace.go:200 [core] [Channel #1] Channel exiting idle mode {"grpc_log": true}
2026-09-13T14:12:37.962Z INFO channelz/trace.go:200 [core] [Channel #1] Resolver state updated: {
"Addresses": [
{
"Addr": "10.96.126.64:4317",
"ServerName": "",
"Attributes": null,
"BalancerAttributes": null,
"Metadata": null
}
],
"Endpoints": [
{
"Addresses": [
{
"Addr": "10.96.126.64:4317",
"ServerName": "",
"Attributes": null,
"BalancerAttributes": null,
"Metadata": null
}
],
"Attributes": null
}
],
"ServiceConfig": null,
"Attributes": null
} (resolver returned new addresses) {"grpc_log": true}
2026-09-13T14:12:37.962Z INFO channelz/trace.go:200 [core] [Channel #1] Channel switches to new LB policy "pick_first" {"grpc_log": true}
2026-09-13T14:12:37.962Z INFO grpclog/prefix_logger.go:42 [pick-first-leaf-lb] [pick-first-leaf-lb 0xc8833d00000] Received new config {
"shuffleAddressList": false
... 45 more lines
$ sleep 30
$ kubectl -n tracing port-forward svc/jaeger 16686:16686 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:16686 -> 16686
Forwarding from [::1]:16686 -> 16686
$ curl -s 'http://localhost:16686/api/services' | jq
Handling connection for 16686
{
"data": [
"err-service",
"jaeger",
"metered"
],
"total": 3,
"limit": 0,
"offset": 0,
"errors": null
}
$ curl -s 'http://localhost:16686/api/traces?service=err-service&limit=10' | jq '.data | length'
Handling connection for 16686
10
$ curl -s 'http://localhost:16686/api/traces?service=ok-service&limit=10' | jq '.data | length'
Handling connection for 16686
0
$ kill $PF1
$ # put the shared pipeline back: left in place, this drops every ordinary trace for the rest of the page
$ kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"processors":{"tail_sampling":null},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","batch"],"exporters":["otlp/jaeger","spanmetrics"]}}}}}}'
opentelemetrycollector.opentelemetry.io/otel patched
$ kubectl -n tracing rollout status deploy otel-collector --timeout=180s
Waiting for deployment spec update to be observed...
Waiting for deployment spec update to be observed...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
deployment "otel-collector" successfully rolled outdecision_wait costs you and why a load balancer in front of two collectors breaks tail sampling.The k8sattributes processor is what turns a span into something you can join against a pod. It needs RBAC to do that, and without it the spans still arrive, just anonymous.
# the tail-sampling exercise above leaves the shared pipeline dropping ordinary traces
kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"processors":{"tail_sampling":null},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","batch"],"exporters":["otlp/jaeger","spanmetrics"]}}}}}}'
# there is no ClusterRoleBinding to back up here: create the RBAC so there is something to take away
kubectl apply -f - <<'EOF'
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata: { name: otel-k8sattributes }
rules:
- { apiGroups: [""], resources: [pods, namespaces, nodes], verbs: [get, list, watch] }
- { apiGroups: ["apps"], resources: [replicasets], verbs: [get, list, watch] }
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata: { name: otel-k8sattributes }
roleRef: { apiGroup: rbac.authorization.k8s.io, kind: ClusterRole, name: otel-k8sattributes }
subjects:
- { kind: ServiceAccount, name: otel-collector, namespace: tracing }
EOF
kubectl -n tracing rollout restart deploy otel-collector
kubectl -n tracing rollout status deploy otel-collector --timeout=180s
kubectl -n tracing run gen-before --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service enriched --traces 3
sleep 20
kubectl -n tracing port-forward svc/jaeger 16686:16686 & PF1=$!
sleep 5
# the collector batches for 5s and Jaeger indexes after that
sleep 30
curl -s 'http://localhost:16686/api/traces?service=enriched&limit=1' | jq -r '[.data[]?.processes[]?.tags[]? | select(.key|startswith("k8s")) | .key] | unique'
kubectl delete clusterrolebinding otel-k8sattributes
kubectl -n tracing rollout restart deploy otel-collector
kubectl -n tracing rollout status deploy otel-collector --timeout=120s
kubectl -n tracing run gen-after --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service unenriched --traces 3
sleep 20
# the collector batches for 5s and Jaeger indexes after that
sleep 30
curl -s 'http://localhost:16686/api/traces?service=unenriched&limit=1' | jq -r '[.data[]?.processes[]?.tags[]? | select(.key|startswith("k8s")) | .key] | unique'
kubectl -n tracing logs deploy/otel-collector --tail=30 | grep -i -m3 forbidden
kubectl apply -f - <<'EOF'
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata: { name: otel-k8sattributes }
roleRef: { apiGroup: rbac.authorization.k8s.io, kind: ClusterRole, name: otel-k8sattributes }
subjects:
- { kind: ServiceAccount, name: otel-collector, namespace: tracing }
EOF
kubectl -n tracing rollout restart deploy otel-collector
kubectl -n tracing rollout status deploy otel-collector --timeout=180s
kill $PF1outputcaptured 2026-09-13
$ # the tail-sampling exercise above leaves the shared pipeline dropping ordinary traces
$ kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"processors":{"tail_sampling":null},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","batch"],"exporters":["otlp/jaeger","spanmetrics"]}}}}}}'
opentelemetrycollector.opentelemetry.io/otel patched (no change)
$ # there is no ClusterRoleBinding to back up here: create the RBAC so there is something to take away
$ kubectl apply -f - <<'EOF'
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata: { name: otel-k8sattributes }
rules:
- { apiGroups: [""], resources: [pods, namespaces, nodes], verbs: [get, list, watch] }
- { apiGroups: ["apps"], resources: [replicasets], verbs: [get, list, watch] }
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata: { name: otel-k8sattributes }
roleRef: { apiGroup: rbac.authorization.k8s.io, kind: ClusterRole, name: otel-k8sattributes }
subjects:
- { kind: ServiceAccount, name: otel-collector, namespace: tracing }
EOF
clusterrole.rbac.authorization.k8s.io/otel-k8sattributes unchanged
clusterrolebinding.rbac.authorization.k8s.io/otel-k8sattributes unchanged
$ kubectl -n tracing rollout restart deploy otel-collector
deployment.apps/otel-collector restarted
$ kubectl -n tracing rollout status deploy otel-collector --timeout=180s
Waiting for deployment spec update to be observed...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
deployment "otel-collector" successfully rolled out
$ kubectl -n tracing run gen-before --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service enriched --traces 3
2026-09-13T14:31:47.568Z INFO traces/traces.go:53 starting gRPC exporter
2026-09-13T14:31:47.568Z INFO grpclog/component.go:69 [core] original dial target is: "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T14:31:47.568Z INFO channelz/trace.go:200 [core] [Channel #1] Channel created for target "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T14:31:47.568Z INFO channelz/trace.go:200 [core] [Channel #1] parsed dial target is: resolver.Target{URL:url.URL{Scheme:"dns", Opaque:"", User:(*url.Userinfo)(nil), Host:"", Path:"/otel-collector.tracing.svc:4317", Fragment:"", RawQuery:"", RawPath:"", RawFragment:"", ForceQuery:false, OmitHost:false}} {"grpc_log": true}
2026-09-13T14:31:47.568Z INFO channelz/trace.go:200 [core] [Channel #1] Channel authority set to "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T14:31:47.568Z INFO traces/traces.go:125 generation of traces is limited {"per-second": 1}
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
2026-09-13T14:31:49.570Z INFO channelz/trace.go:200 [core] [Channel #1] Channel Connectivity change to CONNECTING {"grpc_log": true}
2026-09-13T14:31:49.570Z INFO channelz/trace.go:200 [core] [Channel #1] Channel exiting idle mode {"grpc_log": true}
2026-09-13T14:31:49.573Z INFO channelz/trace.go:200 [core] [Channel #1] Resolver state updated: {
"Addresses": [
{
"Addr": "10.96.126.64:4317",
"ServerName": "",
"Attributes": null,
"BalancerAttributes": null,
"Metadata": null
}
],
"Endpoints": [
{
"Addresses": [
{
"Addr": "10.96.126.64:4317",
"ServerName": "",
"Attributes": null,
"BalancerAttributes": null,
"Metadata": null
}
],
"Attributes": null
}
],
"ServiceConfig": null,
"Attributes": null
} (resolver returned new addresses) {"grpc_log": true}
2026-09-13T14:31:49.573Z INFO channelz/trace.go:200 [core] [Channel #1] Channel switches to new LB policy "pick_first" {"grpc_log": true}
2026-09-13T14:31:49.573Z INFO grpclog/prefix_logger.go:42 [pick-first-leaf-lb] [pick-first-leaf-lb 0x351f223e8000] Received new config {
"shuffleAddressList": false
... 45 more lines
$ sleep 20
$ kubectl -n tracing port-forward svc/jaeger 16686:16686 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:16686 -> 16686
Forwarding from [::1]:16686 -> 16686
$ # the collector batches for 5s and Jaeger indexes after that
$ sleep 30
$ curl -s 'http://localhost:16686/api/traces?service=enriched&limit=1' | jq -r '[.data[]?.processes[]?.tags[]? | select(.key|startswith("k8s")) | .key] | unique'
Handling connection for 16686
[
"k8s.namespace.name",
"k8s.node.name",
"k8s.pod.name",
"k8s.pod.start_time",
"k8s.pod.uid"
]
$ kubectl delete clusterrolebinding otel-k8sattributes
clusterrolebinding.rbac.authorization.k8s.io "otel-k8sattributes" deleted
$ kubectl -n tracing rollout restart deploy otel-collector
deployment.apps/otel-collector restarted
$ kubectl -n tracing rollout status deploy otel-collector --timeout=120s
Waiting for deployment spec update to be observed...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
deployment "otel-collector" successfully rolled out
$ kubectl -n tracing run gen-after --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service unenriched --traces 3
2026-09-13T14:32:59.131Z INFO traces/traces.go:53 starting gRPC exporter
2026-09-13T14:32:59.131Z INFO grpclog/component.go:69 [core] original dial target is: "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T14:32:59.131Z INFO channelz/trace.go:200 [core] [Channel #1] Channel created for target "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T14:32:59.131Z INFO channelz/trace.go:200 [core] [Channel #1] parsed dial target is: resolver.Target{URL:url.URL{Scheme:"dns", Opaque:"", User:(*url.Userinfo)(nil), Host:"", Path:"/otel-collector.tracing.svc:4317", Fragment:"", RawQuery:"", RawPath:"", RawFragment:"", ForceQuery:false, OmitHost:false}} {"grpc_log": true}
2026-09-13T14:32:59.131Z INFO channelz/trace.go:200 [core] [Channel #1] Channel authority set to "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T14:32:59.131Z INFO traces/traces.go:125 generation of traces is limited {"per-second": 1}
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
2026-09-13T14:33:01.133Z INFO channelz/trace.go:200 [core] [Channel #1] Channel Connectivity change to CONNECTING {"grpc_log": true}
2026-09-13T14:33:01.133Z INFO channelz/trace.go:200 [core] [Channel #1] Channel exiting idle mode {"grpc_log": true}
2026-09-13T14:33:01.136Z INFO channelz/trace.go:200 [core] [Channel #1] Resolver state updated: {
"Addresses": [
{
"Addr": "10.96.126.64:4317",
"ServerName": "",
"Attributes": null,
"BalancerAttributes": null,
"Metadata": null
}
],
"Endpoints": [
{
"Addresses": [
{
"Addr": "10.96.126.64:4317",
"ServerName": "",
"Attributes": null,
"BalancerAttributes": null,
"Metadata": null
}
],
"Attributes": null
}
],
"ServiceConfig": null,
"Attributes": null
} (resolver returned new addresses) {"grpc_log": true}
2026-09-13T14:33:01.136Z INFO channelz/trace.go:200 [core] [Channel #1] Channel switches to new LB policy "pick_first" {"grpc_log": true}
2026-09-13T14:33:01.136Z INFO grpclog/prefix_logger.go:42 [pick-first-leaf-lb] [pick-first-leaf-lb 0x27988ddd6000] Received new config {
"shuffleAddressList": false
... 45 more lines
$ sleep 20
$ # the collector batches for 5s and Jaeger indexes after that
$ sleep 30
$ curl -s 'http://localhost:16686/api/traces?service=unenriched&limit=1' | jq -r '[.data[]?.processes[]?.tags[]? | select(.key|startswith("k8s")) | .key] | unique'
Handling connection for 16686
[]
$ kubectl -n tracing logs deploy/otel-collector --tail=30 | grep -i -m3 forbidden
E0913 14:32:53.207760 1 reflector.go:204] "Failed to watch" err="failed to list *v1.Pod: pods is forbidden: User \"system:serviceaccount:tracing:otel-collector\" cannot list resource \"pods\" in API group \"\" at the cluster scope" logger="UnhandledError" reflector="k8s.io/client-go@v0.35.4/tools/cache/reflector.go:289" type="*v1.Pod"
E0913 14:32:54.618610 1 reflector.go:204] "Failed to watch" err="failed to list *v1.Pod: pods is forbidden: User \"system:serviceaccount:tracing:otel-collector\" cannot list resource \"pods\" in API group \"\" at the cluster scope" logger="UnhandledError" reflector="k8s.io/client-go@v0.35.4/tools/cache/reflector.go:289" type="*v1.Pod"
E0913 14:32:57.318475 1 reflector.go:204] "Failed to watch" err="failed to list *v1.Pod: pods is forbidden: User \"system:serviceaccount:tracing:otel-collector\" cannot list resource \"pods\" in API group \"\" at the cluster scope" logger="UnhandledError" reflector="k8s.io/client-go@v0.35.4/tools/cache/reflector.go:289" type="*v1.Pod"
$ kubectl apply -f - <<'EOF'
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata: { name: otel-k8sattributes }
roleRef: { apiGroup: rbac.authorization.k8s.io, kind: ClusterRole, name: otel-k8sattributes }
subjects:
- { kind: ServiceAccount, name: otel-collector, namespace: tracing }
EOF
clusterrolebinding.rbac.authorization.k8s.io/otel-k8sattributes created
$ kubectl -n tracing rollout restart deploy otel-collector
deployment.apps/otel-collector restarted
$ kubectl -n tracing rollout status deploy otel-collector --timeout=180s
Waiting for deployment spec update to be observed...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
deployment "otel-collector" successfully rolled out
$ kill $PF1k8s.pod.name and friends, the unenriched ones do not, and the collector log has a Forbidden line explaining why. Nothing fails; you just lose the ability to ask "which pod".The spanmetrics connector derives RED metrics from traces you are already collecting, which gives you rate, errors and duration for a service that exports no metrics at all. It is the cheapest observability win on this page.
kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"connectors":{"spanmetrics":{"histogram":{"explicit":{"buckets":["100ms","500ms","1s"]}}}},"exporters":{"prometheus":{"endpoint":"0.0.0.0:8889"}},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","batch"],"exporters":["otlp/jaeger","spanmetrics"]},"metrics/spanmetrics":{"receivers":["spanmetrics"],"exporters":["prometheus"]}}}}}}'
kubectl -n tracing rollout status deploy otel-collector --timeout=180s
kubectl -n tracing run gen-metrics --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service metered --traces 20
sleep 30
kubectl -n tracing port-forward deploy/otel-collector 8889:8889 & PF1=$!
sleep 5
curl -s localhost:8889/metrics | grep -E '^(traces_span_metrics|calls)' | head -10
kill $PF1outputcaptured 2026-09-12
$ kubectl -n tracing patch opentelemetrycollector otel --type merge -p '{"spec":{"config":{"connectors":{"spanmetrics":{"histogram":{"explicit":{"buckets":["100ms","500ms","1s"]}}}},"exporters":{"prometheus":{"endpoint":"0.0.0.0:8889"}},"service":{"pipelines":{"traces":{"receivers":["otlp"],"processors":["k8sattributes","batch"],"exporters":["otlp/jaeger","spanmetrics"]},"metrics/spanmetrics":{"receivers":["spanmetrics"],"exporters":["prometheus"]}}}}}}'
opentelemetrycollector.opentelemetry.io/otel patched
$ kubectl -n tracing rollout status deploy otel-collector --timeout=180s
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment spec update to be observed...
Waiting for deployment spec update to be observed...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "otel-collector" rollout to finish: 1 old replicas are pending termination...
deployment "otel-collector" successfully rolled out
$ kubectl -n tracing run gen-metrics --rm -i --restart=Never --image=ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest -- traces --otlp-insecure --otlp-endpoint otel-collector.tracing.svc:4317 --service metered --traces 20
2026-09-13T03:23:08.558Z INFO traces/traces.go:53 starting gRPC exporter
2026-09-13T03:23:08.558Z INFO grpclog/component.go:69 [core] original dial target is: "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T03:23:08.558Z INFO channelz/trace.go:200 [core] [Channel #1] Channel created for target "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T03:23:08.558Z INFO channelz/trace.go:200 [core] [Channel #1] parsed dial target is: resolver.Target{URL:url.URL{Scheme:"dns", Opaque:"", User:(*url.Userinfo)(nil), Host:"", Path:"/otel-collector.tracing.svc:4317", Fragment:"", RawQuery:"", RawPath:"", RawFragment:"", ForceQuery:false, OmitHost:false}} {"grpc_log": true}
2026-09-13T03:23:08.558Z INFO channelz/trace.go:200 [core] [Channel #1] Channel authority set to "otel-collector.tracing.svc:4317" {"grpc_log": true}
2026-09-13T03:23:08.558Z INFO traces/traces.go:125 generation of traces is limited {"per-second": 1}
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
2026-09-13T03:23:10.559Z INFO channelz/trace.go:200 [core] [Channel #1] Channel Connectivity change to CONNECTING {"grpc_log": true}
2026-09-13T03:23:10.559Z INFO channelz/trace.go:200 [core] [Channel #1] Channel exiting idle mode {"grpc_log": true}
2026-09-13T03:23:10.562Z INFO channelz/trace.go:200 [core] [Channel #1] Resolver state updated: {
"Addresses": [
{
"Addr": "10.96.126.64:4317",
"ServerName": "",
"Attributes": null,
"BalancerAttributes": null,
"Metadata": null
}
],
"Endpoints": [
{
"Addresses": [
{
"Addr": "10.96.126.64:4317",
"ServerName": "",
"Attributes": null,
"BalancerAttributes": null,
"Metadata": null
}
],
"Attributes": null
}
],
"ServiceConfig": null,
"Attributes": null
} (resolver returned new addresses) {"grpc_log": true}
2026-09-13T03:23:10.562Z INFO channelz/trace.go:200 [core] [Channel #1] Channel switches to new LB policy "pick_first" {"grpc_log": true}
2026-09-13T03:23:10.562Z INFO grpclog/prefix_logger.go:42 [pick-first-leaf-lb] [pick-first-leaf-lb 0x6af722e0ab0] Received new config {
"shuffleAddressList": false
... 45 more lines
$ sleep 30
$ kubectl -n tracing port-forward deploy/otel-collector 8889:8889 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:8889 -> 8889
Forwarding from [::1]:8889 -> 8889
$ curl -s localhost:8889/metrics | grep -E '^(traces_span_metrics|calls)' | head -10
Handling connection for 8889
traces_span_metrics_calls_total{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_CLIENT",span_name="lets-go",status_code="STATUS_CODE_UNSET"} 0
traces_span_metrics_calls_total{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_SERVER",span_name="okey-dokey-0",status_code="STATUS_CODE_UNSET"} 0
traces_span_metrics_duration_milliseconds_bucket{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_CLIENT",span_name="lets-go",status_code="STATUS_CODE_UNSET",le="100"} 20
traces_span_metrics_duration_milliseconds_bucket{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_CLIENT",span_name="lets-go",status_code="STATUS_CODE_UNSET",le="500"} 20
traces_span_metrics_duration_milliseconds_bucket{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_CLIENT",span_name="lets-go",status_code="STATUS_CODE_UNSET",le="1000"} 20
traces_span_metrics_duration_milliseconds_bucket{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_CLIENT",span_name="lets-go",status_code="STATUS_CODE_UNSET",le="+Inf"} 20
traces_span_metrics_duration_milliseconds_sum{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_CLIENT",span_name="lets-go",status_code="STATUS_CODE_UNSET"} 2.460000000000001
traces_span_metrics_duration_milliseconds_count{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_CLIENT",span_name="lets-go",status_code="STATUS_CODE_UNSET"} 20
traces_span_metrics_duration_milliseconds_bucket{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_SERVER",span_name="okey-dokey-0",status_code="STATUS_CODE_UNSET",le="100"} 20
traces_span_metrics_duration_milliseconds_bucket{collector_instance_id="e12e23aa-5f16-4d41-a030-ad9950fa106a",job="metered",otel_scope_name="spanmetricsconnector",otel_scope_schema_url="",otel_scope_version="",service_name="metered",span_kind="SPAN_KIND_SERVER",span_name="okey-dokey-0",status_code="STATUS_CODE_UNSET",le="500"} 20
$ kill $PF1Jaeger v2 is an OpenTelemetry Collector with storage and query extensions, which means its configuration reads like a collector's. Knowing that changes how you debug it: the pipeline is right there.
kubectl -n tracing port-forward svc/jaeger 16686:16686 & PF1=$!
sleep 5
curl -s http://localhost:16686/api/services | jq
curl -s 'http://localhost:16686/api/operations?service=err-service' | jq '.data[:5]'
kill $PF1
kubectl -n tracing get cm -o name
kubectl -n tracing get deploy jaeger -o jsonpath='{.spec.template.spec.containers[0].args}{"\n"}'
kubectl -n tracing get deploy jaeger -o jsonpath='{.spec.template.spec.containers[0].env}' | jqoutputcaptured 2026-09-12
$ kubectl -n tracing port-forward svc/jaeger 16686:16686 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:16686 -> 16686
Forwarding from [::1]:16686 -> 16686
$ curl -s http://localhost:16686/api/services | jq
Handling connection for 16686
{
"data": [
"err-service",
"jaeger",
"metered"
],
"total": 3,
"limit": 0,
"offset": 0,
"errors": null
}
$ curl -s 'http://localhost:16686/api/operations?service=err-service' | jq '.data[:5]'
Handling connection for 16686
[
{
"name": "okey-dokey-0",
"spanKind": "server"
},
{
"name": "lets-go",
"spanKind": "client"
}
]
$ kill $PF1
$ kubectl -n tracing get cm -o name
configmap/kube-root-ca.crt
configmap/otel-collector-3a402bbd
configmap/otel-collector-a60f73e1
configmap/otel-collector-fa2c2caf
$ kubectl -n tracing get deploy jaeger -o jsonpath='{.spec.template.spec.containers[0].args}{"\n"}'
$ kubectl -n tracing get deploy jaeger -o jsonpath='{.spec.template.spec.containers[0].env}' | jq
[
{
"name": "COLLECTOR_OTLP_ENABLED",
"value": "true"
}
]Self-check
Jaeger shows hundreds of one-span traces. Diagnosis?
Context propagation is broken: each service starts a new trace instead of continuing the caller's. Look for a proxy stripping traceparent, a missing or mismatched propagator (W3C vs B3), a hand-rolled HTTP client that never injects, or an async boundary nobody instrumented.
Head-based versus tail-based sampling: when does the difference bite?
When the interesting traces are rare. Head sampling at 1% will usually miss the one 5-second request that mattered; tail sampling keeps traces that contain errors or exceed a latency threshold, at the cost of buffering complete traces in the collector (and needing all spans of a trace to reach the same collector instance).
Where should sampling and redaction be configured, and why there?
In the collector. It is the platform's control point: one config change applies to every service, whereas SDK-side changes need forty teams to redeploy. This is the same "interface owned by the platform" argument as section 3.1, applied to telemetry.
You annotated a Deployment for auto-instrumentation and nothing changed. Two most likely reasons?
The annotation is on the Deployment's own metadata rather than spec.template.metadata, or the pods have not restarted since (injection happens at admission, at pod creation). Third candidate: the annotation value's <namespace>/<name> does not match an existing Instrumentation CR.
How do traces, metrics and logs tie together in practice?
Exemplars link metric buckets to sample trace IDs; logs carry the trace ID as a field so you can jump from a slow span to its exact log lines; and the collector can emit all three from the same pipeline with the same resource attributes (namespace, pod, service). "Correlated by trace ID and resource attributes" is the sentence.
Your tail-sampling collector runs three replicas behind a Service, and the sampling decisions look random. Why, and what fixes it?
Spans of one trace are spread across replicas, so no single instance sees the whole trace when decision_wait expires. Put a first-tier collector with a loadbalancing exporter using routing_key: traceID in front (or run the sampling tier as one replica, or a StatefulSet the load balancer can address per pod) so every span of a trace lands on the same instance.
What is Jaeger v2 architecturally, and where does the old jaeger-agent go?
A distribution of the OpenTelemetry Collector: one binary, collector-style YAML, roles collector/query/ingester/all-in-one, with Jaeger-specific extensions for storage, the query UI on 16686 and remote sampling. Ingest is plain OTLP on 4317/4318. The agent role is deprecated; run a standard OpenTelemetry Collector as DaemonSet or sidecar where the agent used to be.
A pod has the injected init container and the right annotation, yet no spans appear. Two collector-side reasons and one network reason?
The Instrumentation's exporter.endpoint uses the wrong port for the protocol the agent speaks (Python is HTTP-only, so it must be 4318; Java defaults to gRPC 4317), or the collector pipeline has no traces pipeline wired to that receiver. Network: a default-deny NetworkPolicy in the app namespace blocks egress to the collector's namespace. Read OTEL_EXPORTER_OTLP_ENDPOINT in the pod and test the port from inside it.
How does a metric panel jump to a trace, and what has to be enabled at each layer?
Exemplars: the histogram sample carries a trace_id. The instrumented service or the spanmetrics connector emits them, Prometheus keeps them only with --enable-feature=exemplar-storage, and the Grafana datasource maps trace_id to a Tempo or Jaeger datasource via exemplarTraceIdDestinations. Logs join the same picture when the log line carries the trace ID and Loki's derivedFields links it.
Docs to know your way around
- opentelemetry.io: the Concepts pages (signals, context propagation), the collector's receivers/processors/exporters reference, and the Kubernetes operator's Instrumentation injection docs.
- jaegertracing.io: mostly the UI is self-explanatory; know that the HTTP API exists for scripted verification.
- w3.org/TR/trace-context: one page, and it is the header everything agrees on.
- opentelemetry.io/docs/collector/configuration: sections, connectors and pipelines; the processor READMEs in github.com/open-telemetry/opentelemetry-collector-contrib (tailsamplingprocessor, k8sattributesprocessor) hold the field defaults.
- opentelemetry.io/docs/platforms/kubernetes/operator/automatic: annotation values,
container-names, the Go and Python caveats; the troubleshooting page beside it is the injection diagnostic ladder. - jaegertracing.io/docs/latest/architecture: roles, storage backends and the "we recommend the OpenTelemetry Collector instead of the agent" statement.