
As modern homelabs, developer environments, and enterprise edge clusters shift from monolithic setups toward microservices, understanding why a multi-tier request fails or stalls becomes extraordinarily difficult. Traditional log aggregation answers what an isolated service reported at a specific timestamp, but it cannot illustrate where an end-to-end request spent 80% of its lifecycle across gateway routers, auth middlewares, internal APIs, and database engines. This is the exact blind spot that distributed tracing eliminates.
The industry-standard observability architecture pairs the vendor-neutral OpenTelemetry (OTel) Collector with Jaeger. The OpenTelemetry Collector acts as a high-throughput proxy pipeline that ingests traces, metrics, and logs from client applications via the OTLP (OpenTelemetry Protocol), batches and sanitizes the telemetry, and exports it to backends. Jaeger serves as the distributed tracing storage, search engine, and visualization UI, rendering spans, parent-child waterfalls, and service dependency graphs. In this production-ready guide, we deploy a resilient OTel Collector and Jaeger stack using Docker Compose, complete with end-to-end tracing test verification, memory limiter guards, and reverse proxy security.
Architecture & Telemetry Data Flow
Rather than pointing your applications directly to a specific visualization backend, best practice dictates decoupling ingestion from storage through the OpenTelemetry Collector. Your applications export traces over OTLP (gRPC on port 4317 or HTTP/JSON on port 4318) directly to the Collector. The Collector processes the streams inside internal pipelines (batching, memory limiting, attribute filtering) and exports spans to Jaeger over native OTLP gRPC.
+---------------------------------------------------------------+
| Application Services |
| +-------------------+ +--------------------+ |
| | FastAPI Frontend | | Go Worker Backend | |
| +---------+---------+ +----------+---------+ |
| | | |
| | OTLP/HTTP (4318) | OTLP/gRPC |
| +-----------------+-----------------+ (4317) |
+-------------------------------|-------------------------------+
v
+---------------------------------------------------------------+
| OpenTelemetry Collector (Container) |
| Receivers: otlp (4317 gRPC / 4318 HTTP) |
| Processors: memory_limiter -> batch -> resourcedetection |
| Exporters: otlp (to Jaeger) + debug (stdout logging) |
+-------------------------------+-------------------------------+
|
| OTLP/gRPC (4317)
v
+---------------------------------------------------------------+
| Jaeger Tracing (all-in-one) |
| - Ingests traces from Collector via OTLP |
| - In-Memory or Badgermv store for spans & service graphs |
| - Web UI on port 16686 (Query, Dependencies, Waterfalls) |
+---------------------------------------------------------------+
Prerequisites & Host Directory Setup
To follow this guide, you need a Linux host (Ubuntu 24.04/22.04 LTS or Debian 12) with Docker Engine (v26.0+) and the Docker Compose plugin (v2.27+) installed. Create a clean dedicated directory structure for the observability stack:
sudo mkdir -p /opt/observability/{otel,jaeger-data}
cd /opt/observability
sudo chown -R 10001:10001 /opt/observability/otel
sudo chmod -R 755 /opt/observability
Configuring the OpenTelemetry Collector Pipeline
The OTel Collector configuration defines three core concepts: Receivers (how data enters), Processors (how data is transformed or throttled), and Exporters (where data is dispatched). These components are assembled into distinct Pipelines inside the service block.
Create the collector configuration file at /opt/observability/otel/otel-collector-config.yaml:
# /opt/observability/otel/otel-collector-config.yaml
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
# Safeguards collector container against out-of-memory crashes
memory_limiter:
check_interval: 1s
limit_percentage: 75
spike_limit_percentage: 20
# Batches spans into groups for network efficiency and backend throughput
batch:
send_batch_size: 1024
timeout: 1s
send_batch_max_size: 2048
# Enriches incoming telemetry with host and OS metadata
resourcedetection:
detectors: [env, system]
timeout: 2s
exporters:
# Exports traces to Jaeger over OTLP gRPC
otlp/jaeger:
endpoint: jaeger:4317
tls:
insecure: true
# Debug logger for local verification in container logs
debug:
verbosity: basic
extensions:
health_check:
endpoint: 0.0.0.0:13133
zpages:
endpoint: 0.0.0.0:55679
service:
extensions: [health_check, zpages]
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch, resourcedetection]
exporters: [otlp/jaeger, debug]
telemetry:
logs:
level: "info"
Production Docker Compose Specification
Next, write the docker-compose.yml file in /opt/observability/docker-compose.yml. We configure the otel/opentelemetry-collector-contrib image (the community distribution that includes extended processors like resourcedetection) alongside the official jaegertracing/all-in-one image.
# /opt/observability/docker-compose.yml
services:
otel-collector:
image: otel/opentelemetry-collector-contrib:0.109.0
container_name: otel-collector
restart: unless-stopped
command: ["--config=/etc/otelcol-contrib/config.yaml"]
volumes:
- ./otel/otel-collector-config.yaml:/etc/otelcol-contrib/config.yaml:ro
ports:
- "4317:4317" # OTLP gRPC receiver
- "4318:4318" # OTLP HTTP receiver
- "13133:13133" # OTel Collector Health Check
- "55679:55679" # zPages diagnostic interface
environment:
- OTEL_RESOURCE_ATTRIBUTES=deployment.environment=production,service.namespace=homelab
networks:
- tracing-net
depends_on:
jaeger:
condition: service_healthy
deploy:
resources:
limits:
memory: 512M
cpus: "1.0"
jaeger:
image: jaegertracing/all-in-one:1.60.0
container_name: jaeger
restart: unless-stopped
environment:
- COLLECTOR_OTLP_ENABLED=true
- METRICS_STORAGE_TYPE=prometheus
- MEMORY_MAX_TRACES=100000
volumes:
- ./jaeger-data:/tmp
ports:
- "16686:16686" # Jaeger Web UI
- "14268:14268" # Legacy HTTP collector (optional)
networks:
- tracing-net
healthcheck:
test: ["CMD-SHELL", "wget -q --spider http://localhost:16686/ || exit 1"]
interval: 10s
timeout: 5s
retries: 5
start_period: 15s
deploy:
resources:
limits:
memory: 1024M
cpus: "1.5"
networks:
tracing-net:
name: tracing-net
driver: bridge
Launching the Stack & Health Verification
Launch both services in detached mode and observe container startup logs to ensure both the OTel receiver and Jaeger storage backend initialize without port or syntax collisions:
cd /opt/observability
docker compose up -d
# Verify container statuses and health states
docker compose ps
# Check Collector pipeline initialization
docker compose logs otel-collector | grep -E "Everything is ready|Pipeline is started"
You can also query the built-in OTel health check endpoint over HTTP directly from the terminal:
curl -I http://127.0.0.1:13133/
# Expected HTTP response:
# HTTP/1.1 200 OK
# Content-Type: text/plain; charset=utf-8
# Content-Length: 0
Sending Test Traces via Python OTel SDK
To verify that the entire telemetry pipeline processes spans from application space all the way to Jaeger’s visual index, run a standalone synthetic Python test script that emits a multi-span trace over OTLP/HTTP:
# Install the minimal OpenTelemetry SDK packages inside a virtualenv
python3 -m venv /tmp/otel-test
/tmp/otel-test/bin/pip install opentelemetry-api opentelemetry-sdk opentelemetry-exporter-otlp-proto-http
Create a script named send_trace.py:
# /tmp/send_trace.py
import time
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.sdk.resources import Resource
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
# Define the service identity and environment
resource = Resource.create({
"service.name": "order-gateway",
"service.version": "1.4.2",
"deployment.environment": "staging"
})
provider = TracerProvider(resource=resource)
# Direct exporter pointing to our local OTel Collector on port 4318
exporter = OTLPSpanExporter(endpoint="http://127.0.0.1:4318/v1/traces")
processor = BatchSpanProcessor(exporter)
provider.add_span_processor(processor)
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("ecommerce.checkout")
# Create a parent span with nested child spans to simulate a real transaction
with tracer.start_as_current_span("POST /api/v1/checkout") as parent:
parent.set_attribute("http.status_code", 200)
parent.set_attribute("customer.tier", "enterprise")
time.sleep(0.08)
with tracer.start_as_current_span("auth.verify_token") as auth_span:
auth_span.set_attribute("auth.method", "jwt_ed25519")
time.sleep(0.03)
with tracer.start_as_current_span("inventory.reserve_stock") as inv_span:
inv_span.set_attribute("item.sku", "SRE-KEY-009")
inv_span.set_attribute("inventory.warehouse", "eu-central-1")
time.sleep(0.12)
with tracer.start_as_current_span("stripe.charge_card") as payment_span:
payment_span.set_attribute("payment.gateway", "stripe")
payment_span.set_attribute("payment.currency", "EUR")
time.sleep(0.19)
provider.shutdown()
print("Successfully generated and flushed distributed trace!")
Execute the script:
/tmp/otel-test/bin/python /tmp/send_trace.py
Now open your browser and navigate to http://<your-server-ip>:16686. In the Service dropdown on the left navigation bar, select order-gateway and click Find Traces. You will immediately see the complete transaction waterfall showing the parent POST /api/v1/checkout span and its three children (auth.verify_token, inventory.reserve_stock, and stripe.charge_card) with microsecond timing accuracy.
Production Hardening & Best Practices
- Keep the Collector Memory Limiter First: In your
otel-collector-config.yamlpipelines, thememory_limiterprocessor must always precedebatchor attribute processors. If incoming traffic spikes unexpectedly, the memory limiter drops or backpressures spans before the container triggers an unrecoverable kernel OOM kill. - Isolate OTLP Ports Behind an Internal Network: Applications residing on the same Docker host or Kubernetes node should communicate with the Collector over internal overlay/bridge networks (e.g.,
otel-collector:4317). Only expose ports4317and4318to the host if external nodes or mobile clients push telemetry directly. - Protect Jaeger UI with Reverse Proxy Auth: Jaeger’s built-in web UI does not include native user authentication. Do not expose port
16686directly to the public internet. Place it behind Caddy, Traefik, or Cloudflare Access with HTTP Basic Authentication, Authelia, or Authentik SSO. - Tune In-Memory vs. Persistent Badger Storage: The default
MEMORY_MAX_TRACES=100000environment variable keeps traces in container RAM. For persistent traces across container restarts without deploying an entire Elasticsearch or Cassandra cluster, configure Badger storage by settingSPAN_STORAGE_TYPE=badgerand binding a host directory to/badger.
Troubleshooting Common Failures
1. Error: “connection refused” or “rpc error: code = Unavailable” on Port 4317
Root Cause: Client applications mistakenly attempt to connect to Jaeger on port 4317 instead of the OTel Collector, or the Collector has not bound to 0.0.0.0 inside its container definition.
Solution: Verify that the application’s OTLP exporter points to the Collector service name (http://otel-collector:4317 in Docker or 127.0.0.1:4317 on host). Check that the Collector config explicitly uses endpoint: 0.0.0.0:4317 rather than localhost:4317, which binds exclusively to the loopback interface inside the container.
2. Traces Appear in Collector Logs but Never Show Up in Jaeger UI
Root Cause: Jaeger does not have the OTLP receiver enabled by default in older images, or the Collector exporter is attempting to use TLS against Jaeger’s insecure internal port.
Solution: Ensure your docker-compose.yml sets COLLECTOR_OTLP_ENABLED=true on the Jaeger service. In otel-collector-config.yaml, verify that the otlp/jaeger exporter configuration includes tls: insecure: true.
3. High CPU or Memory Usage Under Heavy Span Load
Root Cause: The batch processor is either absent or configured with an excessively short timeout, causing the Collector to issue synchronous gRPC network calls for every single span emitted by the application.
Solution: Tune the batch processor parameters. Set send_batch_size: 1024 and timeout: 1s. Ensure container resource limits in docker-compose.yml allocate at least 512MB RAM and 1 CPU core to the Collector process.
Conclusion
By pairing the OpenTelemetry Collector with Jaeger in Docker Compose, you establish an extensible, vendor-agnostic observability tier for your infrastructure. Applications remain completely decoupled from the visualization backend: should you later transition to Grafana Tempo, SigNoz, or Datadog, your application code remains untouched—you simply update the exporter block in your otel-collector-config.yaml. With memory safety limiters, automated span batching, and OTLP validation in place, you now possess deep visibility into latency bottlenecks across your entire distributed stack.
Hi, I’m Mark, the author of Clever IT Solutions: Mastering Technology for Success. I am passionate about empowering individuals to navigate the ever-changing world of information technology. With years of experience in the industry, I have honed my skills and knowledge to share with you. At Clever IT Solutions, we are dedicated to teaching you how to tackle any IT challenge, helping you stay ahead in today’s digital world. From troubleshooting common issues to mastering complex technologies, I am here to guide you every step of the way. Join me on this journey as we unlock the secrets to IT success.


