Verifying Correlation ID Propagation
Goal
Create a W3C traceparent by hand and propagate it through a 3-tier service chain, then join the logs by trace ID to reconstruct the path of a single request.
Why it matters
If you explain distributed tracing as "it is needed because it is the third signal", adoption fails. You must know exactly which question it answers. A situation where the p99 of ten services is all normal but the user's screen takes 2 seconds — here metrics are powerless in principle. This is because they are aggregates and cannot reconstruct the path of an individual request. The N+1 problem, where a 4ms query runs 340 times in one request, also never shows up in metrics. This is because each individual query is fast. By reconstructing that path with only three pieces of a header and no SDK, this lab makes you attach OpenTelemetry later knowing what actually happens.
Steps
- With
/opt/app/chain.py, start three services: gateway (8120), orders (8121), and inventory (8122).GET http://127.0.0.1:8120/orderreturns 200. - With
/root/trace/gen.py, create a traceparent and save it on one line to/root/trace/tp.txt. It must be in the format00-<32hex>-<16hex>-01. - Call the gateway with that header. Each of the three services writes, in
/root/trace/svc-<이름>.log(the service name),trace_id=<값>(the value), and all three values must be identical. - Also leave
span_id=<16hex>andparent_span_id=<16hex>in each log. Thespan_idvalues of the three services must all differ. - With
/root/trace/join.sh, extract only the lines with the same trace_id from the three logs and leave 3 lines in time order in/root/trace/joined.log. - If you run
/opt/app/brokenchain.sh, a second chain comes up on 8123/8124/8125 and one of the services drops the header. Each service writes, in/root/trace/broken-<이름>.log(the service name),svc=<이름> in_trace=<값> out_trace=<값>(service name, incoming trace ID, outgoing trace ID). Find the service whose incoming trace ID differs from its outgoing trace ID and write its name on one line in/root/trace/broken.txt. - In
/root/trace/selftime.txt, write two lines:slowest=<서비스명> total_ms=<정수>andself_ms=<정수>(the service name and integers). self is the total minus the child intervals.
Notes
- traceparent format:
00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01 - The trace ID is the same throughout the request, and the span ID changes at every hop.
- The four places where propagation breaks: your own HTTP client, handing off to a thread pool, message queues, and proxies that strip headers.
- Common mistake: making the trace ID uppercase hexadecimal — the specification says lowercase.
Start a 3-tier chain of services
With /opt/app/chain.py, start three services: gateway (8120), orders (8121), and inventory (8122). GET http://127.0.0.1:8120/order returns 200.
/opt/app/chain.py comes up on the port it receives as an argument and takes the next hop's address from an environment variable. Start three, in the order gateway, orders, inventory.
Create the traceparent header
With /root/trace/gen.py, create a traceparent and save it on one line to /root/trace/tp.txt. It must be in the format 00-<32hex>-<16hex>-01.
Join a 2-digit version, a 32-digit trace ID, a 16-digit span ID, and a 2-digit flag with hyphens. All must be lowercase hexadecimal.
Pass the same trace ID to the very end
Call the gateway with that header. Each of the three services writes, in /root/trace/svc-<이름>.log (the service name), trace_id=<값> (the value), and all three values must be identical.
Have each service pass on only the trace ID from the header it received as is, and record it in its own log.
Change the span ID at every hop
Also leave span_id=<16hex> and parent_span_id=<16hex> in each log. The span_id values of the three services must all differ.
In the parent span ID slot, send your own span ID. It is normal for the span IDs in the three logs to be all different.
Join the three logs by trace ID
With /root/trace/join.sh, extract only the lines with the same trace_id from the three logs and leave 3 lines in time order in /root/trace/joined.log.
Filter each service's log file by trace ID and concatenate them in time order. The join result must be exactly 3 lines.
Find the service where propagation broke
If you run /opt/app/brokenchain.sh, a second chain comes up on 8123/8124/8125 and one of the services drops the header. Each service writes, in /root/trace/broken-<이름>.log (the service name), svc=<이름> in_trace=<값> out_trace=<값> (service name, incoming trace ID, outgoing trace ID). Find the service whose incoming trace ID differs from its outgoing trace ID and write its name on one line in /root/trace/broken.txt.
There is one service that deliberately drops the header. Just find the point where the trace ID changes.
Calculate the self time
In /root/trace/selftime.txt, write two lines: slowest=<서비스명> total_ms=<정수> and self_ms=<정수> (the service name and integers). self is the total minus the child intervals.
Subtract the time of the child spans from the total duration. The slowest interval and the self time can be in different places.