One Order Vanished, and There Were Five Services
Goal
From the logs of five services, you link the path of one order from start to end. You split the W3C traceparent into four fields and quarantine invalid values, line up the lines of one trace-id in time order, build the span tree with parent-id, work out the duration of each span, join the place where the header broke with the business key, and then read the sampled flag as a mask.
Why it matters
The customer says "an order disappeared," but there are five services. If you guess by time, among dozens of requests per second there is no basis to choose which one is that order. A correlation id fills that gap, but the problem you actually run into in the field is not a missing id but its breaking at some point in the middle. The investigation includes finding where it broke and proving it by joining the before and after again with the business key. That proof becomes the basis for making the next deployment pass the header through.
Steps
- Create and run
/root/trace/gen_trace.pyto create the logs of five services under/root/trace/raw/. /root/trace/parsed.ndjsonand/root/trace/badtp.ndjson— split every traceparent into four fields, and quarantine values that do not conform to the specification, with reasons./root/trace/one_trace.ndjson— find the trace-id of order ORD-2026-4117 and line up all the lines of that trace in time order./root/trace/tree.ndjson— build the span tree by setting up parent-child relationships with parent-id./root/trace/elapsed.json— work out the duration and own time of each span./root/trace/bridge.ndjsonand/root/trace/gap.json— find the service where the header broke and join its before and after with the order number./root/trace/sampled.json— read the sampled flag as a mask to separate traces that were recorded from those that were not./root/trace/trace_report.md— leave the investigation results as a report.
Notes
- A traceparent is a fixed-length value with four fields,
버전-trace-id-parent-id-trace-flags(the first field is the version), joined by hyphens. The lengths are 2 · 32 · 16 · 2 characters in order, and only lowercase hex characters are allowed. - Do the invalid-value checks in this order: number of fields and length (bad_shape) → lowercase hex (non_lowercase_hex) → whether the version is the current version 00 (bad_version — this lab accepts only 00. The specification pins ff as invalid) → whether the trace-id is all zeros (zero_trace_id) → whether the parent-id is all zeros (zero_parent_id).
- The parent-id of the header I received is the span id of the party that called me, and the parent-id of the header I send is my own span id. The traceparent_in and traceparent_out in the log are those two, respectively.
- trace-flags is an 8-bit field. Do not compare whether it equals
01; mask the least significant bit — both09and03are sampled. - Collect all outputs of this lab under
/root/trace/. They disappear when the session ends, so keep the important ones on screen. - Common mistakes: quietly skipping invalid traceparents, looking only at durations without working out own time, passing over a service with no header as "there is no log," and counting sampled by string comparison.
Get the logs of five services in hand
Create and run /root/trace/gen_trace.py to create edge.jsonl (58 lines), orders.jsonl (72 lines), payments.jsonl (48 lines), stock.jsonl (48 lines), and ledger.jsonl (48 lines) under /root/trace/raw/.
All five files have one JSON object per line (JSON Lines). Most lines carry traceparent_in and traceparent_out, but the lines of one service carry none at all — that service is the protagonist of this lab. The edge also has mixed in health check lines and invalid headers stamped by an old mobile gateway.
Split the traceparent into four fields
In /root/trace/parsed.ndjson, write, one per line, svc, line_no (from 1), field (traceparent_in|traceparent_out), version, trace_id, parent_id, and trace_flags for each traceparent that conforms to the specification. Leave values that do not conform in /root/trace/badtp.ndjson as svc, line_no, field, raw (the original text as it is), and reason.
The material is all five files under /root/trace/raw/ created in step 1. A traceparent is fixed length — the four hyphen-separated fields have lengths 2 · 32 · 16 · 2 in order, and only lowercase hex characters are allowed. Attach the reason in the order of checks set in the instructions. When reading the files, count line numbers from 1 in the original file, and lines with no traceparent at all go into neither.
Line up that order's trace in time order
The order the customer mentioned is ORD-2026-4117. Find that order's traceparent in the edge log to get the trace-id, and in /root/trace/one_trace.ndjson write all the lines that have that trace-id as ts, svc, trace_id, span_id, parent_span_id, and msg in ascending ts order. span_id is the parent-id of that line's traceparent_out, and parent_span_id is the parent-id of traceparent_in (an empty string if absent).
The material is the five files under /root/trace/raw/. The trace-id points to the whole trace and the parent-id points to a single request. So the criterion for gathering is the trace-id. This is an order in which five services were involved, so count how many services appear here — that number is the problem of this lab.
Build the span tree with parent-id
In /root/trace/tree.ndjson, write, one per line, span_id, parent_span_id, svc, depth (the root is 0), and child_count for each span that appeared in the trace from step 3. Sort by depth ascending, and by span_id ascending for ties. There must be exactly one root.
The material is the /root/trace/one_trace.ndjson you made in step 3. One span leaves several lines, so first fold by span_id. depth is obtained by counting up along parent_span_id. child_count is the number of spans that point to me as their parent — if this number is 0 yet you know that service calls another service, that spot is where it broke.
Measure by span where the time went
In /root/trace/elapsed.json, write spans (for each span, span_id, svc, ms, and self_ms, sorted by span_id ascending), slowest_self_svc, and slowest_self_ms. ms is the time difference (integer milliseconds) between that span's first line and last line, and self_ms is that minus the sum of the children's ms.
The material is the /root/trace/one_trace.ndjson you made in step 3. If you look only at durations, the top span is always the largest — naturally, since it holds the children. What you want to know is the time each service spent itself, so you must subtract the children's time. The times are already fixed in the same notation, so convert to milliseconds and subtract. If a span with an unusually large own time shows up, look at step 6 before concluding it is the culprit.
Find where it broke and join with the order number
In /root/trace/bridge.ndjson, write every line of every service that mentions order ORD-2026-4117, as ts, svc, trace_id (an empty string if none), order_id, and linked_by (traceparent|order_id), in ascending ts order. Also in /root/trace/gap.json, write dropped_at (the service that left not a single line with a traceparent), restarted_at (the service that, after it, created a root span with a new trace-id), and trace_ids (the trace-ids involved in this order, in the order they first appeared).
The material is again the five files under /root/trace/raw/. Only three services appeared in the trace from step 3, but this order passed through five services. The key to finding the other two is not the trace context but the business key. dropped_at is the service in whose log no line has a traceparent, and restarted_at is the service that had no received header and so created a new trace-id on its own.
Read the sampled flag as a mask
In /root/trace/sampled.json, write total_traces, sampled_traces, unsampled_traces, flag_values (the number of traces per trace-flags value), and naive_equal_01 (the number of traces whose trace-flags is the string 01). What you count is all distinct trace-ids that appear in the /root/trace/parsed.ndjson made in step 2, and the trace-flags within one trace are all the same.
trace-flags is an 8-bit field. The specification pins down that you should mask, not interpret the hex as a number and compare with that value. In Python it is int(flags, 16) & 1. naive_equal_01 is a number deliberately counted by the wrong method, so putting the two values side by side is the purpose of this step.
Leave the investigation results as a report
In /root/trace/trace_report.md, write four sections: ## 고객은 무엇을 물었나 ## 추적이 어디서 끊겼나 ## 시간은 어디서 갔나 ## 다음 배포에서 고칠 것 (in order, these mean: what the customer asked, where the trace broke, where the time went, and what to fix in the next deployment). It must include as is the order number, the name of the service where the header broke, the largest self_ms, the number of invalid traceparents, and the number of sampled traces.
The sentences you give back to the customer are two: 'the order did not disappear, and it got this far' and 'why it was not visible in the tool.' Do not make up numbers; take them from the outputs of the earlier steps (/root/trace/gap.json, /root/trace/elapsed.json, /root/trace/sampled.json, /root/trace/badtp.ndjson, /root/trace/bridge.ndjson). The last section writes what to change in the next deployment so that this investigation is no longer needed.