TT Lab
Get started
Learn Learning paths Courses

Capital Markets and Settlement

Which hop is slow, and what is that negative number

Continue in TT Lab

Goal

From the segment timestamps of 1,500 orders, you compute per-segment latency, correct with the host clock offsets to separate clock error from real defects, and then count separately the per-segment percentiles and the segments that make up the overall tail, narrowing down even where to fix.

Why it matters

A report that "it is slow" has no place to fix. You have to put a ruler along the path and split it into segments to decide which team will fix it. After splitting, two traps are waiting. The first is negatives. In a segment whose two ends are on different machines, the clock error of the two machines is mixed in as it is, so if the clock difference is larger than the true line time, the value becomes negative. If you dismiss this as a broken measurement, you lose the basis for talking with the counterparty later, and conversely if you lump it all together as the clock's fault, you miss the real one mixed in there. The second is percentiles. The segment with the largest per-segment p99 and the segment that makes up the tail of the overall latency are not the same. The former looks at different orders in each segment, and the latter looks at what was large within a single order. If you mix the two, you spend money in the wrong place.

Steps

  1. Use python3 to create spans.jsonl, offsets.json, spec.json, and budget.json in /root/lat/data.
  2. Count the shape of the data and write it to /root/lat/shape.json.
  3. Compute the uncorrected segment latencies and write them to /root/lat/raw_stats.json and /root/lat/raw_negative.csv.
  4. Correct with the offsets and separate the causes in /root/lat/clock_split.csv, /root/lat/defects.csv, and /root/lat/adjust_summary.json.
  5. Write the per-segment percentiles to /root/lat/percentile.csv.
  6. Pull out the tail of overall latency and write it to /root/lat/tail.csv and /root/lat/tail.json.
  7. Write the exceedance counts of the two budgets to /root/lat/budget_over.csv and /root/lat/budget_summary.json.
  8. Write the per-minute trend to /root/lat/trend.csv and the narrowed-down conclusion to /root/lat/report.json.

Notes

Generate the segment timestamps and the offset reports

Use python3 to create spans.jsonl, offsets.json, spec.json, and budget.json in /root/lat/data. Use the generation script as is, which uses no random numbers.

We cannot bring in the customer's data as it is, so we build synthetic data of the same shape. Without random numbers, the same data comes out no matter who runs it how many times, and you can compare each other's judgments. Writing the percentile definition and the budget next to the data is the key habit of this lab.

Count the shape of the data

Write orders, segments, hosts, complete_orders, incomplete_orders, missing_by_field, and segment_n to /root/lat/shape.json.

The number of orders you can count differs per segment. That is because if one timestamp is missing, the two segments that use it cannot be computed. segment_n is, for each segment, the number of orders that have both end timestamps, and complete_orders is the number of orders that have all six timestamps. Do not mix the two.

Uncorrected segment latency and finding the negatives

Write n, negative, min_us, and max_us for each segment to /root/lat/raw_stats.json, and in /root/lat/raw_negative.csv put the first line segment,host,negative_orders,min_us,max_us and write only the segment and host groups where negatives appeared.

Segment latency is the later timestamp minus the earlier one, and here you do not correct yet. There are two reasons negatives appear. The two ends are on different machines and so the clock difference is mixed in, or the timestamps were stamped backward within the same machine. Do not separate them yet; just count where they are concentrated. Do not write groups with no negatives.

Correct with the offsets and separate clock error from real defects

Put the first line segment,host,raw_negative,adjusted_negative,cause in /root/lat/clock_split.csv and the first line order_id,host,segment,raw_us,adjusted_us in /root/lat/defects.csv, and write raw_negative, explained_by_clock, real_defects, roundtrip_offset_free_orders, and eligible_orders to /root/lat/adjust_summary.json.

The corrected time is the recorded time minus the offset of the machine that stamped it. If no negatives remain after correction, the cause is 시계오차; if all remain, 진짜결함; if only some remain, 섞임 (the three Korean words mean clock error, real defect, and mixed). roundtrip_offset_free_orders is the number of orders for which the value from line send to ack receive is the same before and after correction. Since the same machine stamped both ends, first think about how many should come out.

Compute per-segment percentiles

Put the first line segment,n,p50_us,p95_us,p99_us,max_us,mean_us in /root/lat/percentile.csv and write five segment lines and one total line. The targets are the orders decided by eligible in spec.json.

The percentile definition is in spec.json. After sorting, take the position as p*(n-1)/100 counted from 0, and if the position is not an integer, interpolate between the two neighboring values and round half up. It can all be done with integer arithmetic. total is from receive to ack receive, and since both ends of this value are the same machine, no correction is needed. Round the mean down with integer division.

Count the segments that make up the overall tail

Put the first line segment,tail_orders in /root/lat/tail.csv and write all five segment lines, and write threshold_us, tail_orders, top_p99_segment, top_tail_segment, and same to /root/lat/tail.json.

Tail orders are orders whose corrected total latency is at or above the p99 of total. For each order pick and count its largest segment, and if values are equal, pick the one earlier in the segment order the data defines. top_p99_segment is the segment with the largest per-segment p99 from step 5, and top_tail_segment is the segment picked most often here. If the two differ, same is false. Write 0 for segments that were picked by no order at all.

Compare the exceedance counts of the two budgets

Put the first line budget,segment,over_orders in /root/lat/budget_over.csv and write five segment lines per budget, and write eligible_orders and, for each budget, over_orders, over_ppm, and worst_segment to /root/lat/budget_summary.json.

If a segment value exceeds the upper bound, it is an exceedance. If equal, it is not. over_orders is the number of orders that exceeded in even one segment, not the sum of per-segment exceedance counts. If one order exceeds two segments, count it only once. over_ppm is parts per million, so compute it with integer division. worst_segment is the segment with the most per-segment exceedances, and if equal, the one earlier in the order the data defines.

Write since when it got worse and the hypotheses ruled out

Put the first line minute,segment,n,p99_us in /root/lat/trend.csv and write the per-segment p99 for each minute bucket, and write worst_segment, degraded_minutes, onset_minute, end_minute, supported, and ruled_out to /root/lat/report.json.

The minute bucket is the quotient of the corrected receive time divided by 60 seconds. If some minute's segment p99 is at least 10 times that segment's overall p50, that minute is a degraded minute, and the segment with the most degraded minutes is the worst segment. The judgment rules for the four hypotheses are written in hypotheses of spec.json. Compute by the rules to separate the supported from the ruled out, and write the identifiers of both lists in ascending order.