Capital Markets and Settlement
A report that says it is slow gives you nothing to fix
In one line
Timestamps must be stamped at every place an order passes through before you can say which segment is slow. And if the two ends of a segment are on different machines, the clock error of the two machines gets mixed into that difference as it is.
Why this was needed
When you get a report that "orders are slow," at first there is nothing to do. That is because you do not know what is slow. Whether the gateway is slow, risk checks are taking long, the line is congested, or the exchange is slow to accept — there are four candidates, and the teams that would fix them are all different. If you hold a meeting in this state, each team says its own segment is fine, and with nobody having said anything wrong, nothing gets fixed.
So the first thing you do is put a ruler along the path. You decide the places an order passes through inside our system and leave a timestamp at each place. Receive, risk check done, session encoding done, line send, exchange accept, ack receive. Decide six places and five segments appear, and from then on "it is slow" turns into "this segment is this slow." The moment it turns, the team that will fix it is decided.
How it works
Segment latency is the later timestamp minus the earlier one. It looks easy, but there is a trap here. Whether the same machine stamped the two timestamps changes the meaning of the value.
If the two timestamps are within the same machine, the difference is almost unrelated to how accurate that machine's clock is. Even if the clock is 3 milliseconds ahead, both timestamps are 3 milliseconds ahead, so it disappears when you subtract. So segments within the same machine can be trusted, and here it is better to use a monotonic clock. A wall clock can go backward when time synchronization corrections come in, but a monotonic clock does not go backward.
Conversely, a segment whose two ends are on different machines takes in the clock errors of both machines as they are. The difference between the time our gateway sent it onto the line and the time the exchange stamped as having accepted it is the true line time plus the difference between the two clocks. If the exchange's clock is behind ours, this value comes out negative. It does not mean it traveled faster than light; it is the ruler that is wrong.
But a round trip cancels out. The line-send time and the ack-receive time are both stamped by our machine. The order went to the exchange and back in between, but since the two timestamps came from the same clock, however wrong the exchange's clock is, it does not enter the difference. This property — one-way is contaminated and round trip is clean — is also the basis of how the time synchronization of RFC 5905 measures offsets. So to look at a one-way segment you must receive offset reports and correct, and the segments that still remain negative after correction are real defects. The common mistake here is on the other side. When a whole segment comes out negative, people lump it together as "all the clock's fault," and the real one mixed in there disappears that way.
After collecting, you do not look at the mean. The latency distribution is stretched long to one side, so the mean is dragged around by a small number of slow cases. You look at percentiles instead. But it matters that percentiles have several definitions. The approach of sorting and using the value at that position as it is and the approach of interpolating between two values give different numbers from the same data. The statistics module documentation also lists several methods. If our dashboard and the counterparty's contract use different definitions, then for the same line one side says it is a violation and the other says it is not. So you write the definition next to the data and implement it exactly.
And per-segment percentiles do not explain the overall tail. If you compute p99 per segment, the line segment is usually the largest. But if you pick the worst 1 percent of orders by overall latency and count each of those orders' largest segment, a different segment may come out. Per-segment p99 looks at different orders in each segment, while the overall tail looks at what was large within a single order. The two questions are different, so the answers are different. What decides where to fix is the latter.
The last is the latency budget. You set an upper bound for each segment and count the exceedances. The value of the budget is not a matter of taste but a decision. If you set it loose, only the abnormal gets caught, and if you tighten it, normal variation gets caught too and alerts become thousands a day. At that point nobody looks at those alerts, so you should make two sets of budgets, actually count the exceedances, and then decide.
What it looks like in the field
Once, because the line segment came out negative, the measurement of that segment was discarded entirely. Months later, when we talked with the exchange about response time, we had no numbers for that segment. It was a job that only needed receiving the offset report and subtracting, but the person who saw the negative judged "the measurement is broken" and deleted it from the dashboard, and that was all.
Another time, a line upgrade was decided because the line segment had the largest p99. The reports of slow orders did not decrease even after the upgrade. When we reopened the slowest orders, the segment that was large in those orders was not the line but the wait between the risk check and encoding, and it appeared only in particular time windows. It was the price of treating per-segment p99 and the overall tail as the same question.
What you will do in the next lab
You build as data the segment timestamps for 1,500 orders, per-host clock offset reports, two sets of latency budgets, and the percentile definition. You compute segment latencies to find where negatives appear, and correct with the offsets to separate clock error from real defects. Then you count separately the per-segment percentiles and the segments that make up the overall tail and confirm that the two differ, and compare the exceedance counts of the two budgets. At the end you write down which segment got worse from which minute, together with the hypotheses ruled out by the data.