Capital Markets and Settlement
Why There Is a Sequence Number
In one line
A market data feed runs on a transport that guarantees neither ordering nor arrival, which is why every message carries a sequence number — that number is the only means of telling what is missing from what is late.
Why this was needed
Market data has requirements that are the opposite of other data. A value that has just arrived is more useful than an accurate value that is a few milliseconds late, and one symbol's quote changes dozens of times in a single second. That is why exchange market data usually goes out over UDP multicast, not TCP. There is no retransmission and no ordering guarantee, but in exchange everyone subscribed receives a message at the same time once it is sent.
No guarantee means that two things really happen.
유실 보낸 메시지가 우리에게 오지 않는다
재정렬 나중에 보낸 메시지가 먼저 도착한다
To the receiver, these two look identical. If 102 arrives after sequence number 100, you cannot tell at that moment whether 101 is lost or is still on its way. That is why counting gaps by arrival order always comes out inflated. In this course's lab you will see with real data a case where the arrival-order gap count is 1,138 while only 204 are truly missing.
The reason this difference matters is that what a feed handler does when it finds a gap is request a retransmission. If it makes 1,138 requests, the request traffic eats the line; when the line is congested, latency grows; when latency grows, reordering increases; and when reordering increases, more gaps show up. It grows the problem it created itself. Incidents in which the feed is paralyzed for minutes by this feedback loop right after the open really do recur.
How it works
Judge gaps by the set of sequence numbers. Collect the sequence numbers that arrived, and treat as missing only those within the expected range that never came at all. Late ones enter the set, so they drop out automatically. The judgment comes with a waiting time — usually you wait several tens of milliseconds, and if it is still absent, you then confirm it as missing.
There are three ways to fill what is missing.
A/B 중재 거래소가 같은 내용을 두 채널로 따로 보낸다.
한쪽 구멍은 대개 다른 쪽에 있다. 요청 없이 메워진다.
재전송 요청 빠진 순번을 콕 집어 다시 달라고 한다.
거래소 재전송 버퍼를 넘어가면 못 받는다.
스냅샷 복구 현재 호가창 전체를 다시 받는다.
무겁지만 한 번에 끝난다. 구멍이 길면 이쪽이 싸다.
The order is also the cost order. It is common for an implementation not to try A/B arbitration first and to request a retransmission right away, but in the lab data, 180 of the 204 holes in channel A can be filled just by looking at channel B. 88% of them need no request at all.
Duplicates are also normal. If you receive both the A and B channels, the same message arrives twice. If you take a retransmission, it arrives again. If you do not filter by sequence number, the executed-volume tally doubles, and if you build a trading-value ranking from that value, the whole ranking flips.
Look at latency by the tail, not the average. In the lab data the overall mean is 10,071 microseconds while the median is 1,049 microseconds. A mean ten times the median means a small number of very slow messages dragged the mean up, and that small number is exactly the incident window. If you set the mean as the monitoring metric, this incident passes as nothing on the metric.
For percentiles, you must pin down even the calculation method. The nearest-rank method and the linear-interpolation method give different values on the same data. If our dashboard and the exchange's SLA document use different definitions, then for the same line one side says it is a violation and the other says it is not. Meetings like that really do happen.
What it looks like in the field
One site had a report that "quotes jump only early in the session." The latency dashboard drew only the mean, and even right after the open it was calm at about 12 ms. Taking a capture and splitting it by minute showed that the p99 of two windows, minute 09:00 and minute 09:08, exceeded 380 ms, and 95% of the over-10 ms cases were concentrated in those two windows. The p99 of the other nine windows was 1.6 ms. A single mean was mixing two worlds into one number.
Latency exceeding 300 ms means the quote we are looking at is already a value from the past. If you compute a limit price from that value and place an order, it either does not execute or, conversely, executes at an unfavorable price. If p99 is 310 ms, it means that 1 order in 100 was placed looking at a picture that was 0.3 seconds old. This is why market data quality is money.
What you actually do when the sequence breaks
In a market data feed, a break in the sequence is not an exceptional situation but an everyday occurrence. So how you handle breaks determines the quality of the system.
First, detect the break reliably. If the received sequence number is larger than expected, the stretch in between is empty. If it is smaller, it is a duplicate or out of order. The two must be handled differently.
expected = last + 1
if seq == expected: 정상 처리
elif seq > expected: 갭 — 복구 시작
else: 중복 — 버린다(이미 반영했다)
When a gap appears, change that symbol's state to "suspect." If you keep processing, the calculations run on top of a wrong order book. Worse than giving no value at all is giving a wrong value as if it were normal.
The recovery path is one of three. Request a retransmission (replay), re-receive a snapshot, or switch to a backup feed. Whichever you use, do not discard the real-time messages that arrive during recovery; queue them up, and after applying the snapshot, release that queue in sequence-number order. If you do not keep to this order, a gap appears again right after recovery.
Bring the clocks you measure latency with down to one. Latency is the difference between the time the exchange stamped and the time we received, and if the two clocks are off, you get negative numbers or exaggerated latency. Synchronize with PTP or NTP, and store both timestamps so you can correct later.
Decide in advance what to drop during a surge. When messages per second exceed processing capacity, the queue grows. If you try to process everything then, latency keeps growing, and you end up emitting very old values as if they were the latest. A quote only has meaning when it is the latest, so skipping intermediate values (conflation) is right, and since not a single fill may be dropped, handle fills on a separate path.
Even tracking just two metrics makes a big difference. The number of gaps and the tail value of processing latency. The mean latency does not show a surge.
What you will do in the next lab
You build a capture of an A/B redundant feed with 12,000 sequence numbers yourself, and check how much the gaps counted by arrival order differ from the gaps counted by the set of sequence numbers. You separate what channel B can fill from the real holes, and then split the remaining holes into retransmission and snapshot. You measure latency with nearest-rank percentiles, split it by minute, and pinpoint the windows where the tail is concentrated.