TT Lab
Get started
Learn Learning paths Courses

Capital Markets and Settlement

An Order Is a Line of Events, Not a State

Continue in TT Lab

In one line

An order's final state is not a field stored somewhere. It is the result of folding many events in order, and if the folding rule and the folding order are not fixed, the same log gives different answers.

Why this was needed

Almost every time you join a brokerage support team, it starts with this sentence: "We got the execution notice, but the screen shows it as canceled."

The first reaction is to assume the screen is wrong or the DB was updated late. So you open the orders table and look at the status column. It says cancelled. If you do not know what to do next, the investigation stops right there.

It stops because you believed the order state was a single value. In reality it looks like this.

09:31:04.118  new           수량 300, 지정가 71,500
09:31:04.180  ack           거래소가 접수를 확인
09:44:12.007  partial_fill  150주 체결, 누적 150
09:47:55.310  cancel_req    잔량 취소 요청
09:47:55.402  fill          150주 체결, 누적 300
09:47:55.408  cancelled     잔량 취소 확정

From these six lines you have to extract "what is this order?" If you look only at the last event, it is a cancel. If you look only at the executed quantity, it is fully filled. Both are facts in this log, and which one you call the state is a rule we decide. In a system where the rule is not stated explicitly, several places in the code each use a different rule, and so the screen, the notification email, and the settlement batch give different answers.

How it works

The function that folds the events looks roughly like this. What matters here is not the code but what judgment each line carries.

new           주문 수량이 정해진다. 상태는 '접수'
ack           거래소가 받았다. 상태는 '유효'
reject        거절. 여기서 끝난다
replace       주문 수량이 바뀐다. 상태는 그대로
cancel_req    아무것도 바꾸지 않는다   ← 여기가 첫 번째 함정
cancelled     상태는 '취소'
fill 계열     누적 체결에 더하고, 주문 수량에 닿으면 '체결'

The key point is that cancel_req does not change the state. Requesting a cancel and being canceled are different facts. The request is something we sent, and the confirmation is something the exchange gave us. Implementations that change the state to cancelled at request time are common in practice, and with them an order that got a fill after the request stays recorded as canceled. The customer received the shares, but they are not in our books.

The second trap is that partial fills arrive several times. A 300-share order executing in three pieces, 150 + 100 + 50, is not unusual; it is the default. Each fill carries only its own quantity, and the cumulative total is for us to count. The exchange sometimes sends cum_qty as well, but if that value differs from the one we counted, that in itself is an incident signal.

The third trap is that the same fill arrives twice. When a session drops and reconnects, or when you request a retransmission, fills you already received come again. That is why every fill carries a unique exec_id, and filtering on it is the receiver's responsibility. If you do not filter, the cumulative quantity inflates. An overshoot past the order quantity is noticeable, but one that does not overshoot stays wrong without any alert.

What it looks like in the field

At one brokerage, the nightly settlement batch failed every morning. The log had six messages saying "executed quantity exceeds order quantity," and the person in charge was fixing those six by hand every morning. It had gone on for six months.

The cause was that retransmitted fills were not being filtered. And the real problem was not those six. There were another twenty-some orders whose quantity had inflated for the same reason, but they did not exceed the order quantity and so left no message at all. The six fixed every morning were the tip of the iceberg, and nobody had counted the part below for six months.

Another common sight is the status field and the event log drifting apart and staying that way for years. The status field is updated by the real-time processor while the event log accumulates the raw events. Once the processor misses an event, that order's status field stays wrong forever. The event does not flow in again, and there is usually no procedure to recompute the status field. So when investigating, the first move is not to trust the status field and to rebuild it from the events.

What you will do in the lab that follows

First, the next reading covers clocks and ordering, and then you move on to the lab.

You build an event log of 900 orders yourself and reconstruct each order's final state from that log alone. Then you compare it with the status field held by the OMS and find the orders that disagree. The numbers differ between comparing only the status strings and comparing the executed quantity as well, and that difference is the core of this incident.