Capital Markets and Settlement
Twenty overfill alerts — how many were real?
Goal
In an execution report stream, you find illegal state transitions, quantity equation violations, cumulative regressions, overfills, and average price mismatches one by one, and then strip away those explained by duplicate reports and out-of-order arrival, leaving only the real defects.
Why it matters
Order data is not a size that a person can read by eye to find anomalies. Fortunately this domain has rules, so a filter can be built. There is a state transition table, the quantities obey an equation, and the average price can be recomputed. What is hard is not the checking itself but explaining what you filtered out. Duplicate reports inflate the sum and make it look like an overfill, and a swapped arrival order turns a perfectly good order into an illegal transition. If you hand over "30 anomalies" without separating these three, the person who receives it can do nothing.
Steps
- Use
python3to createtransitions.json,orders.csv, andreports.csvin/root/orders/data. - Arrange the transition table into
/root/orders/reachable.json, write the terminal states to/root/orders/terminal.txt, and write the table's own contradictions to/root/orders/graph_defects.txt. - Write the illegal transitions found when reading in the order received to
/root/orders/illegal.csv. - Write the reports where the quantity equation breaks to
/root/orders/qty_break.csv. - Write the places where the cumulative fill quantity went backward to
/root/orders/cum_back.csv. - Write the orders that exceeded the order quantity to
/root/orders/overfill.csv. - Write the reports whose average price differs from the weighted average to
/root/orders/avgpx_break.csv. - Split the cause per order and write it to
/root/orders/verdict.csv, and summarize the counts by cause in/root/orders/summary.json.
Notes
- Steps 3 through 7 read in the order received (seq). Fixing the order is done in step 8.
- Apply the quantity equation only to live orders. In a terminal state the leaves quantity must be 0. The list of terminal states is in
transitions.json. - Compute the average price with
decimaland round at the fourth decimal place withROUND_HALF_UP. If you usefloat, perfectly good data comes out looking off. - The causes in step 8 are decided in order. If the anomaly disappears once duplicates are removed, it is
중복보고(the Korean word for "duplicate report"); if it additionally disappears after sorting by exchange time, it is순서뒤바뀜(the Korean word for "swapped order"); if it still remains, it is진짜결함(the Korean word for "real defect"), and in thedefectscolumn you write the names of the remaining checks joined with|(for중복보고and순서뒤바뀜it is없음, the Korean word for "none"). - Common mistake 1: applying the equation regardless of state and raising every normal cancel as an anomaly.
- Common mistake 2: judging overfill without removing duplicate reports.
- Public documents: FIX Standards, Fiximate, decimal. The transition table and the twelve orders are synthetic data for this lab.
Generate the transition table and the execution reports for twelve orders
Use python3 to create transitions.json, orders.csv, and reports.csv in /root/orders/data. Use the generation script as is, which uses no random numbers.
We cannot bring in the customer's data as it is, so we build synthetic data of the same shape. Without random numbers, the same data comes out no matter who runs it how many times, and you can compare each other's judgments. The grader converts the data to a canonical form and compares fingerprints, so if you edit it by hand, all the later steps get blocked.
Read the transition table and find the table's own contradictions
Write the list of next states per state to /root/orders/reachable.json, the terminal states one per line to /root/orders/terminal.txt, and the states that are declared terminal but still have outgoing transitions to /root/orders/graph_defects.txt.
A terminal state must have nowhere further to go. If the next-state list of a state the table declares terminal is not empty, the table contradicts itself, and the checker lets wrong reports that pass through that state go through. Sorting the lists makes comparison easier.
Read in the order received and find illegal transitions
Put the first line order_id,from_seq,from_status,to_status in /root/orders/illegal.csv, and write the transitions that are not in the table when each order is read in seq order.
Split by order, read in ascending seq, and compare each pair of adjacent reports. If the later state is not in the earlier state's next-state list, it is an illegal transition. A terminal state has an empty list, so every report that comes after it is caught.
Find the reports where the quantity equation breaks
Put the first line report_id,order_id,ord_status,expected_leaves,leaves_qty in /root/orders/qty_break.csv, and write the reports where the equation breaks for a live order or where the leaves quantity is not 0 in a terminal state.
For a live order the expected leaves quantity is the order quantity minus the cumulative fills, and in a terminal state it is 0. Read the list of terminal states from the data. If you raise every canceled order as a violation, you did not look at the state.
Find the places where the cumulative fill quantity went backward
Put the first line order_id,seq,prev_cum,cum in /root/orders/cum_back.csv, and write the points where, reading in seq order, the cumulative fill became smaller than in the report just before.
The cumulative fill does not decrease. The same value continuing is not a problem; look only for decreases. For seq, write the one from the report on the side that went back.
Find the orders that exceeded the order quantity
Put the first line order_id,order_qty,max_cum,sum_last in /root/orders/overfill.csv, and write the orders whose maximum cumulative fill or sum of executed quantities exceeded the order quantity.
Look at both. One is the case where the cumulative fill exceeds the order quantity, and the other is where adding up all the executed quantities exceeds it. The latter also happens when the same report comes in twice, so in this step do not remove duplicates; count as is. Splitting them is done in step 8.
Recompute the average price and compare
Put the first line report_id,order_id,reported_avg,recomputed_avg in /root/orders/avgpx_break.csv, and write the reports whose average price differs from the weighted average up to that point.
Accumulate the executed quantity and amount up to that report, divide, and round at the fourth decimal place. Use decimal and match it with ROUND_HALF_UP. The average of a report with no fills at all is 0. Write values to the fourth decimal place.
Strip away the duplicate and ordering problems and leave only the real ones
Put the first line order_id,cause,defects in /root/orders/verdict.csv and write one line for each order in which a defect shows, and summarize orders, clean, 중복보고, 순서뒤바뀜, 진짜결함, and defects in /root/orders/summary.json (the three Korean keys mean duplicate report, swapped order, and real defect).
Judge in order. If every check goes quiet once you remove the duplicate reports, it is 중복보고; if it additionally goes quiet when sorted by exchange time, it is 순서뒤바뀜; if it still remains, it is 진짜결함. Deduplication is by report identifier, and sorting is by exchange time and then seq. In the defects column, write only the names of the checks that remained to the end.