TT Lab
Get started
Learn Learning paths Courses

Capital Markets and Settlement

Twenty overfill alerts — how many were real?

Continue in TT Lab

Goal

In an execution report stream, you find illegal state transitions, quantity equation violations, cumulative regressions, overfills, and average price mismatches one by one, and then strip away those explained by duplicate reports and out-of-order arrival, leaving only the real defects.

Why it matters

Order data is not a size that a person can read by eye to find anomalies. Fortunately this domain has rules, so a filter can be built. There is a state transition table, the quantities obey an equation, and the average price can be recomputed. What is hard is not the checking itself but explaining what you filtered out. Duplicate reports inflate the sum and make it look like an overfill, and a swapped arrival order turns a perfectly good order into an illegal transition. If you hand over "30 anomalies" without separating these three, the person who receives it can do nothing.

Steps

  1. Use python3 to create transitions.json, orders.csv, and reports.csv in /root/orders/data.
  2. Arrange the transition table into /root/orders/reachable.json, write the terminal states to /root/orders/terminal.txt, and write the table's own contradictions to /root/orders/graph_defects.txt.
  3. Write the illegal transitions found when reading in the order received to /root/orders/illegal.csv.
  4. Write the reports where the quantity equation breaks to /root/orders/qty_break.csv.
  5. Write the places where the cumulative fill quantity went backward to /root/orders/cum_back.csv.
  6. Write the orders that exceeded the order quantity to /root/orders/overfill.csv.
  7. Write the reports whose average price differs from the weighted average to /root/orders/avgpx_break.csv.
  8. Split the cause per order and write it to /root/orders/verdict.csv, and summarize the counts by cause in /root/orders/summary.json.

Notes

Generate the transition table and the execution reports for twelve orders

Use python3 to create transitions.json, orders.csv, and reports.csv in /root/orders/data. Use the generation script as is, which uses no random numbers.

We cannot bring in the customer's data as it is, so we build synthetic data of the same shape. Without random numbers, the same data comes out no matter who runs it how many times, and you can compare each other's judgments. The grader converts the data to a canonical form and compares fingerprints, so if you edit it by hand, all the later steps get blocked.

Read the transition table and find the table's own contradictions

Write the list of next states per state to /root/orders/reachable.json, the terminal states one per line to /root/orders/terminal.txt, and the states that are declared terminal but still have outgoing transitions to /root/orders/graph_defects.txt.

A terminal state must have nowhere further to go. If the next-state list of a state the table declares terminal is not empty, the table contradicts itself, and the checker lets wrong reports that pass through that state go through. Sorting the lists makes comparison easier.

Read in the order received and find illegal transitions

Put the first line order_id,from_seq,from_status,to_status in /root/orders/illegal.csv, and write the transitions that are not in the table when each order is read in seq order.

Split by order, read in ascending seq, and compare each pair of adjacent reports. If the later state is not in the earlier state's next-state list, it is an illegal transition. A terminal state has an empty list, so every report that comes after it is caught.

Find the reports where the quantity equation breaks

Put the first line report_id,order_id,ord_status,expected_leaves,leaves_qty in /root/orders/qty_break.csv, and write the reports where the equation breaks for a live order or where the leaves quantity is not 0 in a terminal state.

For a live order the expected leaves quantity is the order quantity minus the cumulative fills, and in a terminal state it is 0. Read the list of terminal states from the data. If you raise every canceled order as a violation, you did not look at the state.

Find the places where the cumulative fill quantity went backward

Put the first line order_id,seq,prev_cum,cum in /root/orders/cum_back.csv, and write the points where, reading in seq order, the cumulative fill became smaller than in the report just before.

The cumulative fill does not decrease. The same value continuing is not a problem; look only for decreases. For seq, write the one from the report on the side that went back.

Find the orders that exceeded the order quantity

Put the first line order_id,order_qty,max_cum,sum_last in /root/orders/overfill.csv, and write the orders whose maximum cumulative fill or sum of executed quantities exceeded the order quantity.

Look at both. One is the case where the cumulative fill exceeds the order quantity, and the other is where adding up all the executed quantities exceeds it. The latter also happens when the same report comes in twice, so in this step do not remove duplicates; count as is. Splitting them is done in step 8.

Recompute the average price and compare

Put the first line report_id,order_id,reported_avg,recomputed_avg in /root/orders/avgpx_break.csv, and write the reports whose average price differs from the weighted average up to that point.

Accumulate the executed quantity and amount up to that report, divide, and round at the fourth decimal place. Use decimal and match it with ROUND_HALF_UP. The average of a report with no fills at all is 0. Write values to the fourth decimal place.

Strip away the duplicate and ordering problems and leave only the real ones

Put the first line order_id,cause,defects in /root/orders/verdict.csv and write one line for each order in which a defect shows, and summarize orders, clean, 중복보고, 순서뒤바뀜, 진짜결함, and defects in /root/orders/summary.json (the three Korean keys mean duplicate report, swapped order, and real defect).

Judge in order. If every check goes quiet once you remove the duplicate reports, it is 중복보고; if it additionally goes quiet when sorted by exchange time, it is 순서뒤바뀜; if it still remains, it is 진짜결함. Deduplication is by report identifier, and sorting is by exchange time and then seq. In the defects column, write only the names of the checks that remained to the end.