Capital Markets and Settlement
The restarted process sent the same orders all over again
Goal
From the checkpoint and send log of a sending process that died and came back, you find the reprocessing window, pick out only the orders actually accepted twice by checking against the exchange acknowledgments, and derive a deterministic order id and a deduplication window length from the data and leave a recovery runbook.
Why it matters
Reprocessing cannot be eliminated. If you save the checkpoint after sending, that stretch goes out again when the process dies, and if you save it before sending, that stretch is missed. There is no safe point between the two, so a sending path usually chooses to send again and filters duplicates on the receiving side. To filter, the same signal must give the same order id however many times it is read again, and serial numbers assigned fresh on each run or ids mixed with the time break that condition. And amendments and cancels point to a previous id, so if the id rule wobbles, the chain breaks too and the order that was meant to be canceled remains.
Steps
- Use
python3to create six data files in/root/cap/replay/data. Use the generation script as is. - Compare the checkpoint and the send log and write the reprocessing window in
/root/cap/replay/crash.txt. - Write the orders that will go out again in that window to
/root/cap/replay/replay.csvand/root/cap/replay/replay.txt. - Build order ids from only the signal's stable fields and write them to
/root/cap/replay/clordid.csvand/root/cap/replay/idcheck.txt. - Check against the acknowledgments and write the orders accepted twice to
/root/cap/replay/dup.csvand/root/cap/replay/dup.txt. - Compute each of the two checkpoint save times and write one line each to
/root/cap/replay/modes.csv. - Write the cancels and amendments that cannot point to their original order to
/root/cap/replay/broken.csvand/root/cap/replay/broken.txt. - Leave the deduplication window length in
/root/cap/replay/window.txtand the recovery runbook in/root/cap/replay/runbook.md.
Notes
- The data is in
/root/cap/replay/data.signals.jsonlis the input signals with offsets,sent.csvis the order messages that actually went out,acks.jsonlis the exchange acknowledgments,checkpoint.jsonis how far it committed,incident.jsonis the time it died and the time it came back, andpolicy.jsonis the order id rule and the commit interval. - All times are fixed in the data. If you use
dateto write today, the same data gives a different answer each day. - The order id rule is in
clordidofpolicy.json. Joinfieldsin that order withsep, summarize withhash, cut only the firsthex_lencharacters, and attachprefix. Convert all values to strings before joining. - Signals whose
intentisholddid not cross the threshold, so no order went out. The number of signals and the number of orders differ. - Common mistake 1: counting duplicates from the send log alone. Items rejected in the acknowledgments did not remain at the exchange, so they are not double acceptances.
- Common mistake 2: forgetting the batch boundary in step 6. The checkpoint is saved not per item but every
commit_everyitems. - Specifications: you can check the meaning of
ClOrdIDandOrigClOrdIDat FIX Trading Community standards and FIXimate. All the symbol codes and run names in this lab are synthetic.
Generate the data the incident left behind
Use python3 to create the six files signals.jsonl, sent.csv, acks.jsonl, checkpoint.json, incident.json, and policy.json in /root/cap/replay/data. Use the generation script as is, which uses no random numbers.
An air-gapped network has no samples to download, so you generate the data yourself first. Without random numbers, the same data comes out no matter who runs it how many times, and you can compare each other's judgments. The grader converts the file contents to a canonical form and compares fingerprints, so if you edit the data by hand, all the later steps get blocked.
Find where it started reading again
Write the six lines crashed_run, resumed_run, committed_offset, last_sent_offset, replay_from, and replay_to as key=value in /root/cap/replay/crash.txt.
The checkpoint says how far it had processed, and the send log keeps how far the run that died actually sent. If the two values differ, the stretch between them is the window that gets read again. A restart reads from the offset after the committed one.
Count what goes out again if you restart as is
Put the first line offset,intent,symbol,side,qty,limit_px in /root/cap/replay/replay.csv and write, one per line, the signals in the reprocessing window that send an order, and write the three lines window_signals, window_orders, and window_notional_krw in /root/cap/replay/replay.txt.
The window is from replay_from to replay_to found in the previous step. Counting all the signals in it is different from counting only the signals that send an order. The amount is the sum of order quantity times limit price, adding only the new orders and amendments that create exposure.
Build the same order id even when read again
Put the first line offset,intent,clordid,orig_clordid in /root/cap/replay/clordid.csv and write one line for each signal that sends an order, and write the five lines order_signals, distinct_deterministic_ids, offsets_sent_twice, offsets_with_two_sent_ids, and chain_resolvable in /root/cap/replay/idcheck.txt.
The rule is in clordid of policy.json. Join fields in that order with sep, summarize with sha256, and attach prefix to the first hex_len characters. For a new order leave orig_clordid as a blank cell, and for a cancel or amendment write the value from applying the same rule to the signal that ref_offset points to. chain_resolvable is the number of cancels and amendments whose orig made that way actually exists in this table.
Keep only the orders that were really accepted twice
Put the first line offset,msg_type,clordid_a,clordid_b,symbol,side,qty,notional_krw in /root/cap/replay/dup.csv and write the offsets for which the messages of both runs were accepted, and write the five lines dup_accepted, dup_new, dup_cancel, dup_replace, and dup_notional_krw in /root/cap/replay/dup.txt.
An offset that appears twice in the send log is not necessarily a double acceptance. Keep only those for which both messages are accepted in acks.jsonl. clordid_a is the side earlier in ascending run-name order, and clordid_b is the later side. The amount is quantity times price, and dup_notional_krw adds only the new-order lines.
What differs depending on when you save the checkpoint
Put the first line mode,committed_offset,resume_offset,duplicate_orders,missing_orders in /root/cap/replay/modes.csv and write the two lines commit_after and commit_before.
The checkpoint is saved every commit_every items. With the save-after-sending approach, only up to the last batch finished just before dying is committed, and with the save-before-sending approach, the end offset of the batch being processed is already committed. A restart reads from the offset after the committed one. The former produces duplicates and the latter produces omissions.
The order the cancel pointed to is gone
Put the first line run_id,clordid,offset,msg_type,orig_clordid,reason in /root/cap/replay/broken.csv and write the cancels and amendments that cannot point to their original order, and write the four lines chain_messages, broken_total, unknown_orig, and orig_rejected in /root/cap/replay/broken.txt. reason is unknown_orig or orig_rejected.
There are two ways the chain breaks. If the order id it points to is nowhere in the send log, it is unknown_orig, and if it exists but the exchange rejected that original order, it is orig_rejected. Both have the same result — the order that was meant to be canceled remains. chain_messages is all the cancels and amendments in the send log.
Derive the deduplication window from the data and write the runbook
Write the four lines offsets_sent_twice, max_replay_delay_sec, outage_sec, and recommended_window_sec in /root/cap/replay/window.txt, and in /root/cap/replay/runbook.md write the five sections ## 무슨 일이 있었나, ## 왜 두 번 나갔나, ## 돈으로 얼마인가, ## 복구 절차, and ## 무엇을 고쳐야 하나 in that order, each at least 60 characters (the five Korean section titles mean: what happened, why it went out twice, what it amounts to in money, the recovery procedure, and what to fix). In the body of the runbook, write as numbers the dup_notional_krw value from step 5 and the recommended_window_sec value from this step.
The reprocessing delay is the difference between the time the same offset first went out and the time it went out again. Measure it over all the offsets that went out twice and use the largest value. The recommended window is that value multiplied by dedup_safety_multiple in policy.json and rounded up to a multiple of dedup_round_sec. outage_sec is the difference between the two times in incident.json.