Our Log Says All Succeeded, Their Side Is Missing Twelve
Goal
You deal with the situation where only a few items in a batched request fail. You record per-item results on our side, separate failures to resend from failures to fix, resend only the failures, and finally build a reconciliation table that matches the counts and totals of the sending and receiving sides.
Why it matters
A batch-transfer API gives a 200 for one request while putting per-item results inside the body. From HTTP's point of view, the 200 is correct — a 2xx means the request was received and accepted, not that every item inside the request succeeded. So if the caller looks only at the status code, the few items that failed are left nowhere. Days pile up with no error or alert, and it is discovered while matching against the other side's numbers. There are four steps in the fix. Record per item, split failures into two branches, resend only the failures, and match the counts on both sides. The third is especially important — if you resend the whole batch, items that have already gone in go in twice. The unit of retry is not the batch but the item. The grader does not believe your sentences. It starts the partner server fresh on a port the grader chooses, actually runs your outbox and sender against the grader's database, and compares with the other side's ledger.
Steps
- Create /root/recon/receiver.py, run it on port 8010, and create the 120 items to send in /root/recon/outbox.db with /root/recon/gen_outbox.py.
- Send with
--naiveof /root/recon/send.py, looking only at status codes, and write the difference between our record and the other side's ledger in /root/recon/naive.json. - Make send.py read the per-item results and record them in the
senttable as status, reason, and attempts. - Create /root/recon/classify.py so that it splits failure reasons into what to resend and what to fix.
- Make the
--retryof send.py pick and resend only the items that may be resent. - With /root/recon/recon.py, compare the sending and receiving sides and create /root/recon/recon_result.json.
- Write the handling plan for the remaining items in /root/recon/unresolved.json.
- Report in four sections in /root/recon/recon_report.md.
Notes
- Partner run contract:
python3 /root/recon/receiver.py --port <포트>(the placeholder is the port).POST /batchreceives{"batch_id": ..., "items": [{"item_id", "account", "amount"}]}and returns, with a 200,{"batch_id", "accepted", "rejected", "results": [{"item_id", "status", "reason"}]}.GET /ledgeris{"rows": [...], "count": n, "total": m}. - The partner's rejection reasons are four.
invalid_amount(the amount is not a positive integer),unknown_account(the only known accounts are AC-01 through AC-08),duplicate(already in the ledger), andtemporary_hold(holds items whose amount is a multiple of 97 only the first time). The judging order is this sequence. - Outbox run contract:
python3 gen_outbox.py [--db <경로>](the placeholder is the path) creates 120 rows ofoutbox(item_id, account, amount, batch_id)and an emptysent(item_id, status, reason, attempts). Among the 120, there must be mixed in 5 with a wrong account, 3 with an amount of 0, and 4 with a multiple of 97, and the batches are 4 of 30 items each. - Sender run contract:
python3 send.py --base <URL> --db <sqlite> [--batch-size 30] [--retry] [--naive]outputs{"sent": n, "accepted": n, "rejected": n, "by_reason": {...}, "http_ok": n}. By default it sends the items not yet sent, and--retrysends only those among the previously rejected items that may be resent.--naivedoes not read the per-item results but looks only at status codes and writes nothing in thesenttable. - Classifier run contract:
python3 classify.py --reason <이유>(the placeholder is the reason) outputs{"retryable": true|false, "action": "requeue"|"fix_data"|"ignore"}. For an unknown reason, answer with the safe side (do not resend). - Comparator run contract:
python3 recon.py --base <URL> --db <sqlite> --out <결과 JSON>(the placeholder is the result JSON) outputsoutbox_items,accepted,rejected,not_sent,partner_rows,partner_total,our_accepted_total,only_ours,only_theirs,by_reason,unresolved, andbalanced.balancedis true when there are no items not sent, no items on only one side, and both the counts and totals match. - Remaining-items format:
{"total": n, "by_action": {...}, "items": [{"item_id", "reason", "action", "owner"}]}.owneris the name of a person or team — if there is no owner, that item stays forever. - Common mistakes: looking only at status codes, recording only failures (it cannot be distinguished from never having been sent), resending the whole batch (duplicates arise), and matching only counts and not looking at totals.
- Run the server in the background, wait until
/healthis 200, and then move on. The grader does not look at the process you left running but restarts the scripts directly.
A batch-receiving partner and 120 items to send
Create /root/recon/receiver.py and run it on port 8010, create and run /root/recon/gen_outbox.py to put 120 items in /root/recon/outbox.db. There must be mixed in 5 with a wrong account, 3 with an amount of 0, and 4 with a multiple of 97.
The partner always gives a 200 for the request itself and puts failures in the body's results. It judges rejection in the order amount, account, duplicate, temporary hold. Make the outbox as two tables, what to send and the sent results — if you mix them in one table, "never sent" and "sent and failed" cannot be told apart.
What you miss by looking only at status codes
Add --naive to /root/recon/send.py so that it looks only at status codes and counts everything as a success. Write the difference between that result and the other side's ledger in /root/recon/naive.json as batches, http_ok, assumed_sent, partner_rows, and gap.
The request really did succeed. The failures are inside the body. If our record says 120 succeeded, count how many rows are in the other side's ledger and the difference shows itself. In this step you write nothing in the sent table.
Leave per-item results on our side
Make send.py read the results of the response body and, for each item, leave in the sent table its status, reason, and attempts. You must record both successes and failures. The second run must not send items already sent again.
If you write only failures, "never sent" and "sent and succeeded" cannot be told apart. It is safer not to assume the order of results in the response is the same as the request order but to match and read by item_id. An item not yet sent is one that is in outbox and not in sent.
What to resend and what to fix
Create /root/recon/classify.py so that it judges the reason received with --reason as {"retryable": ..., "action": ...}. action is one of requeue, fix_data, and ignore, and for an unknown reason it answers toward not resending.
A temporary hold resolves if you simply resend. An unknown account and a bad amount give the same answer however many times you send, so a person has to fix the data. An item already in the other side's ledger has nothing to resend or fix. Without this classification, items that cannot be fixed circle the queue forever.
Again by item, not by batch
Make the --retry of send.py pick and resend only the items that may be resent among the previously rejected items and update the result in sent. An item that already succeeded must never be sent again.
If you resend the whole batch, items that have already gone in go in twice. If the other side blocks duplicates, good, and if not, we cause an incident. Call the classifier and pick only the retryable ones, and raise attempts so that later you know which attempt it was.
Match the sending side and the receiving side
With /root/recon/recon.py, compare the outbox, our record, and the other side's ledger and create /root/recon/recon_result.json. You must look not only at counts but also at totals and items on only one side, and balanced is true only when everything matches.
If you match only counts, you miss the case where one is missing and another went in twice. If you subtract the id set we wrote as successful and the other side's ledger id set from each other, which side has it alone comes out right away. Also count whether there are items not yet sent.
The remaining items must have an owner
In /root/recon/unresolved.json, write the items not yet resolved as total, by_action, and items (item_id, reason, action, owner). action must equal the classifier's answer and owner cannot be left empty.
If you create "8 to fix" and do not write who will fix them, those 8 remain forever. The last column of the reconciliation table is the name of a person or team. You can make items by taking out as they are the items left as rejected in the sent table.
Reconciliation report
In /root/recon/recon_report.md, write four sections, ## 무엇을 놓치고 있었나, ## 건별 결과를 읽고 나서, ## 다시 보낼 것과 고칠 것, and ## 대사표와 남은 것 (in order: what we were missing, after reading the per-item results, what to resend and what to fix, the reconciliation table and what remains). The numbers from naive.json and recon_result.json must be in the body.
The reader is someone who said "our log says they all succeeded." First explain that the log is not lying but only looked at the status code of the request. Then show the counts by reason and the two numbers of the reconciliation table.