TT Lab
Get started
Learn Learning paths Courses

Integration and Deployment

The Request Returned 200 and Three Items Vanished

Continue in TT Lab

Summary

A batch API often gives a 200 for one request while putting per-item results inside the body, and if the caller looks only at the status code, the few items that failed vanish without leaving a trace anywhere. So you need to record per-item results, separate what to resend from what to fix, and finally do a reconciliation that matches the counts on both sides.

Why this was needed

"They say we sent 200 settlement records yesterday but only 187 are on their side." The first thing you check in this report is our log. But our log says all 200 succeeded. It is not lying. It is just that all we looked at was the status code of the request.

The response of a batch-transfer API usually looks like this.

{"batch_id": "B-07", "accepted": 27, "rejected": 3,
 "results": [{"item_id": "IT-0031", "status": "rejected", "reason": "unknown_account"},
             {"item_id": "IT-0032", "status": "accepted", "reason": null}]}

The request succeeded. The server received and processed the request, and the result is in the body. From HTTP's point of view, the 200 is correct — the 2xx defined by RFC 9110 means "the request was received, understood, and accepted," not "every item inside the request succeeded."

So this incident happens with no error and no alert. And only after days pile up is it discovered while matching against the other side's numbers.

How it works

There are four steps in the fix.

First, leave per-item results on our side. For every item sent, record status, reason, and 시도 횟수 (the attempt count). Without this record you cannot do the next three steps. A common mistake is recording "only failures," and then "what was never sent" and "what was sent and succeeded" cannot be told apart.

Second, split failures into two branches. Failures that are resolved by resending, and failures that only a person can fix.

Reason If resent What to do
Temporary hold, rate limit, temporary error It resolves Put it back in the queue
Unknown account, bad value The same answer comes Fix the data
Already processed (duplicate) The same answer comes Do nothing

Without this classification, one of two things happens. If you retry everything, the items that cannot be fixed circle in the queue forever, and if you retry nothing, even items that were caught only briefly go to human hands. A rate limit often comes as a 429 of RFC 6585 and gives a Retry-After — that clearly belongs on the resend side. A standard format for putting the failure reason in machine-readable form is the problem details of RFC 9457, which distinguishes branches by type.

Third, resend only the failures. If you resend the whole batch here, items that have already gone in go in twice. If the other side blocks duplicates, good, and if not, we cause an incident. The unit of retry is not the batch but the item.

Fourth, reconcile. Match "the count and total we recorded as successful" on the sending side with "the count and total that actually went in" on the receiving side. The two numbers must be equal, and if they differ, go down to which item is on only one side.

발신함 120건
   ├─ 성공 기록 112건 · 합계 X
   └─ 실패 기록   8건 (고쳐야 함 8)
상대 원장     112행 · 합계 X      ← 건수와 합계가 **둘 다** 맞아야 한다

You must not match only the counts. If one is missing and another goes in twice, the count stays the same. So compare totals or sets as well.

What it looks like in the field

First, assuming the order of results in the response is the same as the request order. Many APIs really do give it in the same order, but if the documentation does not say so, you must not trust it. Match and read by item_id.

Second, throwing a partial failure as an exception. If you raise "3 failed" as an exception and roll back the whole batch, you end up resending even the 27 that succeeded, and that creates duplicates.

Third, there is no limit on the retry count. If an item classified as a transient failure is actually a permanent failure, that item circles the queue forever. Record the attempt count and hand it to a person after a few tries.

Fourth, reconciling only after an incident. Reconciliation must not be an incident-investigation tool but something that runs every day. Only if you can compare yesterday's and today's can you tell since when it went out of sync.

Fifth, the remaining items have no owner. If you create "8 to fix" and do not write who will fix them, those 8 remain forever. The last column of the reconciliation table is a person's name.

What you will do in the next lab

You start a partner server that receives batch transfers and create 120 items to send. First you send by looking only at status codes and confirm in numbers the difference between our record and the other side's ledger. Then you read the per-item results and record them on our side, split the failure reasons into what to resend and what to fix, and pick and resend only what should be resent. Finally you build a reconciliation table that matches the counts and totals of the sending and receiving sides, and write the handling plan for the remaining items.