TT Lab
Get started
Learn Learning paths Courses

Building an EAI Middleware Layer

The Same GUID Arrived Twice — Stop It with a Ledger

Continue in TT Lab

Goal

Attach a GUID-based state ledger (SQLite) to the relay layer to prevent double processing, confirm transactions with undetermined results by lookup, and safely reprocess only confirmed-unsent transactions.

Why it matters

When a timeout occurs, the channel resends with the same GUID, and core banking is not idempotent. If the relay layer does not remember "have I already sent this GUID, do I know the result," a resend is a double transfer. And a reprocessing run that resends an entire failure list is the most common large-scale incident — you have to distinguish transactions whose result is unknown from transactions that could not be sent.

Steps

  1. Write /root/eaimw/dedup/schema.sql. Table tx: guid (primary key), tx_code, state (a CHECK allowing only RECEIVED, SENT, DONE, UNKNOWN and FAILED), rsp_code, request (the request message, BLOB), response (the response message, BLOB), attempts (an integer) and updated_at (a string in the format of datetime('now')). It must not error even when applied twice (IF NOT EXISTS).
  2. Start with cp /opt/lab/fixtures/eaimw/dedup/relay_base.py /root/eaimw/dedup/relay.py. Add a --db <경로> (path) argument and apply the schema.sql in the same directory on startup. Before calling core banking, claim a row by GUID (SENT), and when you receive the result, store DONE and the response message. If a GUID that is DONE comes again, do not call core banking and return the stored response as is.
  3. If the same GUID arrives while the claimed GUID is still being processed (SENT), do not wait and answer E903 immediately. The claim must be atomic (the row count of INSERT ... ON CONFLICT(guid) DO NOTHING).
  4. Record a read timeout (E901) in the ledger as UNKNOWN and a connection failure (E902) as FAILED. If a GUID that is UNKNOWN comes again, do not call core banking and answer E901. If a GUID that is FAILED comes again, it may be sent again (attempts +1).
  5. /root/eaimw/dedup/resolve.py --db <경로> --core <URL> --min-age <초> (path, URL, seconds): only UNKNOWN rows whose last update is older than --min-age seconds are confirmed with core banking's lookup API (GET /v1/transfers/<guid>). On 200, DONE, 0000 and the response message (an R message built from the request header + a 45-byte body built from the lookup result); on 404, FAILED and E902. If there is a connection error during lookup, leave it as it is. Do not send the transfer again.
  6. /root/eaimw/dedup/reprocess.py --db <경로> --core <URL> --max-attempts <N> (path, URL, N): only rows that are FAILED with rsp_code E902 and attempts below N are sent again with the same GUID using the stored request message (after claiming, attempts +1, and the result recorded by the same rules as in step 2). UNKNOWN is never sent.
  7. /root/eaimw/dedup/purge.py --db <경로> --days <N> (path, N): delete only DONE rows older than N days. UNKNOWN, FAILED and SENT are kept regardless of the period.

Notes

The state ledger schema

Write the tx table (guid primary key, state CHECK, original request and response, attempt count, update time) in /root/eaimw/dedup/schema.sql.

Make it safe to apply twice with CREATE TABLE IF NOT EXISTS. CHECK (state IN (...)) blocks misspelled states. The originals are BLOBs.

A resend of a finished transaction — the stored answer as is

Copy relay_base.py and attach the ledger. Claim as SENT before sending, record the result as DONE with the response message, and answer a DONE resend with the stored response as is.

In handle, before calling call_core, claim a row with INSERT ... ON CONFLICT(guid) DO NOTHING, and if it already exists, look at the state with SELECT. Apply schema.sql with executescript on startup.

A resend while processing — E903 immediately

If a GUID in the SENT state comes again, do not wait and answer E903.

If the claim failed and it is not DONE, someone is processing it. You must judge by the row count of a single insert, not the two steps of look up then insert, so that even in concurrent resends only one side claims it.

What you don't know is UNKNOWN, what you couldn't send is FAILED

Record E901 as UNKNOWN and E902 as FAILED. Answer E901 to an UNKNOWN resend without calling, and send again for a FAILED resend.

When recording the result, choose the state by the response code. Re-claiming a FAILED row should also be atomic, by the row count of UPDATE ... WHERE state='FAILED'.

Confirm by lookup — do not send again

/root/eaimw/dedup/resolve.py confirms only old UNKNOWN rows with the lookup API (200→DONE, 404→FAILED/E902, error→leave as is).

Pick only the old ones with updated_at <= datetime('now', '-N seconds'). A transaction that just timed out may still be being processed by core banking. The response message is lhstd.reply(request message, '0000', body).

Reprocess only confirmed-unsent ones with the same GUID

/root/eaimw/dedup/reprocess.py sends again, with the stored request, only rows that are FAILED/E902 and whose attempts are below the limit.

The selection condition is everything — state, rsp_code, attempts. Before sending, re-claim as SENT and raise attempts. Use the GUID of the stored request message as it is.

Clean up the ledger — only what is safe to delete

/root/eaimw/dedup/purge.py --days N deletes only DONE rows older than N days.

Put both state and period in the WHERE of the DELETE. If you delete UNKNOWN, the basis for investigating that transaction is gone.