TT Lab
Get started
Learn Learning paths Courses

Building an EAI Middleware Layer

File Integration Breaks at the Boundaries

Continue in TT Lab

In one line

File integration is not an outdated method but the standard method for bulk, settlement and reconciliation. The heart of safe file reception is three things — take only files that are fully written (a completion flag), check that what you fetched equals what was sent (byte count, checksum, trailer), and get the same result if you run it twice (remote marking and per-file load records). And at the end, check against our ledger (reconciliation).

Why it was needed

Tens of thousands of insurance claims, card purchase slips, payroll transfer statements, end-of-day settlement — these are not called through an API one at a time. A day's worth is made into one file and handed over at a fixed time, and the receiving side loads it at once. There are many cases where the counterparty does not open an API, and above all, a file makes the boundary of the day clear. Because the "claims of September 23" is one file, you check the count and total, and if they are wrong, you simply fetch the whole file again.

But nearly all file integration incidents happen at boundaries. You fetch a file the counterparty is still writing and load only half of it. You load a file broken in transit without knowing. You rerun the batch and load yesterday's file again, so claims are booked twice. The load succeeded, but you learn only at month-end closing that three records differ from our intake ledger. The steps in this module correspond 1:1 to these four incidents.

How it works

SFTP. SFTP is a file transfer protocol that runs over an SSH connection. In production you connect to the other server with sftp user@host, and at that point SSH verifies the other server's host key (a setting that accepts a key it sees for the first time without asking allows man-in-the-middle attacks). An automated batch cannot ask and answer interactively, so it passes a list of commands with -b <배치파일> (batch file), and in batch mode, if one command fails, it stops there (sftp(1)).

This lab runs inside a single Pod. Instead of setting up the partner's SFTP server separately, complete with sshd (the SSH server), it uses sftp's -D option — because what you learn in this lab is not operating an SSH server but the rules for exchanging files. The manual describes -D as "connect directly to a local sftp server without going through ssh." The SFTP protocol (listing, getting, renaming) is the same as in production, and what was cut is the SSH connection, authentication and host key verification. Do not forget that those three are always added when you use the batch in production.

The completion flag. A small marker file uploaded after the file is fully uploaded. The receiving side fetches only files that have a flag. If you write the byte count and sha256 in the flag, even damage in transit is caught. SHA-256 is a hash that gives a completely different value if even one bit of the input differs, so it picks out cases where the size is the same but the content is broken.

Atomic rename. If you write the incoming file straight to its final name, another batch (load) can pick up that file while it is still being received. You receive it fully under a temporary name (.part), verify it, and then rename it to the final name. POSIX's rename() atomically changes the target name within the same filesystem — other processes see only the old state or the new state, never an intermediate one (rename). mv across a different filesystem is copy and then delete, so it does not have this property. That is why the temporary file goes in the same place as the final directory.

The same result when rerun. A batch is certain to be rerun (after a failure, after outage recovery, by someone's mistake). A fetched file is marked remotely by adding .fetched to the end of its name, and loads are recorded per file name (load_log) so that the same file is never loaded twice. A load treats one file as one transaction — it all goes in or nothing goes in. If a row with a duplicate claim number appears partway, it rolls back up to the previous rows. A file that went in only halfway cannot be fixed even by rerunning.

Header and trailer. There are boundaries inside the file too. H (header) says what file it is (date, sending and receiving institutions), and T (trailer) says how many records and what total. The receiving side counts and adds the D lines itself and checks against the trailer. If there is no trailer, the file did not arrive to the end. Line length is measured in bytes — if a line containing a Hangul name is re-saved as UTF-8, the characters look the same but the line gets longer (the same trap as module 1).

Reconciliation. Loading is not the end. You match the partner file against our intake ledger by claim number, and divide the records into those present on both sides with the same amount (matched), those only at the partner, those only at ours, and those with different amounts. A discrepancy is usually a business event — we missed an intake, or the partner did not remove cancellations. The reconciliation result is a report that people read, so it is ordered by claim number, and shows the amounts from both sides.

What it looks like in the field

The most common incident is a batch that trusts only the promise "we upload before 20:00" and fetches at 20:00. One day the partner's server is slow and writing finishes at 20:01, and we load a half file. If you synchronize by time without a completion flag, it will surely happen like this once. The second is the rerun incident. The person handling outage recovery ran the batch once more "just in case," and with no load record, the same file went in twice. The third is not doing reconciliation. Integration judges success by "it finished without errors," but the business judges by "the numbers match." Finally, in production SFTP batches there are cases where a setting that turns off host key verification (StrictHostKeyChecking=no) is put in for convenience, which means exchanging claim data without verifying the other server.

What we do in the next lab

You create a remote directory with the partner fixture and look at a listing with an SFTP batch. Then you grow the receiving script fetch.sh — only files with flags, checksum comparison and quarantine, temporary names and remote marking. Next you build the header and trailer validator check.py, the per-file loader load.py and the ledger reconciliation recon.py, and produce a reconciliation report with today's file.