TT Lab
Get started
Learn Learning paths Courses

System Integration (EAI)

Why File Interfaces Are Still the Best Option

Continue in TT Lab

Summary

File integration has survived for bulk volume, clear boundaries, and easy reprocessing, and in exchange you must deal directly with fixed-length layouts where position is meaning and the problem of picking a file up before it is fully written.

Why still files

There are REST and queues, so why exchange files? The reasons are clear.

So file integration does not disappear at SI sites. If anything, a large part of the nightly batch work is still files.

Fixed-length messages — position is meaning

20260801HONG                    0000012500Y
|      ||                      ||        ||
1-8    9-28                     29-38     39
일자    고객명(20)               금액(10)   여부(1)

A layout specification must have start position, length, type, alignment, and padding character.

Item Convention
Character field Left-aligned, right-padded with spaces
Numeric field Right-aligned, left-padded with 0
Amount As an integer with no decimal point, with the digits stated in the specification
Negative Overlay the sign on the last digit, or a separate sign field
Date 8 digits YYYYMMDD

Trouble starts when Korean text is included. Depending on whether "length 20" is 20 bytes or 20 characters, you get entirely different files. In EUC-KR a Korean character is 2 bytes, and in UTF-8 it is 3. And with fixed length, if even one byte shifts, everything after it breaks.

That is why EUC-KR (or CP949) is still alive in domestic external integration. Korean characters fit neatly into 2 bytes, which makes digit calculation easy. If a received file looks garbled, suspect the encoding first and try converting it.

iconv -f EUC-KR -t UTF-8 SALES_20260801.dat > SALES_utf8.dat

Header and trailer — the file's own self-verification

A well-designed file interface has a header and a trailer.

H20260801SALES        0001          ← 헤더: 구분, 일자, 업무, 파일순번
D20260801HONG        0000012500     ← 데이터
D20260801KIM         0000030000
T0000000002000000042500              ← 트레일러: 건수, 금액 합계

The trailer exists for one reason — to prove by itself that the file arrived intact.

If it was cut during transfer, a line was lost in the middle, or it was corrupted during encoding conversion, this check catches it. If you do not do this check, you process a half-arrived file as normal. And you discover that at month-end closing as "the numbers do not match."

Trailer verification must be done before processing. If you do it during processing, half is already reflected.

Completion flag — the problem of picking up before writing is done

What happens if a batch picks up a file while it is being uploaded over FTP? It processes a half file. If there is trailer verification it gets caught, but if not, it goes straight in.

The standard convention is a completion flag file.

SALES_20260801.dat      ← 데이터 (전송 중일 수 있음)
SALES_20260801.dat.ok   ← 이 파일이 생겨야 처리 대상

After the sender has sent all the data, it creates an empty .ok file. The receiver processes only those that have a .ok. It is simple but very effective.

It must be stated in the specification. If the "flag file convention" exists on only one side, the other side just uploads the data, and our batch processes nothing, forever. And no error occurs at all. Silently doing nothing is discovered the latest.

Filename rules are an interface too

<업무코드>_<기준일자>_<순번>.<확장자>
SALES_20260801_001.dat

Retention after processing

Do not delete processed files; retain them.

/data/if/archive/20260801/SALES_20260801_001.dat.gz

The reason for retention is reprocessing and disputes. To answer "what we sent that day was 12,000 records," you need the original.

Reconciliation — the last step of file integration

Processing the file is not the end. You must check whether the numbers match on both sides.

Reconciliation item Method
Count Sent count == received processed count + error count
Amount total Sent total == received total
Key set Sent key list - received key list = empty set

It is easy to feel reassured if only the count and total match, but if one record is missing and another is inserted twice, the count still matches. So look at the amount total too, and for interfaces where accuracy matters, compare the key set as well.

Leave the reconciliation result as a file. And if there is a mismatch, make it notify the owner automatically. A method where a person opens the reconciliation result every day gets skipped on a busy day, and that skipped day is the day with the problem.

What it looks like in the field

The incidents that occur in file integration are of set kinds.

All three arise from having no procedure to check that the file can be trusted before reading it. So the first step of file integration is not parsing but verification.