Why File Interfaces Are Still the Best Option
Summary
File integration has survived for bulk volume, clear boundaries, and easy reprocessing, and in exchange you must deal directly with fixed-length layouts where position is meaning and the problem of picking a file up before it is fully written.
Why still files
There are REST and queues, so why exchange files? The reasons are clear.
- It is strong at bulk. Sending 1 million records with per-item calls takes hours, but with a single file it takes minutes.
- Boundaries are clear. A unit like "the sales file dated August 1" is natural to business owners. Counts and totals can be reconciled precisely.
- Reprocessing is easy. You just put the file in again. It is more intuitive than queue reprocessing.
- The other side can do only that. A 20-year-old system, an external institution, bank EDI. "Please give it to us via REST" does not work.
So file integration does not disappear at SI sites. If anything, a large part of the nightly batch work is still files.
Fixed-length messages — position is meaning
20260801HONG 0000012500Y
| || || ||
1-8 9-28 29-38 39
일자 고객명(20) 금액(10) 여부(1)
A layout specification must have start position, length, type, alignment, and padding character.
| Item | Convention |
|---|---|
| Character field | Left-aligned, right-padded with spaces |
| Numeric field | Right-aligned, left-padded with 0 |
| Amount | As an integer with no decimal point, with the digits stated in the specification |
| Negative | Overlay the sign on the last digit, or a separate sign field |
| Date | 8 digits YYYYMMDD |
Trouble starts when Korean text is included. Depending on whether "length 20" is 20 bytes or 20 characters, you get entirely different files. In EUC-KR a Korean character is 2 bytes, and in UTF-8 it is 3. And with fixed length, if even one byte shifts, everything after it breaks.
That is why EUC-KR (or CP949) is still alive in domestic external integration. Korean characters fit neatly into 2 bytes, which makes digit calculation easy. If a received file looks garbled, suspect the encoding first and try converting it.
iconv -f EUC-KR -t UTF-8 SALES_20260801.dat > SALES_utf8.dat
Header and trailer — the file's own self-verification
A well-designed file interface has a header and a trailer.
H20260801SALES 0001 ← 헤더: 구분, 일자, 업무, 파일순번
D20260801HONG 0000012500 ← 데이터
D20260801KIM 0000030000
T0000000002000000042500 ← 트레일러: 건수, 금액 합계
The trailer exists for one reason — to prove by itself that the file arrived intact.
- Is the data count == the trailer's count?
- Is the data amount total == the trailer's total?
If it was cut during transfer, a line was lost in the middle, or it was corrupted during encoding conversion, this check catches it. If you do not do this check, you process a half-arrived file as normal. And you discover that at month-end closing as "the numbers do not match."
Trailer verification must be done before processing. If you do it during processing, half is already reflected.
Completion flag — the problem of picking up before writing is done
What happens if a batch picks up a file while it is being uploaded over FTP? It processes a half file. If there is trailer verification it gets caught, but if not, it goes straight in.
The standard convention is a completion flag file.
SALES_20260801.dat ← 데이터 (전송 중일 수 있음)
SALES_20260801.dat.ok ← 이 파일이 생겨야 처리 대상
After the sender has sent all the data, it creates an empty .ok file.
The receiver processes only those that have a .ok. It is simple but very effective.
It must be stated in the specification. If the "flag file convention" exists on only one side, the other side just uploads the data, and our batch processes nothing, forever. And no error occurs at all. Silently doing nothing is discovered the latest.
Filename rules are an interface too
<업무코드>_<기준일자>_<순번>.<확장자>
SALES_20260801_001.dat
- The reference date must be in the file name so that you can identify the file to reprocess
- A sequence number is needed to distinguish cases where several arrive in a day
- Without a filename rule,
sales.txtgets overwritten every day, and you cannot find yesterday's
Retention after processing
Do not delete processed files; retain them.
/data/if/archive/20260801/SALES_20260801_001.dat.gz
- Split them into per-date directories so hundreds of thousands of files do not pile up in one directory
- Compress them. Text files usually shrink by 80–90%
- State the retention period in the specification. Unlimited retention comes back as a disk-full outage
The reason for retention is reprocessing and disputes. To answer "what we sent that day was 12,000 records," you need the original.
Reconciliation — the last step of file integration
Processing the file is not the end. You must check whether the numbers match on both sides.
| Reconciliation item | Method |
|---|---|
| Count | Sent count == received processed count + error count |
| Amount total | Sent total == received total |
| Key set | Sent key list - received key list = empty set |
It is easy to feel reassured if only the count and total match, but if one record is missing and another is inserted twice, the count still matches. So look at the amount total too, and for interfaces where accuracy matters, compare the key set as well.
Leave the reconciliation result as a file. And if there is a mismatch, make it notify the owner automatically. A method where a person opens the reconciliation result every day gets skipped on a busy day, and that skipped day is the day with the problem.
What it looks like in the field
The incidents that occur in file integration are of set kinds.
- We processed a half file. The other side was still writing and we picked it up. That is why you also use a completion flag (
.ok) and pick up only files with a flag. Without this rule, it becomes an irreproducible bug: "on some days fewer records come in." - The encoding changed. A file that used to come in EUC-KR comes in UTF-8 one day. With fixed length, the byte length of Korean text changes, so all the fields after it shift. On screen the names are garbled and the amounts go in as wrong values.
- We did not look at the trailer. The file itself carries the count and total, but if you do not match against them, you load a file cut off during transfer as is and nobody knows.
All three arise from having no procedure to check that the file can be trusted before reading it. So the first step of file integration is not parsing but verification.