What Actually Matters in Synchronous Integration
Summary
The difficulty of synchronous REST integration lies not in the call but in the waiting — if you do not set a timeout, the other system's outage becomes our own system's outage as is.
Why the real risk of synchronous integration is "waiting"
Calling the other system with REST is easy. What is hard is when the other side is slow.
Suppose our screen calls the other side's API synchronously and the other side drags on for 30 seconds.
- Our WAS thread is tied up for 30 seconds
- The user waits in front of the screen. Usually they press refresh
- A refresh creates one more request. The earlier request is still alive
- Threads are exhausted quickly. The other system's outage becomes our system's outage
So synchronous integration must have a timeout. There are two kinds.
연결 타임아웃(connect) : 상대와 TCP 연결이 맺어질 때까지 → 짧게 (1~3초)
응답 타임아웃(read) : 응답을 다 받을 때까지 → 업무에 따라 (5~30초)
Setting a short connect timeout is important. If the other server is dead, the connection should fail immediately. If you set this to 30 seconds, while the other side is dead our threads are tied up 30 seconds at a time.
And timeouts alone are not enough. If the other side stays slow, we stay tied up too. What you need then is a circuit breaker. When the failure rate exceeds a threshold, it stops calling altogether for a set time and fails immediately. There are situations where "failing fast" is better than "waiting for a slow success."
A timeout does not mean "there was no response"
This is the most important point.
When a timeout occurs, we cannot tell whether the other side never received the request or processed it but the response just did not arrive.
A timeout occurred while sending an order. Should we resend?
- If the other side did not receive it, we must resend
- If the other side processed it, a resend is a duplicate order
This problem cannot be solved by retrying. You solve it with idempotency.
The sender attaches a unique key to every request (a message number, Idempotency-Key),
and when the receiver gets the same key twice, it does not process it and returns the original response as is.
Then resending becomes safe. This is covered as a lab in a later module.
Validate before the request — stop it before it goes out
There are things we must validate first, before sending to the other side.
- Required values: are all the required items of the specification present
- Length: does it not exceed the limit, measured in bytes
- Type/format: date format, no letters in numeric fields
- Code values: is it within the code list stated in the specification
If you do not do this, you find out from the other system's error response. Then
- round-trip time is wasted,
- our errors pile up in the other system's logs (emotional drain between integration owners),
- and for a large batch, thousands of errors blow up at once
"We block bad data on our side" is the basic courtesy of integration development.
Keep integration logs separately
If you mix them into application logs, you cannot find them later. Create a dedicated integration log and leave one line per call. The fields should be at least these.
2026-08-19T14:03:22.145+0900|IF-ORD-001|ORDER-SYS|PARTNER-API|0000|182|a1b2c3d4
시각 인터페이스ID 송신 수신 응답코드 소요ms 추적ID
- The trace ID (trace_id) is especially important. When matching against the other system's logs, you find it with this one thing. Send it along in the request header, and write in the specification that the other side should keep it too.
- Keeping the elapsed time gives you a basis for the remark "it has been slow lately."
- If you keep the response code, you can aggregate the error distribution.
And do not keep personal information in logs. Resident registration numbers, account numbers, and card numbers are masked or not kept at all. Integration logs are retained for a long time, so the risk is large.
Do not do bulk processing synchronously
For a requirement like "send 100,000 orders to the other side," you must not use synchronous REST. Even at just 100ms per item, 100,000 items come to about 2.8 hours. If it is cut even once in between, you do not know how far it got.
There are three alternatives.
- File batch — bundle and send at once, and receive the result as a file
- Queue — throw it asynchronously, with the result notified separately
- Bulk API — put N items (usually 100–1000) in one request and receive per-item results
What you must define when using option 3: how to handle partial failure. If 3 of 100 fail, is it a full rollback, or are 97 processed and only 3 returned? If it is not in the specification, the two sides implement it differently. And this difference is discovered days later in the form "the counts do not match."
What it looks like in the field
This failure always comes in the same shape. The other side's API slows down, our screen freezes, the user presses refresh, and that refresh creates one more request. The earlier request is still alive, so the threads are tied up twice over. A few minutes later our WAS runs out of threads, and even screens unrelated to the other side all stop.
At this point, monitoring shows our system's CPU and memory as normal. Because the threads are not working; they are just waiting. So you waste time looking for the cause on our side and not on the other side.
This is why you need to keep a separate integration log. If mixed into application logs, it takes a long time to extract "since when has the other side been slow." If you leave the interface ID, response code, and elapsed time one line at a time in a fixed format, that answer comes in 1 minute.