TT Lab
Get started
Learn Learning paths Courses

System Integration (EAI)

What Actually Matters in Synchronous Integration

Continue in TT Lab

Summary

The difficulty of synchronous REST integration lies not in the call but in the waiting — if you do not set a timeout, the other system's outage becomes our own system's outage as is.

Why the real risk of synchronous integration is "waiting"

Calling the other system with REST is easy. What is hard is when the other side is slow.

Suppose our screen calls the other side's API synchronously and the other side drags on for 30 seconds.

So synchronous integration must have a timeout. There are two kinds.

연결 타임아웃(connect) : 상대와 TCP 연결이 맺어질 때까지  → 짧게 (1~3초)
응답 타임아웃(read)    : 응답을 다 받을 때까지            → 업무에 따라 (5~30초)

Setting a short connect timeout is important. If the other server is dead, the connection should fail immediately. If you set this to 30 seconds, while the other side is dead our threads are tied up 30 seconds at a time.

And timeouts alone are not enough. If the other side stays slow, we stay tied up too. What you need then is a circuit breaker. When the failure rate exceeds a threshold, it stops calling altogether for a set time and fails immediately. There are situations where "failing fast" is better than "waiting for a slow success."

A timeout does not mean "there was no response"

This is the most important point.

When a timeout occurs, we cannot tell whether the other side never received the request or processed it but the response just did not arrive.

A timeout occurred while sending an order. Should we resend?

This problem cannot be solved by retrying. You solve it with idempotency. The sender attaches a unique key to every request (a message number, Idempotency-Key), and when the receiver gets the same key twice, it does not process it and returns the original response as is. Then resending becomes safe. This is covered as a lab in a later module.

Validate before the request — stop it before it goes out

There are things we must validate first, before sending to the other side.

If you do not do this, you find out from the other system's error response. Then

"We block bad data on our side" is the basic courtesy of integration development.

Keep integration logs separately

If you mix them into application logs, you cannot find them later. Create a dedicated integration log and leave one line per call. The fields should be at least these.

2026-08-19T14:03:22.145+0900|IF-ORD-001|ORDER-SYS|PARTNER-API|0000|182|a1b2c3d4
   시각            인터페이스ID  송신     수신      응답코드 소요ms  추적ID

And do not keep personal information in logs. Resident registration numbers, account numbers, and card numbers are masked or not kept at all. Integration logs are retained for a long time, so the risk is large.

Do not do bulk processing synchronously

For a requirement like "send 100,000 orders to the other side," you must not use synchronous REST. Even at just 100ms per item, 100,000 items come to about 2.8 hours. If it is cut even once in between, you do not know how far it got.

There are three alternatives.

  1. File batch — bundle and send at once, and receive the result as a file
  2. Queue — throw it asynchronously, with the result notified separately
  3. Bulk API — put N items (usually 100–1000) in one request and receive per-item results

What you must define when using option 3: how to handle partial failure. If 3 of 100 fail, is it a full rollback, or are 97 processed and only 3 returned? If it is not in the specification, the two sides implement it differently. And this difference is discovered days later in the form "the counts do not match."

What it looks like in the field

This failure always comes in the same shape. The other side's API slows down, our screen freezes, the user presses refresh, and that refresh creates one more request. The earlier request is still alive, so the threads are tied up twice over. A few minutes later our WAS runs out of threads, and even screens unrelated to the other side all stop.

At this point, monitoring shows our system's CPU and memory as normal. Because the threads are not working; they are just waiting. So you waste time looking for the cause on our side and not on the other side.

This is why you need to keep a separate integration log. If mixed into application logs, it takes a long time to extract "since when has the other side been slow." If you leave the interface ID, response code, and elapsed time one line at a time in a fixed format, that answer comes in 1 minute.