TT Lab
Get started
Learn Learning paths Courses

Data Pipelines

Batch or Streaming — What Decides

Continue in TT Lab

One-line summary

Batch processes bounded chunks of data periodically, and streaming keeps processing events that arrive endlessly. The real difference between them is not speed but who decides the boundary.

Why this was needed

The decision "real time is better, so let's use streaming" is made often, and its price usually gets billed six months later. With batch, you just rerun it when it fails, but streaming is a system that keeps running while holding state, so designing reprocessing is much harder.

How it works

The core device of batch is the watermark. You store how far you have processed and, on the next run, fetch only what comes after that.

SELECT * FROM orders
WHERE ordered_at > (SELECT last_ordered_at FROM etl_watermark WHERE job_name = 'orders_archive')
  AND ordered_at < :batch_end;

This simple pattern hides several traps.

In streaming this problem is even more blatant. Because the event time and the processing time differ, you must judge "is it OK to close this window now?", and so the watermark is defined together with an allowed lateness. And you must decide by policy whether to drop late-arriving events or reopen the window.

Delivery guarantees also split into three kinds. At most once (may lose), at least once (may duplicate), and exactly once. Most real systems provide at least once and take the approach of absorbing duplicates on the consuming side so that the result is effectively once. That is why idempotency in the next module becomes important.

What to base the choice on

"Real time is better" is not a criterion. If you weigh four things, the answer is usually decided.

Question Batch fits Streaming fits
When is the result needed In units of hours or days In units of seconds or minutes
What to do with late data Naturally included in the next batch Needs watermark and reprocessing design
Can you recompute everything Easy Hard (state restoration)
Operational burden If it fails, just rerun It must be up all the time

Whether you can undo it is the most important. With batch, you fix the logic and rerun yesterday's, and you are done. To do the same in streaming, you have to rewind the offsets, reset the state, and handle duplicates downstream.

So the common answer in practice is both. You produce a quick approximation with streaming and overwrite it later with the accurate value from batch (the lambda architecture). These days the approach of handling both with the same code (kappa, Flink, Beam) has grown, but the operational complexity is still greater on the streaming side.

Three things about time

Half of what goes wrong in streaming is the definition of time.

A mobile app goes into airplane mode and sends events hours later. If you aggregate by event time, late data arrives into a window that has already been closed. The watermark is the declaration "I will consider that no more data from before this time will come," and what arrives beyond that line is dropped or handled through a separate path.

워터마크 = 지금까지 본 최대 이벤트 시각 − 허용 지연(예: 10분)

Increasing the allowed lateness makes the result more accurate, but the result comes out that much later. This is the one knob that trades accuracy for latency.

Batch runs incrementally too

Even in batch, you do not need to read everything every time. You record the last point you processed and read only what comes after. But you set the boundary so that it overlaps.

-- 워터마크를 그대로 쓰면 경계에 걸친 것을 놓친다
where updated_at >= :last_watermark - interval '10 minutes'
  and updated_at <  :now

The overlapping interval is read again, but if the load is idempotent, it does no harm. The principle that with idempotency, overlapping reads are free is the same here.

What it looks like in the field

There are more situations where batch is better than streaming than you might think. If the source is updated once a day, a real-time-only pipeline has no value. If the person looking at the report looks once in the morning, an early-morning batch is enough. How often the consumer actually looks should be the first question.

Conversely, the price of batch is also clear. The longer the period, the larger the delay when it fails once, and because the amount processed at once is large, resource usage becomes spiky. So splitting a large batch into small chunks and processing them is commonly used. You can rerun only the failed chunk, and the lock time becomes shorter.

What to do in the next lab

You extract order data to a CSV, load it into an archive table, record a watermark and run an incremental load, and then confirm that the result does not change even if you run the same job twice.