Batch or Streaming — What Decides
One-line summary
Batch processes bounded chunks of data periodically, and streaming keeps processing events that arrive endlessly. The real difference between them is not speed but who decides the boundary.
Why this was needed
The decision "real time is better, so let's use streaming" is made often, and its price usually gets billed six months later. With batch, you just rerun it when it fails, but streaming is a system that keeps running while holding state, so designing reprocessing is much harder.
How it works
The core device of batch is the watermark. You store how far you have processed and, on the next run, fetch only what comes after that.
SELECT * FROM orders
WHERE ordered_at > (SELECT last_ordered_at FROM etl_watermark WHERE job_name = 'orders_archive')
AND ordered_at < :batch_end;
This simple pattern hides several traps.
- Whether the boundary value is included. If you mix up
>and>=, one row is duplicated or missed on every run. - Identical timestamps. If several rows come in at the same instant,
>alone causes some of them to be missed forever. It is safer to compare the timestamp and the primary key together, or to keep the watermark pushed back a little and absorb duplicates with an upsert. - Late-arriving data. If the source commits its transaction late, a row with a timestamp earlier than the watermark appears afterward. In this case a time-based watermark misses that row forever.
In streaming this problem is even more blatant. Because the event time and the processing time differ, you must judge "is it OK to close this window now?", and so the watermark is defined together with an allowed lateness. And you must decide by policy whether to drop late-arriving events or reopen the window.
Delivery guarantees also split into three kinds. At most once (may lose), at least once (may duplicate), and exactly once. Most real systems provide at least once and take the approach of absorbing duplicates on the consuming side so that the result is effectively once. That is why idempotency in the next module becomes important.
What to base the choice on
"Real time is better" is not a criterion. If you weigh four things, the answer is usually decided.
| Question | Batch fits | Streaming fits |
|---|---|---|
| When is the result needed | In units of hours or days | In units of seconds or minutes |
| What to do with late data | Naturally included in the next batch | Needs watermark and reprocessing design |
| Can you recompute everything | Easy | Hard (state restoration) |
| Operational burden | If it fails, just rerun | It must be up all the time |
Whether you can undo it is the most important. With batch, you fix the logic and rerun yesterday's, and you are done. To do the same in streaming, you have to rewind the offsets, reset the state, and handle duplicates downstream.
So the common answer in practice is both. You produce a quick approximation with streaming and overwrite it later with the accurate value from batch (the lambda architecture). These days the approach of handling both with the same code (kappa, Flink, Beam) has grown, but the operational complexity is still greater on the streaming side.
Three things about time
Half of what goes wrong in streaming is the definition of time.
- Event time — when it actually happened. The basis for analysis is always this.
- Ingestion time — when it entered the system.
- Processing time — when it was computed. It changes when you reprocess, so it must not be used as the basis.
A mobile app goes into airplane mode and sends events hours later. If you aggregate by event time, late data arrives into a window that has already been closed. The watermark is the declaration "I will consider that no more data from before this time will come," and what arrives beyond that line is dropped or handled through a separate path.
워터마크 = 지금까지 본 최대 이벤트 시각 − 허용 지연(예: 10분)
Increasing the allowed lateness makes the result more accurate, but the result comes out that much later. This is the one knob that trades accuracy for latency.
Batch runs incrementally too
Even in batch, you do not need to read everything every time. You record the last point you processed and read only what comes after. But you set the boundary so that it overlaps.
-- 워터마크를 그대로 쓰면 경계에 걸친 것을 놓친다
where updated_at >= :last_watermark - interval '10 minutes'
and updated_at < :now
The overlapping interval is read again, but if the load is idempotent, it does no harm. The principle that with idempotency, overlapping reads are free is the same here.
What it looks like in the field
There are more situations where batch is better than streaming than you might think. If the source is updated once a day, a real-time-only pipeline has no value. If the person looking at the report looks once in the morning, an early-morning batch is enough. How often the consumer actually looks should be the first question.
Conversely, the price of batch is also clear. The longer the period, the larger the delay when it fails once, and because the amount processed at once is large, resource usage becomes spiky. So splitting a large batch into small chunks and processing them is commonly used. You can rerun only the failed chunk, and the lock time becomes shorter.
What to do in the next lab
You extract order data to a CSV, load it into an archive table, record a watermark and run an incremental load, and then confirm that the result does not change even if you run the same job twice.