The Logs Came From the Future — Five Incidents a Clock Made
The settlement that should have run once ran twice
One-line summary
For recurring jobs, you cannot judge normal operation by "how many times did it run." One skipped run and one run that happened twice hide each other, and the total comes out right.
Why this is needed
Monitoring of periodic jobs usually looks at just two things: did it fail, and how many times did it run. The accident that these two cannot catch is the last piece of this course.
In the early morning of the day daylight saving time ends, the local-time 01:00 settlement runs twice. Both succeed, so there is no failure alert. In the early morning of the day daylight saving time starts, the 02:00 settlement never runs at all. Since nothing happened, there is again no alert. And if you add up the run counts of these two days, they come out exactly equal to the expectation. Monitoring that counts totals structurally cannot catch this accident.
How it works
The root of the problem is how you name the period a job handles. If you name it in local time like "the 01:00 settlement," the same name appears twice on a day when local time occurs twice, and never appears on a day when it does not exist. If you name it by an instant (UTC), each name appears exactly once.
The scheduler side makes the same choice. Kubernetes
CronJob
lets you specify, with .spec.timeZone, in which region's time the schedule is read (stable
since v1.27), and if you do not specify it, it is read in the controller manager's local time. If you explicitly give Etc/UTC,
daylight saving time no longer matters. As a side note, putting TZ or CRON_TZ inside the schedule string is
not officially supported, and if you do, creating the resource is rejected with a validation error.
But even if you pin the time zone to UTC, "exactly once" is not guaranteed. The same document says this — a CronJob creates a Job approximately once per schedule, there are cases where two are created or none is created, and Kubernetes tries to avoid that but cannot prevent it completely. So it insists that the Jobs you define must be idempotent. This is also the last sentence of this course.
Overlap is handled with a separate knob. .spec.concurrencyPolicy takes three values.
Allow (기본) 앞 실행이 안 끝났어도 새 실행을 만든다
Forbid 앞 실행이 안 끝났으면 이번 실행을 건너뛴다
Replace 앞 실행을 새 실행으로 갈아 끼운다
And .spec.startingDeadlineSeconds sets how late after the scheduled time a run is still allowed to start.
A run that cannot start within this value is skipped, and Kubernetes treats it as a
failed Job. The documentation warns separately about two things — if you set this value below 10 seconds,
the job may not be scheduled at all because the controller checks every 10 seconds, and if more than 100 schedules
have been missed, the controller gives up starting.
The Linux systemd timer
has the same distinction. There is OnCalendar, which follows the wall clock, and the OnUnitActiveSec
family, which counts elapsed time, and for the former, if you change the clock, the next firing time
is recomputed.
What you see in the field
The way to actually make things idempotent is usually a run key. You build a string that uniquely identifies what that run handles, and if there is already a record of processing under that key, you do not write again. The criterion for choosing the key is one thing — doing the same work twice must produce the same key, and different work must produce a different key. If you put a value that changes on every run, such as a run ID or a start time, into the key, both of the two runs are seen as new work. In practice this mistake is the most common.
It is also often mixed up that preventing overlap and idempotency are different problems. Overlap prevention stops two runs at the same time, and idempotency makes the second run harmless even when they are far apart in time. A daylight saving time duplicate is an hour apart, so overlap prevention does not catch it. Conversely, for an overlap caused by a previous run that got long, idempotency alone does not prevent resource contention. You need both.
Finally, you have to change the shape of the monitoring. Instead of the run count, compare the list of expected periods with the list of periods actually processed. If missing periods and duplicated periods each show up, totals no longer hide each other.
This monitoring is worth more because it also catches accidents unrelated to clocks. A node died and one run was skipped, the scheduler was stopped during a deployment, a previous run got too long and the next one was skipped — the causes differ, but the symptom is always the same: "that period was not processed." Monitoring that matches period lists does not ask about the cause and looks only at that one symptom. It covers more accidents with far fewer rules than building a separate alert for each cause.
And the whole course can be compressed into one sentence. A record with a time written on it is not a fact but a claim of the machine that wrote it, and code that judges deadlines trusts that claim and decides someone else's fate. So code that handles time must make two things clear: what it stores and what it merely displays, and which calculations rest on "what time is it now" and which rest on "how much time has passed." Just keeping these two distinctions prevents four of the five incidents seen in this course.
What you will do in the next lab
In the lab you read in turn a bundle of logs stamped with off clocks, send records kept only in local time, timer records that keep the wall clock and the monotonic clock together, a certificate containing a validity period, and the run history of a periodic settlement. You measure and write down the size of the error, the overlapping interval, and the skipped runs yourself, and write three functions that handle time correctly. The grader recomputes those numbers from the materials and compares them, and calls the functions you wrote directly to check their properties.