TT Lab
Get started
Learn Learning paths Courses

The Logs Came From the Future — Five Incidents a Clock Made

The elapsed time came out negative

Continue in TT Lab

One-line summary

"What time is it now" and "how much time has passed" must be asked of different clocks. If you try to answer both with one, elapsed time goes negative on the day the clock is set.

Why this is needed

Code that measures elapsed time usually looks like this. You note the current time at the start, and at the end you subtract it from the current time. This code almost always gives the right value. The only moment it is wrong is when someone adjusts the clock in between.

And clocks do get adjusted. When an operator finds a machine with a large error, they force a sync, and when a daemon meets a large error, it makes the time jump in one go. The same thing happens when a virtual machine is restored from a snapshot or wakes from sleep. A job that straddled that moment gets a negative elapsed time or, conversely, one inflated by several minutes.

A negative value is at least noticeable. The real problem is the code that uses that value. If you set a lock deadline as "now + 30 seconds" and the clock jumps forward 5 minutes, the deadline looks already passed, and someone else's lock is released. Conversely, if it jumps back 5 minutes, a 30-second lock is not released for five and a half minutes. In both cases what the lock was meant to protect breaks, and neither leaves the word "clock" in any log.

How it works

The operating system divided these two into different clocks from the start. The description in clock_gettime(2) states the difference precisely.

CLOCK_REALTIME    설정 가능한 시계. 1970-01-01 UTC 부터의 초를 센다.
                  관리자가 시각을 바꾸면 **불연속적으로 점프** 하고,
                  NTP 의 주파수 조정에도 영향을 받는다.
                  → "지금 몇 시인가" 에 답한다.

CLOCK_MONOTONIC   설정할 수 없는 시계. 리눅스에서는 부팅 이후의 초다.
                  시각을 바꿔도 **점프하지 않는다**(주파수 조정에는 영향을 받는다).
                  연속한 두 번의 호출이 뒤로 가는 일은 없다.
                  시스템이 잠들어 있던 시간은 세지 않는다.
                  → "얼마나 지났는가" 에 답한다.

CLOCK_BOOTTIME    CLOCK_MONOTONIC 과 같은데 절전 시간까지 포함한다.

The last line matters more than it seems. If you measure elapsed time on a laptop or a machine that uses sleep, CLOCK_MONOTONIC does not count the time spent asleep, so after three hours on the wall clock, the elapsed time can come out as 10 minutes. You must choose according to what you want to measure.

Languages expose the same distinction as is. Python's time module separates time.time() and time.monotonic(), and Go's time package even packs a monotonic clock reading into the value returned by time.Now() and uses that one when subtracting two Time values. The language handles the choice for you in the situations where you need to choose.

So the rule for code comes down to this. Measure times to display or store with the wall clock, and values to compute the interval between two points with the monotonic clock. When leaving a record, it is good to write both. The wall clock is read by people, and the monotonic clock is read by calculations.

The same principle applies when setting time limits. A database lock timeout, an HTTP request timeout, and the wait between retries are all "how much time has passed," so they must run on the monotonic clock. Conversely, certificate validity periods and token expiry are matters of matching "what time is it now" with others, so you have no choice but to use the wall clock. That story comes next.

What you see in the field

This problem shows up most expensively in distributed locks. The lock service answers "this lock expires in 30 seconds" as a wall-clock time, and the party holding the lock also checks that time against its wall clock. If the two machines' clocks differ, there is a window in which one side believes the lock is still valid and the other believes it has already expired. During that window two processes touch the same resource at once. That is why lock protocols usually exchange the remaining time rather than an absolute time.

The same thing happens in observability metrics. A service that measures response time by subtracting wall-clock times puts negative or absurdly large values into the histogram on the day the clock is set. Most metrics libraries either discard negatives or pile them into the last bucket, and either way that day's p99 becomes an untrustworthy value.

Finally, there is one signal you can spot in code review: code that does not even consider that a variable holding elapsed time could receive a negative value. A comparison like if elapsed > timeout quietly passes on a negative. On a day the clock goes backward, that code becomes code with no time limit.

What to check in the next quiz

Check which question each of the two clocks answers, what the monotonic clock is and is not affected by, and which one to use for locks and time limits.