TT Lab
Get started
Learn Learning paths Courses

The Logs Came From the Future — Five Incidents a Clock Made

A perfectly good certificate was rejected

Continue in TT Lab

One-line summary

Certificates, tokens, and caches all check "what time is it now" against someone else. If my clock is wrong, a perfectly good thing is rejected and an expired thing passes.

Why this is needed

Most clock errors cause nothing. Logs being off by a few seconds is merely inconvenient to read. But at the point where a deadline is judged, the same error immediately becomes a rejection. And that rejection does not point to the cause. The message says "the certificate is not yet valid," and the certificate is a perfectly normal one that was just received.

What especially confuses people is that it heals itself. A machine whose clock is 12 seconds behind fails only for 12 seconds after getting a new certificate and is fine after that. A few errors appear only right after a deployment and cannot be reproduced. Such things are usually filed away as "a transient network problem" and repeat as they are on the next deployment.

How it works

An X.509 certificate contains its validity period as two timestamps: notBefore and notAfter. RFC 5280, section 4.1.2.5 defines this period as "from notBefore through notAfter, inclusive." The verifying side uses its own clock to see whether now falls inside that period. So the same certificate passes on some machines and is rejected on others.

openssl x509 -in edge-1.pem -noout -startdate -enddate
# notBefore=Nov  1 09:38:00 2025 GMT
# notAfter=Feb  1 09:38:00 2026 GMT

A machine whose clock is behind gets caught at the front end. The notBefore of a newly issued certificate is in the future relative to that machine's "now," so it decides the certificate is not yet valid. The time it is caught is exactly the amount it is behind. Conversely, a machine whose clock is ahead gets caught at the back end. It sees expiry earlier than others. The two symptoms are in opposite directions, and knowing which one tells you immediately which way the clock is wrong.

Tokens have the same structure. The exp claim of RFC 7519 says it must not be accepted "on or after" that time, and nbf says it must not be accepted before that time. And both sections add the same sentence — implementers may provide some leeway to account for clock skew, and it is usually no more than a few minutes.

This is where the relationship between lifetime and error shows. If a machine whose clock is 277 seconds ahead verifies a token with a 300-second lifetime, it looks expired from 23 seconds after issuance. As the error approaches the lifetime, the usable time converges to 0. So the shorter the lifetime, the harsher the clock accuracy requirement. If you decide to use 1-minute tokens, you must first set a clock error budget.

쓸 수 있는 시간 = 수명 - (검증하는 쪽 시계가 앞선 만큼)
  수명 300초 · 오차 +277초  →   23초
  수명 300초 · 오차 +  5초  →  295초
  수명  60초 · 오차 +277초  →  아예 못 쓴다

Caches sit on the same axis. If you give the expiry time in a response as an absolute time, it depends on the receiver's clock, and if you give it as remaining seconds, the receiver counts with its own monotonic clock and is not affected by the error. This is why HTTP has Cache-Control: max-age separately from the Expires header and gives precedence to the former.

What you see in the field

In environments with automatic renewal, this problem interlocks with the deployment cycle. If you swap certificates early, with generous margin, instead of just before expiry, then even with a few minutes of clock error you are caught at neither end. This is where the practice of setting the renewal point at around two thirds of the lifetime comes from.

There is also a set order for investigating authentication failures. First see whether the message says "not yet valid" or "expired". If the former, the verifying side's clock is behind or the issuance just finished; if the latter, the verifying side's clock is ahead or it really has expired. Then compare the times on both machines directly. These two steps settle most incidents, yet in practice the hand reaches first for reissuing the certificate. And the same thing happens again at the next renewal.

You can also cover it up by setting a large leeway, but this has a cost. It means accepting expired tokens for that long. In a design that uses reissuance as revocation, revocation takes effect that much later. Keep the leeway as a temporary measure until the clock is fixed, and make the real fix monitoring the error.

One more thing. A failure caused by a deadline is not symmetric between the two ends. A front-end problem (not yet valid) appears briefly right after issuance and heals itself, while a back-end problem (early expiry), once it starts, continues until the clock is fixed or the certificate is renewed. So the former gets reported as "sporadic errors that can't be reproduced" and the latter as "suddenly everything died." The same cause, yet the report titles are opposite. If you think of this asymmetry first when you receive an incident, you can immediately decide which end to look at.

And the points where deadlines are judged are easy to find in code. Everywhere that reads "now" and compares it with a time given by someone else is such a point. Certificate verification, token verification, cache freshness, expiry of signed URLs, and timestamp windows that prevent replay attacks all have the same shape. Once you build a list of these points, how much clock error budget to set becomes a calculation, not a guess — the shortest lifetime on that list sets the upper limit of the budget.

What to check in the next quiz

Check which end of the validity period a behind-clock machine and an ahead-clock machine each get caught at, how token lifetime and clock error combine to reduce the usable time, and what the cost is of the fix that increases the leeway.