Tokens Are Supposed to Expire
Summary
It is normal for an access token to live briefly and die, so a client must refresh it ahead of expiry, leave margin for clock skew, refresh and retry exactly once when it receives a 401, and make sure that even with many workers the refresh happens only once.
Why this was needed
If you follow a report like "the nightly batch fails with 401 on two or three records every time," the cause is almost always the same. A structure that obtains a token once and, when a 401 arrives, gets a new one then. This structure necessarily fails once at every expiry moment. If that failure is a user request, the user sees it.
A token living briefly is not a defect but a design. RFC 6749 defines a structure where the access token's lifetime is short and it is obtained again with a refresh token, and RFC 6750 specifies how to carry that token in Authorization: Bearer and the WWW-Authenticate response on failure. The lifetime must be short so that the window in which a leaked token can be used is short — that is the purpose.
How it works
Reading inside a token and trusting it are different. A JWT (RFC 7519) is three pieces separated by dots, and the first two pieces are not encrypted but plaintext encoded in base64url. Anyone can open them and anyone can change them. So it is fine for the client to read exp to decide "when to refresh," but it must not judge permissions on the basis of that value. Signature verification is the resource server's job.
There are three time-related claims. iat (issued-at time), nbf (do not use before this time), and exp (do not use after this time). RFC 7519 says that when handling exp and nbf, you may allow a margin (leeway) that accounts for small clock skew. It is common for the clocks of the server and client to differ by a few seconds, and without margin, those few seconds cause a token just received to be rejected as "not yet valid."
Refresh before expiry. The rule is one line. If 남은 시간 <= 여유 (remaining time <= margin), refresh now. A large margin makes refreshes frequent, and a small one raises the risk of being caught at the expiry moment. Between 10% and 20% of the lifetime is a common starting point.
Accept a 401 only once. A 401 can come even if you refresh ahead — when the server has revoked the token early, permissions have changed, or the clock has drifted greatly. In that case, refresh and retry exactly once. If you put no limit on the count here, a loop forms. When the problem is permissions and not the token (many servers answer with a 401 in a situation that should really be a 403), it stays 401 even after refreshing, and the client keeps obtaining tokens endlessly and pounds the authentication server.
Prevent a refresh stampede. If 20 workers share the same token, at the expiry moment all 20 simultaneously decide "we have to refresh." From the authentication server's point of view, 20 times the usual requests arrive at once, and when that server slows down, refreshes take longer and more workers pile in. The solution is to make the refresh happen only once.
토큰 필요
│
├─ 캐시를 본다 ── 아직 쓸 만한가? ── 예 ─▶ 그대로 쓴다 (대부분 여기서 끝난다)
│ 아니오
└─ 잠금을 잡는다 ──▶ 캐시를 **다시** 본다 ── 남이 갈아 뒀는가? ── 예 ─▶ 그것을 쓴다
아니오 ─▶ 내가 갱신한다
The key is checking once more after taking the lock. Because while waiting, another worker may already have refreshed. If you leave out this one line, the lock only lines them up in order and the number of refreshes stays the same.
What it looks like in the field
First, many servers mix up 401 and 403. If a server gives a 401 in a situation where permissions are lacking, the client concludes "the token is the problem" and repeats refreshing. On our side, capping the retry count at 1 is the only defense.
Second, keeping the token only in process memory. If workers are scattered per process, they cannot share the cache and issuance happens as many times as there are processes. For a partner with a limit on issuance counts, that alone becomes an outage.
Third, not syncing the clock. If a container's clock is off by 30 seconds, however well you set the margin, it is pushed toward the drifted side. If a 401 occurs only on a particular node, look at that node's clock first.
Fourth, the token is left in logs. An access token is itself a credential. If a debug log that prints whole request headers remains, everyone who can read that log can call that API.
What you will do in the next lab
You start an authentication server that issues a short-lived token, obtain one token, and open it to read iat, nbf, and exp. You build a judge that tells apart four branches, expired, not yet valid, near expiry, and normal, taking the clock skew margin into account, and with that judgment build a client that refreshes ahead of expiry while calling continuously. Next you check, against a server that gives a 401 whatever you bring, that the retry stops at one, and make it so that even if six workers start at the same time, issuance happens only once. Finally you write these values as a policy table.