The Nightly Batch Loses a Few Calls to 401 Every Run
Goal
You build a client that keeps calling someone else's API while holding a short-lived access token. It refreshes ahead of expiry, leaves a clock skew margin, accepts a 401 only once and retries after refreshing, and makes the refresh happen only once even with many workers.
Why it matters
The cause of "the nightly batch fails with 401 on two or three records every time" is almost always the same. A structure that obtains a token once and gets a new one when a 401 arrives necessarily fails once at every expiry moment. A token living briefly is not a defect but a design. The lifetime must be short so that the window of use when it leaks is short. So the side that must be fixed is not the server but our client. And when fixing, you must look at three things together. Clocks are not exactly the same, so you need a margin; a 401 may not be a token problem, so you need a limit on the retry count; and with several workers the refreshes pile up at the expiry moment, so you must bundle the refresh into one. The grader does not believe your sentences. It starts the authentication server directly on a port the grader chooses, counts the issuances and resource calls on the server side, and checks how many times your client actually refreshed and how many times it called again.
Steps
- Create /root/token/authsrv.py, run it on port 8015, obtain one token, and save it to /root/token/token.json.
- Create /root/token/decode.py so that it reads the header and payload inside the token and writes them to /root/token/claims.json.
- Create /root/token/should_refresh.py so that it tells apart four branches, expired, not yet valid, near expiry, and normal, taking the clock skew margin into account.
- Create /root/token/client.py so that it refreshes ahead of expiry while calling several times and not a single 401 occurs.
- Make client.py, when it receives a 401, refresh and call again exactly once.
- Make client.py ensure that even if several workers start at the same time, the token issuance happens only once.
- Write a policy table in /root/token/token_policy.json.
- Report in four sections in /root/token/token_report.md.
Notes
- Authentication server run contract:
python3 /root/token/authsrv.py --port <포트> [--ttl <초>] [--always-401](the placeholders are the port and seconds).POST /oauth/tokenreturns{"access_token": <JWT>, "token_type": "Bearer", "expires_in": <초>}, andGET /api/datareturns 200 ifAuthorization: Bearer <토큰>is valid (the placeholder is the token), otherwise 401{"error": "invalid_token", "reason": ...}.GET /statsis{"issued": n, "api_requests": n}. The JWT is HS256 and the claims are iss, sub, iat, nbf, exp, and jti. - Reader run contract:
python3 decode.py --in <토큰 응답 JSON> --out <claims JSON>(the placeholders are the token response JSON and the claims JSON) outputs{"header": {...}, "payload": {...}, "ttl_s": exp - iat}. base64url goes around with its padding stripped, so make the length a multiple of 4 before decoding. What you do here is reading, not verification. - Judge run contract:
python3 should_refresh.py --exp <epoch> --now <epoch> --skew-s <초> [--nbf <epoch>](the placeholders are epoch values and seconds) outputs{"valid": ..., "refresh": ..., "remaining_s": ..., "reason": "ok"|"near_expiry"|"expired"|"not_yet_valid"}. The judging order is nbf first, then expiry, then near expiry.remaining_sisexp - nowand is negative after expiry. A token that is near expiry (near_expiry) can still be used. - Client run contract:
python3 client.py --base <URL> --calls <n> [--interval-ms <m>] [--skew-s <s>] [--workers <w>] [--cache <파일>] [--start-expired](the placeholder is the cache file) outputs{"calls": n, "ok": k, "unauthorized": u, "refreshes": r, "retried_after_401": x}. If--workersis greater than 1, calls are made concurrently,--cacheis the file that holds the token, and--start-expiredstarts after planting an expired token in the cache. - The places to prevent a refresh stampede are the cache and the lock. If you do not check the cache once more after taking the lock, the lock only lines them up in order and the number of refreshes does not decrease.
- Policy table format:
{"token_ttl_s": ..., "refresh_skew_s": ..., "max_retry_on_401": ..., "cache_path": ..., "stampede_guard": ..., "clock_sync": ..., "notes": ...}.refresh_skew_smust be 1 or more and less thantoken_ttl_s, andmax_retry_on_401is 1 or less. - Common mistakes: refreshing only after a 401 arrives, not putting a limit on the retry count (a loop forms with a server that answers a permission problem with a 401), putting on only the lock and leaving out the second check, and leaving the token in logs as is.
- Run the server in the background, wait until
/healthis 200, and then move on. The grader does not look at the process you left running but restarts the scripts directly.
Get a short-lived token
Create /root/token/authsrv.py, run it on port 8015, obtain one token with POST /oauth/token, and save the whole response to /root/token/token.json.
A JWT is the base64url-encoded header and payload and an HMAC signature joined by dots. You can build it with only Python standard library hmac, hashlib, and base64. If the server counts issuances and resource calls, you can later check how many times the client actually refreshed.
Open up a token
Create /root/token/decode.py so that it takes the JWT out of the token response received with --in, reads the header and payload, and, with --out, writes {"header": ..., "payload": ..., "ttl_s": exp - iat} to /root/token/claims.json.
The first two pieces of a JWT are not encrypted but plaintext encoded in base64url. base64url goes around with the padding (=) stripped, so you must make the length a multiple of 4 before decoding. Remember that what you do here is reading, not verification — you must not judge permissions on the basis of this value.
When should it be refreshed
Create /root/token/should_refresh.py so that it judges with --exp, --now, --skew-s, and --nbf. The judging order is nbf, expiry, near expiry, and the reason is not_yet_valid, expired, near_expiry, or ok. A token near expiry can still be used, so valid is true.
Clocks are not exactly the same. For nbf you must allow margin on our side so that a token just received is not rejected as "not yet valid." The near-expiry judgment needs just one line, 남은 시간 <= 여유 (remaining time <= margin). remaining_s must become negative after expiry.
Refresh ahead of expiry
Create /root/token/client.py so that it calls the resource --calls times, each time checking whether the token is near expiry and, if so, refreshing first. Even when calling for longer than the token lifetime, unauthorized must be 0.
If you import the previous step's judge as a module, the rule lives in only one place. Keep the token in a file and ask "may I use this?" just before each call. If you call for more than 3 seconds with a 3-second-lifetime token, you can check whether refreshes actually happened from the server's issuance count.
Accept a 401 only once
Make client.py, when it receives a 401, refresh and call again exactly once. If it is still 401 even after refreshing, that call must end as a failure. You must be able to start holding an expired token with --start-expired.
Many servers give a 401 in a situation where permissions are lacking. With such a server, if there is no limit on the retry count, refreshing and 401 repeat endlessly and pound the authentication server. Check from the server's call count that the resource call goes out exactly twice per call.
Twenty workers refresh at the expiry moment all at once
Make client.py ensure that even when it makes several calls concurrently with --workers, the token issuance happens only once. Use a cache file and a lock, and you must check the cache once more after taking the lock.
With only a lock, the refreshes just happen in line and the count stays the same. While waiting, someone else may have already refreshed, so re-read the cache after taking the lock and, if there is a usable token, use that. In Python, use fcntl file locks.
Pin it down with a policy table
Write token_ttl_s, refresh_skew_s, max_retry_on_401, cache_path, stampede_guard, clock_sync, and notes in /root/token/token_policy.json. refresh_skew_s must be 1 or more and less than the lifetime, and max_retry_on_401 is 1 or less.
Leave the basis of the values in notes in one line. It is enough to include what percentage of the lifetime you set the margin at and why you capped the retry at 1. This table is also a document that tells us what we must change together when the partner changes the lifetime later.
Token inspection report
In /root/token/token_report.md, write four sections, ## 토큰이 어떻게 생겼나, ## 언제 갈아야 하나, ## 401 을 받으면 무엇을 하나, and ## 워커가 여럿일 때 (in order: what the token looks like, when to refresh, what to do on a 401, when there are several workers). The values from claims.json and token_policy.json must be in the body.
The reader is someone asking why the nightly batch dies with 401. Do not stop at "because the token expired"; show with numbers that expiry is normal and that the problem was the way we handled expiry.