TT Lab
Get started
Learn Learning paths Courses

Integration and Deployment

The Nightly Batch Loses a Few Calls to 401 Every Run

Continue in TT Lab

Goal

You build a client that keeps calling someone else's API while holding a short-lived access token. It refreshes ahead of expiry, leaves a clock skew margin, accepts a 401 only once and retries after refreshing, and makes the refresh happen only once even with many workers.

Why it matters

The cause of "the nightly batch fails with 401 on two or three records every time" is almost always the same. A structure that obtains a token once and gets a new one when a 401 arrives necessarily fails once at every expiry moment. A token living briefly is not a defect but a design. The lifetime must be short so that the window of use when it leaks is short. So the side that must be fixed is not the server but our client. And when fixing, you must look at three things together. Clocks are not exactly the same, so you need a margin; a 401 may not be a token problem, so you need a limit on the retry count; and with several workers the refreshes pile up at the expiry moment, so you must bundle the refresh into one. The grader does not believe your sentences. It starts the authentication server directly on a port the grader chooses, counts the issuances and resource calls on the server side, and checks how many times your client actually refreshed and how many times it called again.

Steps

  1. Create /root/token/authsrv.py, run it on port 8015, obtain one token, and save it to /root/token/token.json.
  2. Create /root/token/decode.py so that it reads the header and payload inside the token and writes them to /root/token/claims.json.
  3. Create /root/token/should_refresh.py so that it tells apart four branches, expired, not yet valid, near expiry, and normal, taking the clock skew margin into account.
  4. Create /root/token/client.py so that it refreshes ahead of expiry while calling several times and not a single 401 occurs.
  5. Make client.py, when it receives a 401, refresh and call again exactly once.
  6. Make client.py ensure that even if several workers start at the same time, the token issuance happens only once.
  7. Write a policy table in /root/token/token_policy.json.
  8. Report in four sections in /root/token/token_report.md.

Notes

Get a short-lived token

Create /root/token/authsrv.py, run it on port 8015, obtain one token with POST /oauth/token, and save the whole response to /root/token/token.json.

A JWT is the base64url-encoded header and payload and an HMAC signature joined by dots. You can build it with only Python standard library hmac, hashlib, and base64. If the server counts issuances and resource calls, you can later check how many times the client actually refreshed.

Open up a token

Create /root/token/decode.py so that it takes the JWT out of the token response received with --in, reads the header and payload, and, with --out, writes {"header": ..., "payload": ..., "ttl_s": exp - iat} to /root/token/claims.json.

The first two pieces of a JWT are not encrypted but plaintext encoded in base64url. base64url goes around with the padding (=) stripped, so you must make the length a multiple of 4 before decoding. Remember that what you do here is reading, not verification — you must not judge permissions on the basis of this value.

When should it be refreshed

Create /root/token/should_refresh.py so that it judges with --exp, --now, --skew-s, and --nbf. The judging order is nbf, expiry, near expiry, and the reason is not_yet_valid, expired, near_expiry, or ok. A token near expiry can still be used, so valid is true.

Clocks are not exactly the same. For nbf you must allow margin on our side so that a token just received is not rejected as "not yet valid." The near-expiry judgment needs just one line, 남은 시간 <= 여유 (remaining time <= margin). remaining_s must become negative after expiry.

Refresh ahead of expiry

Create /root/token/client.py so that it calls the resource --calls times, each time checking whether the token is near expiry and, if so, refreshing first. Even when calling for longer than the token lifetime, unauthorized must be 0.

If you import the previous step's judge as a module, the rule lives in only one place. Keep the token in a file and ask "may I use this?" just before each call. If you call for more than 3 seconds with a 3-second-lifetime token, you can check whether refreshes actually happened from the server's issuance count.

Accept a 401 only once

Make client.py, when it receives a 401, refresh and call again exactly once. If it is still 401 even after refreshing, that call must end as a failure. You must be able to start holding an expired token with --start-expired.

Many servers give a 401 in a situation where permissions are lacking. With such a server, if there is no limit on the retry count, refreshing and 401 repeat endlessly and pound the authentication server. Check from the server's call count that the resource call goes out exactly twice per call.

Twenty workers refresh at the expiry moment all at once

Make client.py ensure that even when it makes several calls concurrently with --workers, the token issuance happens only once. Use a cache file and a lock, and you must check the cache once more after taking the lock.

With only a lock, the refreshes just happen in line and the count stays the same. While waiting, someone else may have already refreshed, so re-read the cache after taking the lock and, if there is a usable token, use that. In Python, use fcntl file locks.

Pin it down with a policy table

Write token_ttl_s, refresh_skew_s, max_retry_on_401, cache_path, stampede_guard, clock_sync, and notes in /root/token/token_policy.json. refresh_skew_s must be 1 or more and less than the lifetime, and max_retry_on_401 is 1 or less.

Leave the basis of the values in notes in one line. It is enough to include what percentage of the lifetime you set the margin at and why you capped the retry at 1. This table is also a document that tells us what we must change together when the partner changes the lifetime later.

Token inspection report

In /root/token/token_report.md, write four sections, ## 토큰이 어떻게 생겼나, ## 언제 갈아야 하나, ## 401 을 받으면 무엇을 하나, and ## 워커가 여럿일 때 (in order: what the token looks like, when to refresh, what to do on a 401, when there are several workers). The values from claims.json and token_policy.json must be in the body.

The reader is someone asking why the nightly batch dies with 401. Do not stop at "because the token expired"; show with numbers that expiry is normal and that the problem was the way we handled expiry.