401 Means I Don't Know Who You Are; 403 Means I Do and Still No
Summary
The cheapest tool for narrowing an authentication problem is two status codes, and the cheapest tool for narrowing a certificate problem is how many certificates the server actually sends.
Why this was needed
An FDE is someone who handles other people's keys. And authentication problems are among the trickiest of outage reports to reproduce. Because they fail differently per token, per user, and per time of day.
But there is one distinction that cuts this complexity in half.
A 401 is an authentication failure. "I cannot verify who you are." The token is missing, malformed, expired, or its signature does not match.
A 403 is an authorization failure. "I know who you are, but you are not allowed to do this." Identity verification passed, and permissions are lacking.
If you distinguish the two, the investigation goes in opposite directions. For a 401 you look at the credentials themselves — whether the token is being delivered properly and has not expired. For a 403 there is no need to touch the credentials, and you look at permission settings and scopes. If you lump them together as a "permission error" without this distinction, you dig in the wrong place for hours.
One more thing. If a 401 is returned but the response does not indicate a re-authentication path, the client falls into infinite retries. Because it keeps sending the same token without knowing how to refresh it.
And within 401 too, the reasons split. A missing token, an unknown token, and an expired token are all 401, but the response differs. So you must look at the reason string in the response body too. If you judge by the status code alone, in a situation where you should answer "issue a new token," you end up answering "check your configuration."
How it works
Certificate problems have a few fixed failure modes, each with different evidence.
Expiry — notAfter has passed. But there is something you must check together here. The client's clock. There are quite a lot of cases where an "certificate has expired" error occurs on an embedded device with no clock, a VM that booted after being stopped for a long time, or a container whose clock has drifted from the host, and then the server certificate is perfectly fine. You need the habit of printing the expiry date and the current time together.
Name mismatch — the name you connected with is not in the certificate. What people often get wrong here is looking at the CN. Modern clients do not look at the CN at all and check only subjectAltName. The judgment that "we put the domain in the CN, so it is fine" has been wrong since 2017.
Missing intermediate certificate — this is the type an FDE meets most often, and the symptom is distinctive. The padlock is normal in the browser, but only server-to-server calls fail. The cause is the server, not the client. The server is not sending the intermediate certificate along, and the browser follows the link inside the certificate to download the missing intermediate itself and cover for the server's mistake. curl and most language runtimes do not follow that link.
So the way to judge is simple. You just count how many certificates the server sends. If only the leaf certificate comes, the cause is confirmed at that moment, and the place to fix is the server, not the client.
What it looks like in the field
There is a shortcut an FDE is often tempted by here. Turning verification off.
The symptom disappears immediately. And if it is deployed to production in that state, it is left exposed to man-in-the-middle attacks. Worse, if you develop with verification off, you work in an environment that does not exist in production, so other problems are hidden as well. The hidden problems blow up together later.
Use turning off verification only to confirm "this problem is indeed the cause" after you have pinned down the cause, and turn it back on right after that confirmation.
Rules for handling other people's credentials
The statement that an FDE handles other people's keys carries more weight than technology. There are a few things to observe while receiving and using a customer's tokens and certificates, and if you break them, however well you diagnose, that project is over.
Receive only the minimum privilege needed. If what you want to investigate is one lookup API, a read token is enough. "Please give us an administrator token because it is convenient" is a sentence we must not say first, and even if the customer offers it first, if you cannot explain why that much is needed, it is better to give it back. When an incident happens, what could be done with that token becomes the scope of the investigation.
Set a lifetime and return it when done. When the investigation is over, request revocation and leave a record of that. A token still alive on a laptop half a year after the project ended is the most common path to an incident.
Decide where it lives. As seen in the previous course, environment variables can be read by other processes on the same host, and they also remain in shell history and command-line arguments. Even for temporary use, it is better to put it in a file and narrow the permissions. And do not open that file during screen sharing.
Filter it out of logs and reports. When you paste the diagnostic process, it is common for authentication headers to go along with it. You need the habit of skimming once before pasting, and for a script shared with the team, it is more reliable to make it not print the value in the first place.
And record turning off verification. I said earlier to use it only to confirm the cause, and I add one thing to that. If you turned it off, write one line in the report about why and when you turned it back on. Writing it down reduces the chance of that code being left by mistake, and above all, when the customer later finds that trace, there is something to explain.
What you will do in the next lab
You put in each of four tokens to produce 401 and 403 by hand, obtain the expiry reason string, extract the needed values from an expired certificate and from a certificate with a different name, and produce a diagnostic report.