TT Lab
Get started
Learn Learning paths Courses

The certificate was renewed, but the browser still showed the old one

A setup that issues fine and only fails at renewal

Continue in TT Lab

In one line

ACME is a protocol in which, once you prove over HTTP or DNS that "I control this domain," the CA issues a certificate, and cert-manager handles that proof and the renewal inside the cluster on your behalf. There is a kind of configuration mistake where issuance works fine but only renewal blows up, and we actually ran into it in the homelab.

Why you need this

When people renewed certificates by hand, expiry incidents were frequent. Short validity periods and automatic renewal were the answer, and that required a way for the CA to verify control of a domain without a human. RFC 8555 (ACME) is that protocol. The problem is that automation produces failures that look like successes. The first issuance went fine. After that, we changed one gateway setting. There were no symptoms — the certificate still had two months left. Only at renewal time does the challenge fail, and by then nobody thinks of that configuration change as the reason.

How it works

An ACME client places an order with the CA using its account key, and the CA gives a challenge for each domain. In HTTP-01, when you GET http://<도메인>/.well-known/acme-challenge/<token>, the body must equal the key authorization — token || '.' || base64url(Thumbprint(accountKey)) (replace the placeholders with your domain and the challenge token). The RFC says to send this request to TCP port 80, and notes that the validation server may follow redirects (SHOULD). According to Let's Encrypt's explanation, the actual implementation follows up to 10 redirects, but only to http: and https: on ports 80 and 443, and it does not verify the certificate when it is bounced to HTTPS. DNS-01 puts a value derived from the account key in a TXT record at _acme-challenge.<도메인>, and wildcard certificates can be obtained only through DNS-01.

GET /.well-known/acme-challenge/zk3r9Qm...   Host: www.lab.internal
200 OK
zk3r9Qm....Q7pVw2x...          ← token.thumbprint, 이것이 본문 전체

cert-manager turns this procedure into a controller. An Issuer/ClusterIssuer holds the agreement with the CA (account key and solvers), and a Certificate resource declares the certificate you want — secretName (the Secret where the result is stored), dnsNames, issuerRef (use kind: ClusterIssuer to use it from other namespaces too), duration (90 days by default), and renewBefore. As the documentation says, duration and renewBefore are Go duration strings, so only h, m, and s can be used and d (days) cannot, and duration must be greater than renewBefore. If you do not set renewBefore, renewal happens at the 2/3 point of the issued certificate's lifetime, and because the actual lifetime can come out shorter than requested, renewBeforePercentage is recommended over an absolute value. During a challenge, the HTTP-01 solver creates a temporary Ingress or HTTPRoute to send the challenge path to the solver Pod, and the documentation states that when you use the Gateway API, that route must attach to a Gateway that has a port 80 listener.

What it looks like in the field

This is a case from the homelab. We turned on an HTTP→HTTPS redirect on the gateway (Cilium Gateway). The site ran fine. But that redirect also bounced requests for /.well-known/acme-challenge/ to HTTPS, and because the HTTPS listener had no solver route, the end of that path was a 404. Issuance had already completed, so there were no symptoms, and the system was in a state where the renewal 30 days before expiry would silently fail. We actually measured "wouldn't the solver route win because its path is more specific?" — we brought up an Exact-path HTTPRoute imitating the solver and hit that path, and the redirect won (measured on Cilium 1.20.1). Path specificity did not take priority. The fix was to do the redirect in the app rather than at the gateway, with the challenge path as an exception, and the standing check is one line — if curl -sI http://<도메인>/.well-known/acme-challenge/x | head -1 returns 301, renewal is blocked, and if it returns 404 or 200, all is well. A 404 because the file does not exist is fine. It only must not be a redirect.

One more thing. Because this configuration had never gone through a renewal, we did not wait for expiry and instead forced a renewal by setting an Issuing condition in the status of the Certificate. A few seconds later a new CertificateRequest was created, and we made the verdict from the date on the certificate retrieved with openssl s_client having changed. Automation can be trusted only when there is an actual issuance history.

What you will do in the next lab

You start an edge (/opt/app/edge.py) with HTTP and HTTPS listeners and an HTTP-01 validation server (/opt/app/acme_va.py) inside a Pod. You put in the key authorization file and get the validation to pass as VALID, then turn on the redirect to reproduce the incident where 301 → 404 gives INVALID, and then make only the challenge path an exception to get VALID again while keeping the redirect. You write probe.sh, which checks whether renewal is blocked, and have it checked in both directions against the grader's server, and you write the cert-manager Certificate manifest with duration, renewBefore, and ClusterIssuer.