TT Lab
Get started
Learn Learning paths Courses

Tomcat & nginx Operations

90% of Certificate Incidents Are the Chain and Expiry

Continue in TT Lab

In one line

Most certificate outages come not from cryptography but from two things — a chain with the intermediate certificate not attached, and an expiry date nobody was watching.

Why this is a problem

These two are dangerous because they don't reproduce on a developer PC. Browsers cache intermediate certificates or download and fill them in on their own, so the screen comes up fine even when the chain is missing. But server-to-server integrations (Java clients, batches) make no such correction and simply fail. "It works in my browser" comes from here.

Expiry is simpler and more painful. The moment the validity period passes it is a total outage, it starts at dawn, and the fix is swapping one file, but getting that file issued takes days. So it is a matter of monitoring, not response — something that ends with one script counting the remaining days and one alert, yet somewhere experiences it every year.

The moments you meet certificates at an SI site

Certificates usually arrive like this. The customer's security team emails a single .pfx file and a password. Or three files arrive, server.crt, server.key and chain.crt. Sometimes only a single .cer arrives. And it says "please apply it before go-live."

The knowledge needed at this point is not cryptography but how to check what a file contains.

# 인증서 내용 보기 (주체, 발급자, 유효기간, SAN)
openssl x509 -in server.crt -noout -subject -issuer -dates -ext subjectAltName

# 키와 인증서가 짝인지 확인 (두 해시가 같아야 함)
openssl x509 -in server.crt -noout -modulus | openssl md5
openssl rsa  -in server.key -noout -modulus | openssl md5

# pfx 를 crt/key 로 분해
openssl pkcs12 -in cert.pfx -clcerts -nokeys  -out server.crt
openssl pkcs12 -in cert.pfx -nocerts -nodes   -out server.key

If the key and certificate are not a pair, nginx doesn't even start and leaves only the terse message key values mismatch. If you know these two command lines, it is over in 30 seconds.

The chain — why it works on a developer PC but not between servers

A certificate usually has three tiers.

Root CA (브라우저·OS 에 이미 들어 있음)
   └─ Intermediate CA (중간 인증서)
        └─ 서버 인증서 (우리 것)

The server must send its own certificate + the intermediate certificate together. The other side already has the Root. What happens if you leave out the intermediate certificate?

So the symptom appears like this. "It works fine in a developer PC browser, but only the integrated counterpart system gets an SSL error." This is the typical picture of a missing chain. Checking takes one line.

echo | openssl s_client -connect api.example.com:443 \
  -servername api.example.com -showcerts 2>/dev/null | grep -c 'BEGIN CERTIFICATE'

If it prints 1, there is no chain. When healthy, it prints 2 or more.

For nginx, you give ssl_certificate a file with the server certificate followed by the intermediate certificate, concatenated. The reverse order doesn't work.

cat server.crt intermediate.crt > fullchain.pem

Without a SAN, today's browsers reject it

In the old days you wrote the domain in the CN (Common Name). Now, without the SAN (Subject Alternative Name) extension, recent browsers reject the certificate. The CN has become reference only.

When making a self-signed certificate, it is common to forget the SAN and repeat "why doesn't this work". When you create a private CA for development/test environments, always include the SAN.

subjectAltName = DNS:labhub.local, DNS:*.labhub.local, IP:127.0.0.1

Expiry — the most common and the most absurd outage

To give one real case, in a large SSO migration project a SAML signing certificate expired and caused a 47-minute total login outage. Certificate expiry

So expiry monitoring must be done by a script, not a person.

openssl x509 -in server.crt -noout -enddate
# notAfter=Nov 12 09:00:00 2026 GMT

# 임계일 이내면 실패로 종료 (checkend 는 초 단위)
openssl x509 -in server.crt -noout -checkend $((30*86400)) \
  || echo "30일 이내 만료"

-checkend is hardly known, but when you build a monitoring script, it is exactly for this purpose. You have to manage as a list not only server certificates but also the counterpart's certificates, client certificates, SAML signing certificates and code signing certificates. Without a list, you will surely miss one.

Practical defaults for nginx TLS configuration

server {
    listen 8443 ssl;
    server_name labhub.local;

    ssl_certificate     /etc/nginx/tls/fullchain.pem;
    ssl_certificate_key /etc/nginx/tls/server.key;

    ssl_protocols       TLSv1.2 TLSv1.3;
    ssl_ciphers         HIGH:!aNULL:!MD5;
    ssl_prefer_server_ciphers off;
    ssl_session_cache   shared:SSL:10m;
    ssl_session_timeout 10m;
}

HTTP → HTTPS redirect

server {
    listen 8088;
    server_name labhub.local;
    return 301 https://$host:8443$request_uri;
}

Using, instead of rewrite, return 301 is recommended. It is faster and the intent is clearer. And if you set up the redirect, also consider HSTS. But once HSTS is stamped into a browser it is hard to undo, so for internal systems it is safer to start with a short max-age and lengthen it.