TT Lab
Get started
Learn Learning paths Courses

Debugging in Practice

It Works on Our Server

Continue in TT Lab

One-line summary

When the same code fails only here, the culprit is usually one of these: environment variables, encoding, permissions, or the fact that the shell's values differ from the process's values.

Why this is needed

The sentence an FDE hears most often is "it works on our server." And that statement is usually true. The code is the same, and the environment differs.

An environment difference is harder to find than a code defect. You can see code by reading it, but there is nothing to read in an environment. That is why you need an order.

How it works

You look at four things, in the order they actually come up most often in the field.

First, environment variables. What tells you that it fails when one is missing is usually the error message itself, such as missing env API_TOKEN. The problem is when you get the answer "but I set it" even after seeing this message, and almost always the place where it was set and the place where it is read are different. A value exported in a shell goes only to that shell's child processes. A service started by systemd cannot see that value.

Second, encoding. It blows up overwhelmingly often with Korean customer data. CSVs made in a Windows environment are often EUC-KR (CP949), not UTF-8, and if you read such a file as UTF-8, the characters get garbled or it dies with a decoding error. The file command does not always guess the encoding correctly, so if Korean text looks garbled, the fastest thing is to try converting with iconv -f EUC-KR -t UTF-8.

The trap here is relying on automatic conversion. A string read with the wrong encoding passes quietly, gets stored in the database, and comes back weeks later as a report that search does not work. Encoding must be settled at the point of reading.

Third, permissions. This is especially so for credential files. If a token file is left at 644, every other user on the same server can read it, and in a security review this one thing erodes trust in the whole project. An FDE handles the keys to someone else's house, so this weighs especially heavily. Least privilege is not courtesy but a survival rule.

Fourth, the shell's value and the process's value being different. This is what fools people most often. Taking the open file limit as an example, even if ulimit -n in the shell shows 1048576, the limit of the actually running service can be 1024. The place to check is not the shell but /proc/<PID>/limits.

The same trap repeats in several layers. /etc/security/limits.conf applies only to PAM login sessions and not to systemd services. free and nproc inside a container show the host's values, while the real limits are in cgroups. There is one principle — measure from the running process, not from the configuration file.

What you see in the field

Building one environment check script replaces all of this. It is a script that checks, in order, the required environment variables, the existence and permissions of files, the locale, and whether the data is accessible, and stops at the first failure while stating the reason.

The value of such a script lies not in diagnosis time but in the quality of the conversation. When the sentence "envcheck says API_TOKEN is missing" is exchanged instead of "it doesn't work," the round trips with the customer's contact drop from three to one.

Moving toward removing the differences

Finding environment differences well matters, but the better path is to reduce the places where differences can arise. Pushing the handed-over system bit by bit in that direction is also the value an FDE leaves behind.

Make a list of what it depends on. If what this program needs is written down nowhere, there is no way to know what is missing in a new environment. The required environment variables, the paths and addresses it must reach, the required commands and their versions. This list becomes the content of the check script built earlier.

Have defaults, but do not slide past them quietly. Falling back to a reasonable default when a value is missing is convenient, but if that fact is not left in the log, you cannot later explain "why are we getting different results?" Just leaving one line saying a default was used greatly reduces investigation time.

Check everything once at startup and die. If something required is missing, rather than failing somewhere in the middle of a run, it is better to check everything immediately at startup, report everything that is missing at once, and stop. Then what would have taken three round trips ends in one.

Write the environment as code. What was configured by hand cannot be reproduced, and what cannot be reproduced leaves no way to explain why two environments differ. Whether it is a container image or a configuration management tool, if the procedure that creates that environment remains as a file, the sentence "it fails only here" itself becomes hard to sustain.

If you demand all four at once, the customer feels burdened. So it is better to start from one incident you have just experienced. The proposal "we spent half a day on API_TOKEN this time, so let's put that into the check script first" is easy to accept, and a script that grows one line at a time becomes quite a useful document a few months later.

What you will do in the next lab

Starting from a state where the environment check script fails, you capture the failure output, identify the missing environment variable, revive an EUC-KR Korean CSV as UTF-8, and apply least privilege to the token file to get the check to pass.