The Characters Are Broken: Decide, Don't Guess
In one line
A file does not have its encoding written in it. So the encoding is not something you find out but something you judge by ruling out candidates in order, and the last question of that judgment is "can this file be brought back to life?"
Why this was needed
You opened the roster the customer sent and the name column was all boxes or unknown symbols. The first thing most people do here is click through the encoding menu of the editor one by one. That method works in some cases, but the problem is that you end up judging only by eye whether it was right. With 300 files, the eye is useless.
More troublesome is the fact that there are files that look fine when opened but are wrong. The Korean is clearly visible but the lookup fails. The names look the same but the join gives 0 rows. The places where judging by eye does not work are always like this.
A text file is just a sequence of bytes, and which table to use to read those bytes is not written inside the file. The Unicode Consortium's UTF-8 and BOM FAQ drives home that a program that follows the specification must not interpret an invalid or non-conforming byte sequence as characters, and RFC 3629 defines which byte sequences are valid UTF-8. This strictness gives us one thing for free. Not just any bytes are UTF-8. So the very fact that "it reads as UTF-8" is strong evidence.
How it works
The judgment goes down in four steps. From the top, it goes from the most certain.
- Look at the BOM. If the file starts with
EF BB BF, it is UTF-8, and those three bytes are not data. The codec in Python that strips these three bytes when reading isutf-8-sig(the codecs documentation). If you read without stripping them, the first column name becomesmember_idand the header comparison quietly fails. - Try strict UTF-8 decoding. If it succeeds, it is almost certainly UTF-8. A single Hangul character is 3 bytes, and even the high bits of the continuation bytes are fixed, so the probability that the Hangul bytes of another encoding pass as UTF-8 by chance is very low.
- If it fails, set up candidates. For Korean material, they are CP949 and EUC-KR. CP949 is an extension of EUC-KR, so whatever reads as EUC-KR also reads as CP949. So instead of telling the two apart, you read everything as CP949 and tell them apart only when necessary.
- Cross-check a sample. Decide on a yardstick by which a machine can judge whether the characters read out with a candidate make sense. For a Korean roster, the yardstick is the proportion of precomposed Hangul syllables (U+AC00 to U+D7A3) among the characters other than ASCII. Without this yardstick, you move on with just "the decoding ended without an error", and that is not enough.
Next comes the real subject of this lab. When the judgment is finished, the file falls into three groups.
바이트
├─ 올바른 인코딩으로 읽힌다 → 정상. 읽어서 UTF-8 로 통일한다
├─ 한 번 잘못 읽혀 그대로 굳었다(mojibake) → 잘못 읽은 그 표로 되돌려 다시 읽으면 살아난다
└─ 잘못 읽으면서 글자가 뭉개졌다 → 못 되살린다. 원본을 다시 받아야 한다
The middle one is double encoding. If you read UTF-8 bytes as latin-1, a single Hangul character looks like three Latin letters. Here not a single byte has been lost — latin-1 is a table that attaches one character to every byte from 0 to 255, so the way back is not blocked. That is why one line, 깨진문자열.encode("latin-1").decode("utf-8"), brings the original text back (the placeholder is the garbled string).
The one on the right is a lost file. This happens when the side doing the wrong reading is lenient. If you replace bytes that cannot be read with U+FFFD (the replacement character) and move on, several different bytes are crushed into the same single character. Nothing is left in that file from which to recover the original bytes. This distinction is a judgment, not a matter of effort — if you see U+FFFD, that file has to be received again.
What it looks like in the field
First, it looks the same but the join gives 0 rows. The Hangul syllable "bak" (as in the surname Park) can be written as one precomposed syllable (U+BC15) or as three jamo (U+1107 U+1161 U+11A8). On screen they look identical but they are different strings. This is especially common with file names and data created on macOS. UAX #15 defines the normalization that converts between the two; the combining side is NFC and the decomposing side is NFD. Before comparing, you must normalize both sides to one form. In Python it is unicodedata.normalize.
Second, invisible characters get mixed into keys. A value copied from a web page comes with an NBSP (U+00A0) put in to prevent line breaks, or a zero-width space (U+200B) put in to set line-break positions. Neither is removed by strip() — in Python the NBSP is treated as whitespace and is removed, but in Java and in some DB functions it remains. A zero-width character is not removed anywhere. So for a value used as a key, scan it code point by code point and remove only what is on the list.
Third, you fix the encoding and overwrite the original. If you save the restored file in the place of the original, you have nowhere to return to if the judgment was wrong. Leave the bytes you received as they are, and make the fixed copy in a different directory.
Fourth, you look at one file and assert about the whole. Even within one delivery, it is common for the encoding to differ file by file. It is because each system exports along a different path. Judge file by file, and leave the result as a table.
What really matters in practice
- Turn guessing into a procedure. Instead of clicking through the editor menu, build code that goes down in the same order. That way you get the same answer even with 300 files.
- Say separately what can be restored and what is lost. "It is broken" is not a report. "Of the three files, two were restored and for one you need to give us the original again" is a report.
- Normalize before comparing. Normalization is for comparison, not for fixing the data. Leave the original as the original.
- Remove invisible characters by a list. If you count and keep what you removed, you can explain it when the counts do not match later.
What you will do in the next lab
You run a reproducer yourself that produces the same roster in seven states to set up an inbox, and then grow a judging tool, enc_probe.py, one step at a time. Starting from the BOM and the UTF-8 decodability check, you add candidate encodings and a sample check, restore double encoding, and separate the files left with U+FFFD as lost. Then you confirm with numbers why names that arrived as NFD do not match at all, and count the invisible characters by code point and remove them. The grader makes its own files with different names each time, actually runs your tool, and checks the judgments.