TT Lab
Get started
Learn Learning paths Courses

Working With Customer Data

Every Name Is a Box: Build a Detector

Continue in TT Lab

Goal

You build enc_probe.py, a tool that judges the encoding of a file the customer sent without guessing. Starting from the BOM and UTF-8 decodability, you go down to candidate encodings and a sample check, restore double encoding, single out files that cannot be restored, and prove with your own data how normalization and invisible characters wrecked a comparison.

Why it matters

A text file does not have its encoding written in it. So to the question "what encoding is it?", you can answer only by a procedure. Clicking through the editor menu works only when there are three files, and above all it makes you judge only by eye whether it was right. When the judgment is finished, the file falls into three groups. Files that read correctly, files that were once read wrongly and hardened but can be reversed, and files whose characters were crushed so that the original bytes cannot be known. Spending time on the last group is not effort but waste — that file has to be received again. You must be able to make this distinction in code to be able to say what to request from the customer. And there are places where the characters look fine but do not match. If the same Hangul is stored split into precomposed and decomposed jamo forms, the screen is the same and the bytes differ. The invisible NBSP and zero-width space are the same. The situation where a join gives 0 rows and the cause is not visible on screen comes from here. The grader does not trust your wording. It sets up files it made in a temporary directory, actually runs your tool, and checks the judgments. The names and the row counts change on every run.

Steps

  1. Create and run /root/enc/gen_inbox.py to produce seven files under /root/enc/inbox. The same roster of 24 people is stored in seven states.
  2. Build detect in /root/enc/enc_probe.py so that it judges the encoding in the order of the BOM, UTF-8 decodability, candidate encodings, and sample check.
  3. Add decode so that it reads with the judged encoding and outputs UTF-8 to standard output, and save the whole inbox unified under the same names below /root/enc/norm.
  4. Add repair so that it restores double encoding, and make detect answer such a file with mojibake set to true and the verdict mojibake.
  5. Add lossy to detect so that the verdict of a file left with U+FFFD is lost, and write the inbox judgment in /root/enc/lost.json.
  6. Add keys so that it outputs the name column as comparison keys normalized to NFC, and write in /root/enc/join_report.json how many names of the NFD version and the UTF-8 version match before and after normalization.
  7. Add scan so that it counts invisible characters by code point, have keys strip them out, and write /root/enc/invisible_report.json.
  8. Process all seven inbox files in one go to produce /root/enc/enc_report.json and /root/enc/enc_report.md.

Notes

Build the same roster in seven states

Create and run /root/enc/gen_inbox.py to produce seven files under /root/enc/inbox. The contents are the same 24 people, and only the stored bytes differ.

First create /root/enc, and write the files with python3 inside it. All there is to it is to make the text and then store it as bytes with a different encode each time. You must open the file in binary mode ("wb"), not text mode, so that the encoding is not applied twice.

Judge without guessing

Create /root/enc/enc_probe.py so that detect <파일> (the placeholder is the file) outputs, as one chunk of JSON, the result judged in the order of the BOM, UTF-8 decodability, candidate encodings, and sample check. The response must have path, bom, utf8_ok, encoding, and verdict.

The order is the judgment. First see whether the first three bytes are EF BB BF, if not try strict UTF-8 decoding, and if that fails too, try reading as CP949. Finally, check with the precomposed Hangul ratio whether the characters read out make sense. That it decoded without an error is not enough.

Read as judged and unify to UTF-8

Add decode <파일> (the placeholder is the file) so that it outputs the text read with the judged encoding to standard output as UTF-8 (with the BOM stripped). Using that result, save unified copies under the same names below /root/enc/norm.

You can simply reuse the judgment result. A file with a BOM must be read as utf-8-sig so that no invisible character is left in the first column name. Leave the originals as they are and make the unified copies in a different directory — there has to be somewhere to return to when the judgment was wrong.

Restore characters that were read wrongly once and hardened

Add repair <파일> (the placeholder is the file) so that it restores double encoding and outputs it to standard output, and add mojibake to the detect response so that the verdict of such a file is mojibake. If it cannot be restored, the exit code is 3.

If you read UTF-8 bytes as latin-1, a single Hangul character becomes three Latin letters. Not a single byte was lost, so you can go back the same way. The grounds for judging that it was restored are that the precomposed Hangul ratio of the restored side went up — if you drop this check, you end up touching even fine files.

Separate what can be restored from what is already lost

Add lossy to the detect response so that the verdict of a file left with U+FFFD is lost. Then walk the inbox and write two lists, lost and recoverable, and a reason in /root/enc/lost.json.

When the side doing the wrong reading is lenient, bytes that cannot be read are crushed into a single replacement character. Since several different bytes become the same character, no information is left to reverse it. This is not a matter of effort but a judgment — that file has to be received again. lost and recoverable are lists holding only file names.

Names that look the same but do not match

Add keys <파일> (the placeholder is the file) so that it outputs the second column (the name) as an array of comparison keys normalized to NFC. And write in /root/enc/join_report.json, as rows, matched_raw, and matched_nfc, how many names of the NFD version and the UTF-8 version match before normalization and how many after.

A single Hangul character is stored either as one precomposed syllable or as three jamo. They look the same on screen but are different strings. Normalization is for comparison, not for fixing the original — leave the original files as they are. matched_raw is the number of unnormalized NFD names that are found as they are in the set of names of the UTF-8 version.

Invisible characters got mixed into the keys

Add scan <파일> (the placeholder is the file) so that it counts invisible characters by code point, and have keys replace NBSP with a normal space and strip out zero-width characters. Write the result in /root/enc/invisible_report.json as counts, rows_affected, and keys_fixed.

Zero-width characters are not removed by strip(). You have to decide what to remove as a list of code points and filter by scanning one character at a time. rows_affected is the number of lines that contained at least one invisible character, and keys_fixed is the number of keys that came to match a name of the UTF-8 version after the cleanup.

Report on one sheet for the inbox

Process all seven inbox files in one go and write files, ok, mojibake, and lost in /root/enc/enc_report.json, and write /root/enc/enc_report.md in four sections: ## 무엇을 받았나 ## 어떻게 판정했나 ## 되살린 것과 잃은 것 ## 보내는 쪽에 요청할 것 (the Korean headings mean "What we received", "How we judged it", "What was restored and what was lost", and "What to request from the sender").

files is an object keyed by file name that holds encoding and verdict. In the report, write the number of restored files and the number of lost files as numbers — the purpose of this document is what the customer has to send again. You can simply call the judging function you built in the earlier steps.