TT Lab
Get started
Learn Learning paths Courses

CS for Building Good Services — Relearning Textbook Ideas by Measuring

A String Has Three Lengths, and Whose Today Is Today?

Continue in TT Lab

In one line

The single line "nicknames up to 10 characters" hides three lengths: the UTF-8 bytes the storage counts, the code points that Python's len() counts, and the characters a person counts. On top of that, "the same name" and "today" are also places where the human eye and the machine's count part ways, and you will reconcile them with numbers.

Why this was needed

The signup screen says "up to 6 characters," the server checks with len(name) <= 6, and the storage column measures in bytes. In an English-only test the three are the same and there is no problem. Then one family emoji comes in: on the screen it is one character, but len() gives 5 and the bytes are 18. If you truncate a notification preview by bytes, a question mark in a diamond is printed at the end, and a Korean name pasted from a Mac is the same name yet passes the duplicate check.

Dates have the same shape. Events are stamped in UTC, but users live in "today" in their own time zone. If you group daily statistics by UTC date, a Seoul user's activity before nine in the morning slips into the previous day.

How it works

Three lengths. These are examples measured with Python 3.12 on this Pod.

String Code points UTF-8 bytes Characters a person sees
a 1 1 1
ß 1 2 1
가 (precomposed) 1 3 1
가 (decomposed into two jamo, NFD) 2 6 1
é (e + combining accent) 2 3 1
👍🏽 (skin tone modifier) 2 8 1
Family emoji (three people joined with ZWJ) 5 18 1

The Python documentation defines str as an "immutable sequence of code points," so len() is the number of code points. The bytes are decided by the encoding. UTF-8 in RFC 3629 writes one character in 1–4 octets, every continuation octet has the form 10xxxxxx, and the first octet tells you the total length. So, as the RFC says, you can easily find a character boundary anywhere in an octet stream. When cutting at a byte limit, if you just step back until you reach a position that is not 10xxxxxx, no broken character is left.

The characters a person sees are the grapheme clusters of UAX #29. They are made of rules such as not breaking before a combining mark or ZWJ (GB9), the rule that binds Hangul jamo L·V·T into one syllable (GB6–GB8), the emoji ZWJ sequence rule (GB11), and the flag pair rules (GB12·GB13). The standard library's unicodedata has no function that splits at these boundaries. So in the lab you write a simplified rule yourself. Combining marks, variation selectors, skin tone modifiers, and ZWJ attach to what comes before, the character after a ZWJ attaches too, and jamo are bound by the L·V·T rule. Flags, Prepend, and SpacingMark are not handled, and the condition in GB11 that "the preceding character must be a pictograph" is left out. In a real service you would use a library that implements UAX #29, and you would also need to check which Unicode version that library follows.

The same name. UAX #15 separates two kinds of sameness. Canonical equivalence is the relationship of "the same abstract character that always looks the same when rendered correctly," as with the precomposed 가 and the jamo sequence ㄱ+ㅏ. NFC and NFD reconcile this. Compatibility equivalence is weaker and allows different shapes: fullwidth K and K, the ligature fi and fi, and NFKC and NFKD reconcile these. The same document warns that KC and KD erase formatting distinctions and so "should not be applied blindly to arbitrary text." So leave the original as it is and use them only for the comparison key.

For case, the example in the str.casefold documentation says it all. ß is already lowercase, so lower() leaves it alone, but casefold() turns it into ss. Straße and STRASSE stay apart under lower and meet under casefold.

Whose today is today? The datetime documentation has two sentences. Even if you add a timedelta to an aware datetime, "no time zone adjustments are done." If you subtract aware datetimes that have the same tzinfo, the tzinfo is ignored and the difference in wall-clock time is returned. The example in the zoneinfo documentation shows the result. If you add timedelta(days=1) to 12:00 in Los Angeles, it is 12:00 on the next day even after daylight saving time ends. Conversely, if you add 24 hours at the moment converted to UTC, the local time shifts by an hour on the transition day. "Every day at 09:00" means creating 09:00 for each local date and converting it to UTC, and the real interval between two reservations must be subtracted after converting to UTC for the 23 hours and 25 hours to show. zoneinfo reads the system's IANA database or the tzdata package, so whether that data exists in the container image is also worth checking.

What it looks like in the field

Encoding detection, mojibake, and stripping invisible characters are covered in depth by the fde-dat-encoding lab in the "Working With Customer Data" course. Separating, with fold, a time that does not exist and a time that occurs twice on a daylight saving transition day is handled by the "The Logs Came From the Future — Five Incidents a Clock Made" course. Why a batch that silently swallows truncated strings with errors="ignore" and ends with 0 is dangerous is covered, as a matter of tool discipline, by the "It failed, but the exit code was 0" course. This module is the place in between where you pin down, as service rules, the counting of length, truncation, sameness, and dates.

What you will do in the next lab

You will count the bytes, code points, and characters of 360 nicknames, implement the simplified grapheme rule, and then cut without breaking by the byte limit and the character limit. You will layer on NFC, casefold, and NFKC one after another to count the colliding names, group UTC events into local dates, convert a "every day at 09:00" reservation that crosses daylight saving time into UTC, and finally gather it all into a single signup review function.