TT Lab
Get started
Learn Learning paths Courses

CS for Building Good Services — Relearning Textbook Ideas by Measuring

Pin Down Length, Truncation, Equality and Dates with 360 Nicknames

Continue in TT Lab

Goal

From a list of nicknames, you will count UTF-8 bytes, code points, and the characters a person sees separately, cut without breaking by a byte limit and a character limit, and count username collisions with normalization and case folding. Then you will group UTC events into the user's local date, convert a "every day at 09:00" reservation that crosses daylight saving time into UTC instants, and gather it all into a single signup review function.

Why it matters

If the screen, the server, and the storage count different lengths, the rule "up to 6 characters" becomes three rules. If you cut by bytes, characters break, and if you cut by code points, a family emoji is left with one person and an accent comes off. "The same name" can be counted as one only after you decide which of canonical equivalence (NFC), compatibility equivalence (NFKC), and case folding (casefold) you treat as equal. Dates are the same: if you do not decide "whose today," the daily numbers leak into the next date. So the grader does not look only at the numbers you wrote; it runs your functions again on inputs that are not the materials and compares them with the reference implementation.

Materials

They are under /opt/fixtures/svccs/text/. Only read them. All CSVs are UTF-8 and the first line is the header (read them with the csv module, passing newline="").

names.csv    user_id,nickname      기존 회원 닉네임 360줄(완성형·조합형 한글, 결합 문자, 이모지, 전각, ß …)
signups.csv  row,nickname          새 가입 신청 80줄. 빈 칸도 있다
users.csv    user_id,tz            이벤트를 낸 사용자와 IANA 시간대 이름
events.csv   event_id,user_id,ts_utc   UTC 로 찍힌 이벤트. ts_utc 는 2026-03-07T23:40:00Z 꼴
params.json  byte_limit · max_graphemes · schedule{zones, start, days, local_time}

Grapheme rule (this lab's simplified rule)

문자열을 앞에서부터 코드 포인트 하나씩 본다. 맨 앞 코드 포인트는 새 글자를 시작한다.
그 뒤로는 아래 가운데 하나라도 맞으면 앞 글자에 붙이고, 아니면 새 글자를 시작한다.
 1. 지금 코드 포인트가 결합 표시(unicodedata.category 가 Mn 또는 Me)·변형 선택자
    (U+FE00~U+FE0F)·피부색 수식자(U+1F3FB~U+1F3FF)·ZWJ(U+200D) 가운데 하나
 2. 바로 앞 코드 포인트가 ZWJ
 3. 한글 묶음 — 앞이 L 이고 지금이 L·V·LV·LVT / 앞이 V 나 LV 이고 지금이 V·T /
    앞이 T 나 LVT 이고 지금이 T
    L = U+1100~U+115F, V = U+1160~U+11A7, T = U+11A8~U+11FF,
    완성형 U+AC00~U+D7A3 은 (코드 − 0xAC00) % 28 == 0 이면 LV, 아니면 LVT
국기(지역 표시 문자 쌍)·Prepend·SpacingMark 는 다루지 않는다(UAX #29 전체가 아니다).

Steps

  1. In /root/svccs/text/lengths.json, write, for the nicknames in names.csv, rows (the number of lines), utf8_bytes (the sum of UTF-8 bytes), code_points (the sum of code points), rows_multibyte (the number of lines where the byte count ≠ the code point count), and max_utf8_bytes (the maximum bytes of one line).
  2. In /root/svccs/text/textkit.py, create graphemes(s): it returns the list of strings split according to the "Grapheme rule" above (joining them gives s). And in /root/svccs/text/graphemes.json, write graphemes (the sum of grapheme counts), rows_cp_ne_graphemes (the number of lines where the code point count ≠ the grapheme count), and longest_cluster_cp (the maximum number of code points in one grapheme).
  3. In the same file, create cut_bytes(s, limit): the longest prefix that is at most limit bytes in UTF-8 without splitting a code point. Let L be the value stored in params.json under byte_limit. In /root/svccs/text/cut.json, write byte_limit, rows_over_limit (the number of lines whose bytes exceed L), rows_naive_broken (the number of lines for which strictly decoding s.encode("utf-8")[:L] raises UnicodeDecodeError), and utf8_bytes_after_cut (the byte sum after applying cut_bytes(s, L) to every line).
  4. In the same file, create fit(s, max_bytes, max_graphemes): the longest prefix that holds only whole graphemes (by the rule of step 2) and keeps both grapheme count ≤ max_graphemes and bytes ≤ max_bytes. With B = byte_limit and G = max_graphemes, in /root/svccs/text/fit.json write max_bytes, max_graphemes, rows_changed (the number of lines where the result of fit differs from the original), split_by_cut_bytes (the number of lines where the end of cut_bytes(s, B) is not a grapheme boundary), and split_by_cp_slice (the number of lines where the end of s[:G] is not a grapheme boundary). A grapheme boundary is a code point position that arises when you join the graphemes from the front (including 0 and len(s)).
  5. In the same file, create username_key(s) = NFKC(casefold(NFKC(s))). In /root/svccs/text/unique.json, write the numbers of distinct values: distinct_raw (the original), distinct_nfc (NFC), distinct_nfc_lower (NFC then lower), distinct_nfc_casefold (NFC then casefold), and distinct_key (username_key), plus collision_pairs (the number of pairs that can be made among lines with the same key, the sum of n·(n−1)/2 over each key) and collision_groups (the lists of user_ids that share the same key, only those with 2 or more).
  6. In the same file, create local_date(ts_utc, tz): it converts a UTC time of the form found in events.csv into the local date "YYYY-MM-DD" of the IANA time zone tz. Group each event by that user's time zone in users.csv, and in /root/svccs/text/daily.json write events (the number of lines), by_local_date (local date → number of events), by_utc_date (first 10 characters of ts_utc → number of events), and events_on_other_date (the number of events where the local date ≠ the UTC date).
  7. In the same file, create daily_at(tz, start_date, days, hhmm): for days days starting from start_date ("YYYY-MM-DD"), it returns the instant of each day's local hhmm (of the form "09:00") as a list of "YYYY-MM-DDTHH:MM:SSZ" (UTC) strings. Using the value in params.json under schedule, in /root/svccs/text/schedule.json write, for each time zone, {"utc": 목록, "gaps_hours": 이웃한 두 순간의 실제 간격(시간, 넷째 자리) 목록, "short_days": 24 미만인 간격 수, "long_days": 24 초과인 간격 수} (utc is the list, gaps_hours is the list of the actual gaps in hours between adjacent instants to four decimal places, short_days is the number of gaps under 24, and long_days is the number of gaps over 24).
  8. In the same file, create check_signup(nickname, taken, max_bytes, max_graphemes), review signups.csv in order following the "Signup rule" below, and in /root/svccs/text/signups.json write the counts per status ok·taken·too_many_chars·too_many_bytes·empty and a verdicts list in line order (the limits are byte_limit·max_graphemes).

Signup rule

shown = NFC(nickname)                        저장·표시는 NFC 로 한다
1. shown 이 공백(str.isspace)·ZWJ·규칙 1의 결합류로만 되어 있거나 비었으면  "empty"
2. graphemes(shown) 의 개수가 max_graphemes 를 넘으면                    "too_many_chars"
3. shown 의 UTF-8 바이트가 max_bytes 를 넘으면                           "too_many_bytes"
4. username_key(nickname) 이 taken 안에 있으면                          "taken"
5. 아니면                                                              "ok"
taken 은 처음에 names.csv 모든 닉네임의 키이고, "ok" 를 받은 신청의 키가 차례로 더해진다.
채점기는 taken 을 목록(list)으로 넘길 수도 있다 — in 으로만 쓰세요.

Notes

Bytes and code points are different

Write the number of lines, the sum of UTF-8 bytes, the sum of code points, the number of lines where the two differ, and the maximum bytes for the nicknames in names.csv to /root/svccs/text/lengths.json as rows, utf8_bytes, code_points, rows_multibyte, and max_utf8_bytes.

A Python str is a sequence of code points, so len() is the number of code points. Bytes are the length of s.encode("utf-8"). A precomposed Hangul syllable is 3 bytes, and ASCII is 1 byte.

Count the characters a person sees

Create graphemes(s) in /root/svccs/text/textkit.py following the grapheme rule, and write graphemes, rows_cp_ne_graphemes, and longest_cluster_cp to /root/svccs/text/graphemes.json. The grader calls the function directly and compares it on variant strings that mix combining characters, ZWJ, and jamo.

You only need to decide whether to attach the current code point to the preceding grapheme. unicodedata.combining() returns 0 for variation selectors (U+FE0F) and ZWJ, so it alone is not enough: look at the category and the ranges together. A vowel jamo (V) after a leading consonant jamo (L) attaches, while a leading consonant (L) after a final consonant (T) starts a new grapheme.

Cut by the byte limit without breaking a character

Create cut_bytes(s, limit) in textkit.py, and with byte_limit write byte_limit, rows_over_limit, rows_naive_broken, and utf8_bytes_after_cut to /root/svccs/text/cut.json. The grader calls cut_bytes with variant strings and several limits.

In UTF-8 a code point is 1 to 4 bytes, and every continuation byte has the form 10xxxxxx. Add code points one at a time while adding their bytes, and stop just before you would exceed the limit. If you cut by bytes and then read with errors="replace", U+FFFD is left at the end.

Cut by grapheme

Create fit(s, max_bytes, max_graphemes) in textkit.py, and write max_bytes, max_graphemes, rows_changed, split_by_cut_bytes, and split_by_cp_slice to /root/svccs/text/fit.json. The grader calls fit with variant strings and several limits.

Add the graphemes split by graphemes(s) one at a time, and stop just before either of the two limits, grapheme count or bytes, would be exceeded. Even a result cut safely at code point level can leave only one person of a family emoji or pull an accent off, and that count is split_by_cut_bytes.

Decide on the same name

Create username_key(s) = NFKC(casefold(NFKC(s))) in textkit.py, and write distinct_raw, distinct_nfc, distinct_nfc_lower, distinct_nfc_casefold, distinct_key, collision_pairs, and collision_groups to /root/svccs/text/unique.json. The grader also calls username_key with variant strings.

NFC reconciles precomposed and jamo-decomposed forms, casefold folds ß into ss, and NFKC reconciles compatibility characters such as fullwidth characters, ligatures, and the Kelvin sign. Each time you add a rule, the number of distinct values should go down. The number of pairs is n·(n−1)/2 for every n with the same key.

Group UTC events by local date

Create local_date(ts_utc, tz) in textkit.py, group events.csv by the date in each user's time zone, and write events, by_local_date, by_utc_date, and events_on_other_date to /root/svccs/text/daily.json. The grader calls local_date with variant events in other time zones and dates.

Read ts_utc as a UTC aware datetime (the trailing Z means UTC), move it with astimezone(ZoneInfo(tz)), and then look at .date(). Seoul is ahead of UTC, so if you group by UTC date, Seoul's early morning slips into the previous day.

Turn a daily 09:00 local reservation into UTC

Create daily_at(tz, start_date, days, hhmm) in textkit.py, and with the schedule in params.json write utc, gaps_hours, short_days, and long_days for each time zone to /root/svccs/text/schedule.json. The grader calls daily_at with other time zones and periods.

If you convert the first instant to UTC and add 24 hours at a time, the local time shifts by an hour after a daylight saving transition. Build hhmm fresh for each local date and convert it to UTC. Subtract the UTC instants from each other for the gaps: if you subtract aware datetimes that share the same ZoneInfo, the tzinfo is ignored and you always get 24 hours.

Gather the signup review into one function

Create check_signup(nickname, taken, max_bytes, max_graphemes) in textkit.py following the signup rule, review signups.csv in order, and write ok, taken, too_many_chars, too_many_bytes, empty, and verdicts to /root/svccs/text/signups.json. The grader also calls check_signup with variant signups and taken lists.

The order of the checks is part of the rule. The byte limit is measured after converting to NFC: a Korean name that arrives in decomposed form measures three times larger if you measure the original as is. If you do not add the key of an accepted signup to taken, two signups that arrive on the same day both pass.