CS for Building Good Services — Relearning Textbook Ideas by Measuring
Pin Down Length, Truncation, Equality and Dates with 360 Nicknames
Goal
From a list of nicknames, you will count UTF-8 bytes, code points, and the characters a person sees separately, cut without breaking by a byte limit and a character limit, and count username collisions with normalization and case folding. Then you will group UTC events into the user's local date, convert a "every day at 09:00" reservation that crosses daylight saving time into UTC instants, and gather it all into a single signup review function.
Why it matters
If the screen, the server, and the storage count different lengths, the rule "up to 6 characters" becomes three rules. If you cut by bytes, characters break, and if you cut by code points, a family emoji is left with one person and an accent comes off. "The same name" can be counted as one only after you decide which of canonical equivalence (NFC), compatibility equivalence (NFKC), and case folding (casefold) you treat as equal. Dates are the same: if you do not decide "whose today," the daily numbers leak into the next date. So the grader does not look only at the numbers you wrote; it runs your functions again on inputs that are not the materials and compares them with the reference implementation.
Materials
They are under /opt/fixtures/svccs/text/. Only read them. All CSVs are UTF-8 and the first line is the header (read them with the csv module, passing newline="").
names.csv user_id,nickname 기존 회원 닉네임 360줄(완성형·조합형 한글, 결합 문자, 이모지, 전각, ß …)
signups.csv row,nickname 새 가입 신청 80줄. 빈 칸도 있다
users.csv user_id,tz 이벤트를 낸 사용자와 IANA 시간대 이름
events.csv event_id,user_id,ts_utc UTC 로 찍힌 이벤트. ts_utc 는 2026-03-07T23:40:00Z 꼴
params.json byte_limit · max_graphemes · schedule{zones, start, days, local_time}
Grapheme rule (this lab's simplified rule)
문자열을 앞에서부터 코드 포인트 하나씩 본다. 맨 앞 코드 포인트는 새 글자를 시작한다.
그 뒤로는 아래 가운데 하나라도 맞으면 앞 글자에 붙이고, 아니면 새 글자를 시작한다.
1. 지금 코드 포인트가 결합 표시(unicodedata.category 가 Mn 또는 Me)·변형 선택자
(U+FE00~U+FE0F)·피부색 수식자(U+1F3FB~U+1F3FF)·ZWJ(U+200D) 가운데 하나
2. 바로 앞 코드 포인트가 ZWJ
3. 한글 묶음 — 앞이 L 이고 지금이 L·V·LV·LVT / 앞이 V 나 LV 이고 지금이 V·T /
앞이 T 나 LVT 이고 지금이 T
L = U+1100~U+115F, V = U+1160~U+11A7, T = U+11A8~U+11FF,
완성형 U+AC00~U+D7A3 은 (코드 − 0xAC00) % 28 == 0 이면 LV, 아니면 LVT
국기(지역 표시 문자 쌍)·Prepend·SpacingMark 는 다루지 않는다(UAX #29 전체가 아니다).
Steps
- In
/root/svccs/text/lengths.json, write, for the nicknames innames.csv,rows(the number of lines),utf8_bytes(the sum of UTF-8 bytes),code_points(the sum of code points),rows_multibyte(the number of lines where the byte count ≠ the code point count), andmax_utf8_bytes(the maximum bytes of one line). - In
/root/svccs/text/textkit.py, creategraphemes(s): it returns the list of strings split according to the "Grapheme rule" above (joining them gives s). And in/root/svccs/text/graphemes.json, writegraphemes(the sum of grapheme counts),rows_cp_ne_graphemes(the number of lines where the code point count ≠ the grapheme count), andlongest_cluster_cp(the maximum number of code points in one grapheme). - In the same file, create
cut_bytes(s, limit): the longest prefix that is at most limit bytes in UTF-8 without splitting a code point. Let L be the value stored inparams.jsonunderbyte_limit. In/root/svccs/text/cut.json, writebyte_limit,rows_over_limit(the number of lines whose bytes exceed L),rows_naive_broken(the number of lines for which strictly decodings.encode("utf-8")[:L]raises UnicodeDecodeError), andutf8_bytes_after_cut(the byte sum after applyingcut_bytes(s, L)to every line). - In the same file, create
fit(s, max_bytes, max_graphemes): the longest prefix that holds only whole graphemes (by the rule of step 2) and keeps both grapheme count ≤ max_graphemes and bytes ≤ max_bytes. With B =byte_limitand G =max_graphemes, in/root/svccs/text/fit.jsonwritemax_bytes,max_graphemes,rows_changed(the number of lines where the result offitdiffers from the original),split_by_cut_bytes(the number of lines where the end ofcut_bytes(s, B)is not a grapheme boundary), andsplit_by_cp_slice(the number of lines where the end ofs[:G]is not a grapheme boundary). A grapheme boundary is a code point position that arises when you join the graphemes from the front (including 0 and len(s)). - In the same file, create
username_key(s)=NFKC(casefold(NFKC(s))). In/root/svccs/text/unique.json, write the numbers of distinct values:distinct_raw(the original),distinct_nfc(NFC),distinct_nfc_lower(NFC then lower),distinct_nfc_casefold(NFC then casefold), anddistinct_key(username_key), pluscollision_pairs(the number of pairs that can be made among lines with the same key, the sum of n·(n−1)/2 over each key) andcollision_groups(the lists of user_ids that share the same key, only those with 2 or more). - In the same file, create
local_date(ts_utc, tz): it converts a UTC time of the form found inevents.csvinto the local date"YYYY-MM-DD"of the IANA time zone tz. Group each event by that user's time zone inusers.csv, and in/root/svccs/text/daily.jsonwriteevents(the number of lines),by_local_date(local date → number of events),by_utc_date(first 10 characters of ts_utc → number of events), andevents_on_other_date(the number of events where the local date ≠ the UTC date). - In the same file, create
daily_at(tz, start_date, days, hhmm): for days days starting from start_date ("YYYY-MM-DD"), it returns the instant of each day's local hhmm (of the form"09:00") as a list of"YYYY-MM-DDTHH:MM:SSZ"(UTC) strings. Using the value inparams.jsonunderschedule, in/root/svccs/text/schedule.jsonwrite, for each time zone,{"utc": 목록, "gaps_hours": 이웃한 두 순간의 실제 간격(시간, 넷째 자리) 목록, "short_days": 24 미만인 간격 수, "long_days": 24 초과인 간격 수}(utc is the list, gaps_hours is the list of the actual gaps in hours between adjacent instants to four decimal places, short_days is the number of gaps under 24, and long_days is the number of gaps over 24). - In the same file, create
check_signup(nickname, taken, max_bytes, max_graphemes), reviewsignups.csvin order following the "Signup rule" below, and in/root/svccs/text/signups.jsonwrite the counts per statusok·taken·too_many_chars·too_many_bytes·emptyand averdictslist in line order (the limits arebyte_limit·max_graphemes).
Signup rule
shown = NFC(nickname) 저장·표시는 NFC 로 한다
1. shown 이 공백(str.isspace)·ZWJ·규칙 1의 결합류로만 되어 있거나 비었으면 "empty"
2. graphemes(shown) 의 개수가 max_graphemes 를 넘으면 "too_many_chars"
3. shown 의 UTF-8 바이트가 max_bytes 를 넘으면 "too_many_bytes"
4. username_key(nickname) 이 taken 안에 있으면 "taken"
5. 아니면 "ok"
taken 은 처음에 names.csv 모든 닉네임의 키이고, "ok" 를 받은 신청의 키가 차례로 더해진다.
채점기는 taken 을 목록(list)으로 넘길 수도 있다 — in 으로만 쓰세요.
Notes
- The grader imports
textkit.py. Put code that reads files and produces results in a separate script or inpython3 - <<'PY'. - The byte count is
len(s.encode("utf-8"))and the code point count islen(s). Normalization isunicodedata.normalize("NFC", s), and time zones arezoneinfo.ZoneInfo(tz)andastimezone. - Common mistakes: writing
len()as the byte count, decoding a byte slice witherrors="replace"so that U+FFFD is left at the end, judging duplicates without NFC, grouping daily by UTC date, building a reservation by addingtimedelta(hours=24)to a UTC instant, and subtracting aware datetimes of the same tzinfo and writing that the gap is always 24 hours. - The outputs disappear when the session ends. Keep them elsewhere if you need them.
Bytes and code points are different
Write the number of lines, the sum of UTF-8 bytes, the sum of code points, the number of lines where the two differ, and the maximum bytes for the nicknames in names.csv to /root/svccs/text/lengths.json as rows, utf8_bytes, code_points, rows_multibyte, and max_utf8_bytes.
A Python str is a sequence of code points, so len() is the number of code points. Bytes are the length of s.encode("utf-8"). A precomposed Hangul syllable is 3 bytes, and ASCII is 1 byte.
Count the characters a person sees
Create graphemes(s) in /root/svccs/text/textkit.py following the grapheme rule, and write graphemes, rows_cp_ne_graphemes, and longest_cluster_cp to /root/svccs/text/graphemes.json. The grader calls the function directly and compares it on variant strings that mix combining characters, ZWJ, and jamo.
You only need to decide whether to attach the current code point to the preceding grapheme. unicodedata.combining() returns 0 for variation selectors (U+FE0F) and ZWJ, so it alone is not enough: look at the category and the ranges together. A vowel jamo (V) after a leading consonant jamo (L) attaches, while a leading consonant (L) after a final consonant (T) starts a new grapheme.
Cut by the byte limit without breaking a character
Create cut_bytes(s, limit) in textkit.py, and with byte_limit write byte_limit, rows_over_limit, rows_naive_broken, and utf8_bytes_after_cut to /root/svccs/text/cut.json. The grader calls cut_bytes with variant strings and several limits.
In UTF-8 a code point is 1 to 4 bytes, and every continuation byte has the form 10xxxxxx. Add code points one at a time while adding their bytes, and stop just before you would exceed the limit. If you cut by bytes and then read with errors="replace", U+FFFD is left at the end.
Cut by grapheme
Create fit(s, max_bytes, max_graphemes) in textkit.py, and write max_bytes, max_graphemes, rows_changed, split_by_cut_bytes, and split_by_cp_slice to /root/svccs/text/fit.json. The grader calls fit with variant strings and several limits.
Add the graphemes split by graphemes(s) one at a time, and stop just before either of the two limits, grapheme count or bytes, would be exceeded. Even a result cut safely at code point level can leave only one person of a family emoji or pull an accent off, and that count is split_by_cut_bytes.
Decide on the same name
Create username_key(s) = NFKC(casefold(NFKC(s))) in textkit.py, and write distinct_raw, distinct_nfc, distinct_nfc_lower, distinct_nfc_casefold, distinct_key, collision_pairs, and collision_groups to /root/svccs/text/unique.json. The grader also calls username_key with variant strings.
NFC reconciles precomposed and jamo-decomposed forms, casefold folds ß into ss, and NFKC reconciles compatibility characters such as fullwidth characters, ligatures, and the Kelvin sign. Each time you add a rule, the number of distinct values should go down. The number of pairs is n·(n−1)/2 for every n with the same key.
Group UTC events by local date
Create local_date(ts_utc, tz) in textkit.py, group events.csv by the date in each user's time zone, and write events, by_local_date, by_utc_date, and events_on_other_date to /root/svccs/text/daily.json. The grader calls local_date with variant events in other time zones and dates.
Read ts_utc as a UTC aware datetime (the trailing Z means UTC), move it with astimezone(ZoneInfo(tz)), and then look at .date(). Seoul is ahead of UTC, so if you group by UTC date, Seoul's early morning slips into the previous day.
Turn a daily 09:00 local reservation into UTC
Create daily_at(tz, start_date, days, hhmm) in textkit.py, and with the schedule in params.json write utc, gaps_hours, short_days, and long_days for each time zone to /root/svccs/text/schedule.json. The grader calls daily_at with other time zones and periods.
If you convert the first instant to UTC and add 24 hours at a time, the local time shifts by an hour after a daylight saving transition. Build hhmm fresh for each local date and convert it to UTC. Subtract the UTC instants from each other for the gaps: if you subtract aware datetimes that share the same ZoneInfo, the tzinfo is ignored and you always get 24 hours.
Gather the signup review into one function
Create check_signup(nickname, taken, max_bytes, max_graphemes) in textkit.py following the signup rule, review signups.csv in order, and write ok, taken, too_many_chars, too_many_bytes, empty, and verdicts to /root/svccs/text/signups.json. The grader also calls check_signup with variant signups and taken lists.
The order of the checks is part of the rule. The byte limit is measured after converting to NFC: a Korean name that arrives in decomposed form measures three times larger if you measure the original as is. If you do not add the key of an accepted signup to taken, two signups that arrive on the same day both pass.