TT Lab
Get started
Learn Learning paths Courses

Working With Customer Data

The Validator Is Green and the Numbers Still Look Wrong

Continue in TT Lab

Goal

You build profile.py, a tool that finds data whose format is right but whose values are strange, using a per-column profile, distribution changes between deliveries, and the first-digit distribution. You put in a column where Benford does not work as a counterexample, and make the code itself say that this is a lead and not evidence.

Why it matters

There are cases where the schema is right and the missing values are as the rules say, yet the numbers are strange. Then "it's strange" alone does not move anyone. You need grounds, and those grounds are not a single value but a distribution. A single profile alone makes judgment hard. This is because strangeness shows up not as an absolute value but as a difference from the previous delivery. So you keep the profile for every delivery and compare, and you write the thresholds in the code — a threshold that is not written down cannot be fixed by the next person. The first-digit distribution is a cheap signal that works well, but it is also the most misunderstood signal. That the distribution is off means a person should look, not that someone tampered with it, and in a column with a narrow value range even honest data is always off. So you must judge applicability first. The grader does not trust your wording. It sets up deliveries it made in a temporary directory, actually runs your tool, and checks the profile and the judgments against the values it computed itself. The counterparties and amounts change on every run.

Steps

  1. Create and run /root/prof/gen_feed.py to produce 2026-04.csv, 2026-05.csv, and 2026-06.csv under /root/prof/feed. All three pass the format validation.
  2. Build profile <파일> (the placeholder is the file) in /root/prof/profile.py so that for each column it outputs the type, missing, unique ratio, value length, mode, and range.
  3. Add digits <파일> <칼럼> so that it outputs the value length distribution and the values that deviate from the modal length.
  4. Add drift <파일A> <파일B> so that it judges the distribution change between deliveries by thresholds.
  5. Add benford <파일> <칼럼> so that it outputs the difference between the first-digit distribution and the expected distribution, compare two deliveries, and write it in /root/prof/benford.json.
  6. Add an applicability judgment to benford so that for a column with a narrow value range, applicable is false and the verdict is not_applicable. Write the result in /root/prof/narrow.json.
  7. Overlay several signals to build a list of investigation targets, and write what to check next for each column in /root/prof/leads.json.
  8. Summarize everything on one sheet and produce /root/prof/prof_report.json and /root/prof/prof_report.md.

Notes

Build three months of deliveries

Create and run /root/prof/gen_feed.py to produce 2026-04.csv, 2026-05.csv, and 2026-06.csv under /root/prof/feed. Two months are ordinary, and in one month the format is right but the values are strange.

Make the amounts so that the number of digits is evenly spread — if you draw the exponent uniformly and make it a power of 10, the first-digit distribution becomes natural. In the tampered month, mix in a few amounts that look hand-written by a person, and put in together some lines with an empty counterparty and some lines with a different phone number digit count. If you use a fixed-seed generator instead of the standard library's random numbers, the same file comes out whoever runs it.

Have each column introduce itself

Build profile <파일> (the placeholder is the file) in /root/prof/profile.py so that it outputs rows and columns. For each column include type, nulls, null_rate, distinct, unique_ratio, len_min, len_max, modal_length, and top, and add min and max for numeric columns.

The denominator of a ratio is the total number of rows, and lengths and distinct values are counted only over the values other than missing. top must be in descending order of count but, when counts are equal, in dictionary order of the value, so that the same answer comes out every run.

Pull out and look at what deviates from the value length

Add digits <파일> <칼럼> (the placeholders are the file and the column) so that it outputs lengths, modal_length, outliers, and sample_outliers. sample_outliers is the first 3 of the distinct values, sorted.

Even for values that pass the format check, if the length differs, the producer has changed. If you pull out a few of the deviating values and look at them directly, the cause is usually visible right away — for a phone number, the area code dropped out or the hyphens vanished.

Compare with the previous delivery

Add drift <파일A> <파일B> (the placeholders are the two files) so that for each column it measures four kinds of change and puts those that exceed a threshold in flags. The response also includes flagged, the list of columns that have at least one flag.

Strangeness shows up not as an absolute value but as a difference. Look at the missing ratio at 0.05, the unique ratio at 0.20, the share of the mode at 0.10, and whether the modal length changed. If you pull the thresholds out as constants, there is one place to fix when the data changes later.

Measure the first-digit distribution

Add benford <파일> <칼럼> (the placeholders are the file and the column) so that it outputs observed, expected, max_deviation, deviation_digit, and verdict, and write the amount results for the ordinary month and the tampered month in /root/prof/benford.json as normal and suspect.

The expected ratio for digit d is log10(1 + 1/d). The first digit is the first nonzero digit encountered, skipping the sign, leading zeros, and the decimal point, and a value of 0 is not counted. Amounts a person made up by hand cluster at particular digits, so the difference widens greatly.

A column where Benford does not work

Add span and applicable to benford, and if there are fewer than 200 values or the maximum-to-minimum ratio is under 100, treat it as not applicable and make the verdict not_applicable. Write the score column result in /root/prof/narrow.json as column, n, span, applicable, verdict, and why.

If the values lie only in a narrow interval, the first digit is restricted to a few possibilities, and even honest data deviates greatly from the expected distribution. A Benford run without an applicability judgment keeps reporting perfectly fine data, and then nobody reads those reports. In why, write in a sentence or two why it cannot be applied.

Overlay signals to make investigation targets

Overlay the distribution changes and the Benford results and write leads, not_leads, and note in /root/prof/leads.json. Each entry of leads holds column, signals, and next_check, and not_leads lists the columns to which Benford cannot be applied.

A column caught by only one signal and a column caught by three must be handled differently. next_check must not be 'it is strange' but 'what to check next' — for example, which staff member the hand-written-looking amounts are concentrated on, or which system the phone numbers with changed digit counts came from.

Hand over the leads on one sheet

Write baseline, target, rows, flagged, benford, not_applicable, and leads in /root/prof/prof_report.json, and write /root/prof/prof_report.md in four sections: ## 무엇을 봤나 ## 지난 전달분과 무엇이 달라졌나 ## 첫 자리 숫자가 말해 주는 것 ## 무엇을 더 확인해야 하나 (the Korean headings mean "What we looked at", "What changed from the previous delivery", "What the first digits tell us", and "What else needs checking").

benford is an object holding the column, max_deviation, and verdict of the amount column. In the third section of the report, write the Benford result and, without fail, why it is not evidence, and in the fourth section, write what to check next for each column. You can just read and combine the files you left in the earlier steps.