The Validator Is Green and the Numbers Still Look Wrong
Goal
You build profile.py, a tool that finds data whose format is right but whose values are strange, using a per-column profile, distribution changes between deliveries, and the first-digit distribution. You put in a column where Benford does not work as a counterexample, and make the code itself say that this is a lead and not evidence.
Why it matters
There are cases where the schema is right and the missing values are as the rules say, yet the numbers are strange. Then "it's strange" alone does not move anyone. You need grounds, and those grounds are not a single value but a distribution. A single profile alone makes judgment hard. This is because strangeness shows up not as an absolute value but as a difference from the previous delivery. So you keep the profile for every delivery and compare, and you write the thresholds in the code — a threshold that is not written down cannot be fixed by the next person. The first-digit distribution is a cheap signal that works well, but it is also the most misunderstood signal. That the distribution is off means a person should look, not that someone tampered with it, and in a column with a narrow value range even honest data is always off. So you must judge applicability first. The grader does not trust your wording. It sets up deliveries it made in a temporary directory, actually runs your tool, and checks the profile and the judgments against the values it computed itself. The counterparties and amounts change on every run.
Steps
- Create and run /root/prof/gen_feed.py to produce
2026-04.csv,2026-05.csv, and2026-06.csvunder /root/prof/feed. All three pass the format validation. - Build
profile <파일>(the placeholder is the file) in /root/prof/profile.py so that for each column it outputs the type, missing, unique ratio, value length, mode, and range. - Add
digits <파일> <칼럼>so that it outputs the value length distribution and the values that deviate from the modal length. - Add
drift <파일A> <파일B>so that it judges the distribution change between deliveries by thresholds. - Add
benford <파일> <칼럼>so that it outputs the difference between the first-digit distribution and the expected distribution, compare two deliveries, and write it in /root/prof/benford.json. - Add an applicability judgment to
benfordso that for a column with a narrow value range,applicableis false and the verdict isnot_applicable. Write the result in /root/prof/narrow.json. - Overlay several signals to build a list of investigation targets, and write what to check next for each column in /root/prof/leads.json.
- Summarize everything on one sheet and produce /root/prof/prof_report.json and /root/prof/prof_report.md.
Notes
- Execution contract:
python3 /root/prof/profile.py <명령> <파일> [인자](the placeholders are the command, the file, and the arguments). There are four commands: profile, digits, drift, and benford. On success the exit code is 0, if the file does not exist it is 3, and if the command or the number of arguments is wrong it is 2. - The notations counted as missing are the blank cell and
NULL,NA,N/A, and-, and case is not distinguished. profileresponse:{"rows": 정수, "columns": {칼럼: {"type":…, "nulls":…, "null_rate":…, "distinct":…, "unique_ratio":…, "len_min":…, "len_max":…, "modal_length":…, "top": [[값, 건수], …3개]}}}(the Korean words in the code mean "integer", "column", "value", "count", and "3 items").minandmaxare added for numeric columns. Ratios are rounded at the fourth decimal place.topis in descending order of count, and when counts are equal, in dictionary order of the value. The denominator of a ratio is the total number of rows, and lengths and distinct values are counted only over the values other than missing.digitsresponse:column,rows,lengths(the counts keyed by length as a string),modal_length,outliers, andsample_outliers(the first 3 of the distinct values, sorted).driftresponse:{"columns": {칼럼: {…_before, …_after, "flags": [...]}}, "flagged": [칼럼...]}(the Korean word in the code means "column"). It looks at four things —null_rateif the missing ratio difference is 0.05 or more,unique_ratioif the unique ratio difference is 0.20 or more,modal_lengthif the modal length differs, andtop_shareif the difference in the share of the mode is 0.10 or more. These thresholds are assumptions of this lab.benfordresponse:column,n,observed,expected,max_deviation,deviation_digit, andverdict. From step 6,spanandapplicableare added. The expected ratio for digit d islog10(1 + 1/d), rounded at the fourth decimal place.- The judgment rules are also an assumption of this lab — it is applicable only if there are at least 200 values and the maximum-to-minimum ratio is at least 100; when applicable, if the maximum deviation is 0.05 or more it is
suspect, otherwiseordinary, and when not applicable it isnot_applicable. - The first digit is the first nonzero digit encountered, skipping the sign, leading zeros, and the decimal point. If the value is 0, it is not counted.
- Official documents: NIST chi-square goodness-of-fit test · python statistics · python math
- Background (not a standard document): Benford's law — it is for skimming the name and the history, and the basis for the judgment procedure is the NIST document above.
- Common mistakes: writing the Benford result as if it were evidence, applying it as is to a column with a narrow value range, leaving the thresholds only in the code and not writing them in a document, and reaching a conclusion from a single signal.
Build three months of deliveries
Create and run /root/prof/gen_feed.py to produce 2026-04.csv, 2026-05.csv, and 2026-06.csv under /root/prof/feed. Two months are ordinary, and in one month the format is right but the values are strange.
Make the amounts so that the number of digits is evenly spread — if you draw the exponent uniformly and make it a power of 10, the first-digit distribution becomes natural. In the tampered month, mix in a few amounts that look hand-written by a person, and put in together some lines with an empty counterparty and some lines with a different phone number digit count. If you use a fixed-seed generator instead of the standard library's random numbers, the same file comes out whoever runs it.
Have each column introduce itself
Build profile <파일> (the placeholder is the file) in /root/prof/profile.py so that it outputs rows and columns. For each column include type, nulls, null_rate, distinct, unique_ratio, len_min, len_max, modal_length, and top, and add min and max for numeric columns.
The denominator of a ratio is the total number of rows, and lengths and distinct values are counted only over the values other than missing. top must be in descending order of count but, when counts are equal, in dictionary order of the value, so that the same answer comes out every run.
Pull out and look at what deviates from the value length
Add digits <파일> <칼럼> (the placeholders are the file and the column) so that it outputs lengths, modal_length, outliers, and sample_outliers. sample_outliers is the first 3 of the distinct values, sorted.
Even for values that pass the format check, if the length differs, the producer has changed. If you pull out a few of the deviating values and look at them directly, the cause is usually visible right away — for a phone number, the area code dropped out or the hyphens vanished.
Compare with the previous delivery
Add drift <파일A> <파일B> (the placeholders are the two files) so that for each column it measures four kinds of change and puts those that exceed a threshold in flags. The response also includes flagged, the list of columns that have at least one flag.
Strangeness shows up not as an absolute value but as a difference. Look at the missing ratio at 0.05, the unique ratio at 0.20, the share of the mode at 0.10, and whether the modal length changed. If you pull the thresholds out as constants, there is one place to fix when the data changes later.
Measure the first-digit distribution
Add benford <파일> <칼럼> (the placeholders are the file and the column) so that it outputs observed, expected, max_deviation, deviation_digit, and verdict, and write the amount results for the ordinary month and the tampered month in /root/prof/benford.json as normal and suspect.
The expected ratio for digit d is log10(1 + 1/d). The first digit is the first nonzero digit encountered, skipping the sign, leading zeros, and the decimal point, and a value of 0 is not counted. Amounts a person made up by hand cluster at particular digits, so the difference widens greatly.
A column where Benford does not work
Add span and applicable to benford, and if there are fewer than 200 values or the maximum-to-minimum ratio is under 100, treat it as not applicable and make the verdict not_applicable. Write the score column result in /root/prof/narrow.json as column, n, span, applicable, verdict, and why.
If the values lie only in a narrow interval, the first digit is restricted to a few possibilities, and even honest data deviates greatly from the expected distribution. A Benford run without an applicability judgment keeps reporting perfectly fine data, and then nobody reads those reports. In why, write in a sentence or two why it cannot be applied.
Overlay signals to make investigation targets
Overlay the distribution changes and the Benford results and write leads, not_leads, and note in /root/prof/leads.json. Each entry of leads holds column, signals, and next_check, and not_leads lists the columns to which Benford cannot be applied.
A column caught by only one signal and a column caught by three must be handled differently. next_check must not be 'it is strange' but 'what to check next' — for example, which staff member the hand-written-looking amounts are concentrated on, or which system the phone numbers with changed digit counts came from.
Hand over the leads on one sheet
Write baseline, target, rows, flagged, benford, not_applicable, and leads in /root/prof/prof_report.json, and write /root/prof/prof_report.md in four sections: ## 무엇을 봤나 ## 지난 전달분과 무엇이 달라졌나 ## 첫 자리 숫자가 말해 주는 것 ## 무엇을 더 확인해야 하나 (the Korean headings mean "What we looked at", "What changed from the previous delivery", "What the first digits tell us", and "What else needs checking").
benford is an object holding the column, max_deviation, and verdict of the amount column. In the third section of the report, write the Benford result and, without fail, why it is not evidence, and in the fourth section, write what to check next for each column. You can just read and combine the files you left in the earlier steps.