In Front of an Unfamiliar System
The answers do not match the system - reconcile it field by field
Goal
You reconcile the nine intake answers received before going on site against the values measured on that customer's real system. You harden the verdict rules into a file, build a reconciler that judges into five verdicts (exact match, within range, mismatch, no answer, unverifiable), and have the machine generate the question to ask again for each mismatched field.
Why it matters
The intake answers are the customer's memory and the system is the fact. The person who wrote the answers is usually the one who installed that system three years ago, and since then the retention period has shrunk, one more integration has been attached, and the backup interval has changed. A plan that takes the answers as they are as its premise collapses as soon as the work starts. So you reconcile with rules, not by eye. If you decide in advance how to compare each field and write it in a file, you can rerun the same command two weeks later and see what changed in the meantime. A reconciliation done by eye cannot be rerun. The most important thing is to write the fields you do not know as not known. If you write a field you could not confirm as a match, nobody looks at that field again, and if you write it as a mismatch, you argue with the customer for nothing. Fields with no answer and fields that cannot be verified are counted separately and left as a list. The grader does not trust the wording you write down. It sets up answers and facts that it made itself in a temporary directory, actually executes your reconciler with different values each time, and checks the verdicts and the summary against the values it computes itself.
Steps
- Create and run /root/intake/gen_site.py to produce /root/intake/answers.json (nine fields) and /root/intake/site/ (the version file, the configuration, 14 days of logs, and the production DB).
- Measure the system directly and write seven fields in /root/intake/facts.json: app_version, timezone, retention_days, log_days, daily_orders_max, integrations, and backup_interval_hours.
- Decide the comparison method for each field and write it in /root/intake/rules.json as a fields list. A field answered as a range is range, a field answered as a list is set, a field answered as an approximate figure is tolerance (tolerance_pct 20), and the rest are exact. A field with no place to measure gets source set to null.
- Create /root/intake/reconcile.py so that it judges match_exact and mismatch with the two rules exact and set, and outputs them as JSON together with a summary.
- Add range and tolerance so that a field that falls within the range is judged match_in_range. A boundary value counts as within range.
- Make a field whose answer is empty be judged unanswered, and a field that has an answer but no measured value be judged unverifiable. A field that is both is unanswered.
- Add
--questions <경로>(the placeholder is the path) so that it writes a JSON array containing one question sentence to ask again for each field that did not match. - Produce /root/intake/report.json and /root/intake/questions.json with the real answers and the real facts, and report in four sections in /root/intake/intake_report.md.
Notes
- Execution contract:
python3 /root/intake/reconcile.py --answers <A> --facts <F> --rules <R> --out <보고 JSON> [--questions <질문 JSON>](the placeholders are the answers, facts, and rules files, the report JSON, and the questions JSON) writes a one-line summary JSON to standard output and ends with exit code 0. If it cannot read the input, the exit code is 3. - Report JSON:
{"summary": {"fields": 정수, "match_exact": 정수, "match_in_range": 정수, "mismatch": 정수, "unanswered": 정수, "unverifiable": 정수}, "findings": [...]}(the Korean word in the placeholders means "integer"). findings is sorted by field name ascending, and each entry has five keys: field, rule, answer, fact, and verdict. - The verdict names are exactly five: match_exact, match_in_range, mismatch, unanswered, and unverifiable.
- Rules file:
{"fields": [{"name": …, "kind": "exact|range|tolerance|set", "source": 문자열 또는 null, "tolerance_pct": 정수}]}(the Korean words in the code mean "string or null" and "integer"). For a field that is not tolerance, you do not need to set tolerance_pct. - range is a field whose answer is
{"min": …, "max": …}, and it includes the boundary values. tolerance counts as within range when잰 값과 답변의 차이 <= 답변 × tolerance_pct / 100(the Korean text means "difference between the measured value and the answer <= answer × tolerance_pct / 100"). If the measured value is exactly equal to the answer, it is match_exact even for a tolerance field. - Measure the values of the facts file yourself. The version is
site/app/VERSION, the configuration issite/app/app.ini(configparser), the retention period and time zone are in the settings table ofsite/data/app.db, the log retention days come from the file names insite/logs, the maximum daily orders come from the orders table, and the backup interval is the time gap in the backup_log table. - Common mistakes: comparing a list including its order, hiding the tolerance in the code, writing a field you could not confirm as a match, and giving an accusing sentence instead of a question sentence.
- tolerance_pct 20 and the priority "a field with no answer comes before unverifiable" are assumptions of this lab. They are not values a standard decides but values the team agrees on and writes in the rules file.
- Reference documents: RFC 2119 and RFC 8174 describe how to fix the strength of a rule sentence with words, RFC 3339 describes the notation of dates and times, and the Python json documentation and the SQLite date functions describe the tools used in this lab.
Get the answers and the system in hand
Create and run /root/intake/gen_site.py to produce /root/intake/answers.json (nine fields) and /root/intake/site/. site holds the version file, app.ini, 14 days of logs, and the production DB (settings, orders, backup_log).
What you hold on site is only one answers file and one system. Here we make those two ourselves. First create /root/intake, and inside it make the files and the sqlite DB with python3. The answers are values the customer wrote from memory, so they differ from the system in several fields.
Measure directly on the system
Write seven fields in /root/intake/facts.json: app_version, timezone, retention_days, log_days, daily_orders_max, integrations, and backup_interval_hours. All the values must be measured directly from site/, and do not put any field other than the seven.
The version is the single line of site/app/VERSION, the time zone and the retention period are in the settings table of site/data/app.db, and the integration list is in the integrations section of site/app/app.ini. log_days is the number of distinct dates in site/logs, daily_orders_max is the maximum of the orders table counted by date, and backup_interval_hours is the gap between adjacent times in backup_log.
Harden the verdict rules into a file
In /root/intake/rules.json, write all nine answer fields as a fields list. A field answered as a range gets kind range, a field answered as a list gets set, a field answered as an approximate figure (daily_orders_max) gets tolerance with tolerance_pct 20, and the rest are exact. For a field with no measured value, set source to null; for a field with one, write as a string where you measured it.
The rules must be decided before you reconcile. If you decide them while reconciling, you get the result you want to see. Which field has no place to measure depends on whether its name is in facts.json. In source, write evidence such as a path or a table name that you can later show to the customer.
Separate exact match from mismatch
Create /root/intake/reconcile.py so that it judges with the two rules exact and set. If the values are equal it is match_exact, and if they differ it is mismatch. set does not look at order. The report JSON holds summary and findings.
Sort findings by field name ascending and put the five keys field, rule, answer, fact, and verdict in each entry. summary counts each of the five verdict names and puts the total number of fields in fields. If you sort before comparing lists, they are not shaken by order.
Count within range separately
Add range and tolerance so that a field that falls within the range is judged match_in_range. range is a field whose answer is min and max, and it includes the boundary values. tolerance applies when the difference is within tolerance_pct percent of the answer, and if the measured value is exactly equal to the answer, it is match_exact.
If you write all the approximate-figure fields as mismatches, the real problem is buried in the same color. Whether to include or exclude the boundary values is decided by the rules, and this lab includes them. Do not hide the tolerance in the code; read and use tolerance_pct from the rules file.
Fields with no answer and unverifiable fields
Make a field whose answer is empty or entirely absent be judged unanswered, and a field that has an answer but no measured value be judged unverifiable. A field that is both is unanswered.
If you write a field you could not confirm as a match, nobody looks at that field again, and if you write it as a mismatch, you argue with the customer for nothing. Treat the case where the key is entirely absent from the answers the same as the case where the value is null. If you put the priority at the very front of the code, the remaining rules cannot touch these two.
Generate questions instead of accusations
Add --questions <경로> (the placeholder is the path) so that it writes a JSON array holding one entry for each field that did not match. Each entry contains field, verdict, answer, and fact, along with one ask sentence. ask contains the field name and must end with a question mark.
If you walk in holding a table saying twelve fields are wrong, the customer becomes defensive first. If you turn the same content into questions, it becomes a conversation. Keep a separate sentence template for each kind of verdict — for a mismatch, ask which side is right; for unverifiable, ask where to look; for a field with no answer, ask who decides the value.
Run it with the real answers and report
Produce /root/intake/report.json and /root/intake/questions.json with the real answers and the real facts, and write /root/intake/intake_report.md in four sections: ## 무엇을 대조했나 ## 어긋난 칸 ## 답이 없는 칸 ## 다시 물어볼 것 (the Korean headings mean "What was reconciled", "The mismatched fields", "The unanswered fields", and "What to ask again"). The names of the mismatched fields and the unanswered fields must all appear in the report.
Do not write the report by hand; generate it from report.json and questions.json. That way, when you rerun the same command two weeks later, the report is updated along with it. Take the numbers as they are from the summary, and take the field names from findings.