TT Lab
Get started
Learn Learning paths Courses

In Front of an Unfamiliar System

The answers do not match the system - reconcile it field by field

Continue in TT Lab

Goal

You reconcile the nine intake answers received before going on site against the values measured on that customer's real system. You harden the verdict rules into a file, build a reconciler that judges into five verdicts (exact match, within range, mismatch, no answer, unverifiable), and have the machine generate the question to ask again for each mismatched field.

Why it matters

The intake answers are the customer's memory and the system is the fact. The person who wrote the answers is usually the one who installed that system three years ago, and since then the retention period has shrunk, one more integration has been attached, and the backup interval has changed. A plan that takes the answers as they are as its premise collapses as soon as the work starts. So you reconcile with rules, not by eye. If you decide in advance how to compare each field and write it in a file, you can rerun the same command two weeks later and see what changed in the meantime. A reconciliation done by eye cannot be rerun. The most important thing is to write the fields you do not know as not known. If you write a field you could not confirm as a match, nobody looks at that field again, and if you write it as a mismatch, you argue with the customer for nothing. Fields with no answer and fields that cannot be verified are counted separately and left as a list. The grader does not trust the wording you write down. It sets up answers and facts that it made itself in a temporary directory, actually executes your reconciler with different values each time, and checks the verdicts and the summary against the values it computes itself.

Steps

  1. Create and run /root/intake/gen_site.py to produce /root/intake/answers.json (nine fields) and /root/intake/site/ (the version file, the configuration, 14 days of logs, and the production DB).
  2. Measure the system directly and write seven fields in /root/intake/facts.json: app_version, timezone, retention_days, log_days, daily_orders_max, integrations, and backup_interval_hours.
  3. Decide the comparison method for each field and write it in /root/intake/rules.json as a fields list. A field answered as a range is range, a field answered as a list is set, a field answered as an approximate figure is tolerance (tolerance_pct 20), and the rest are exact. A field with no place to measure gets source set to null.
  4. Create /root/intake/reconcile.py so that it judges match_exact and mismatch with the two rules exact and set, and outputs them as JSON together with a summary.
  5. Add range and tolerance so that a field that falls within the range is judged match_in_range. A boundary value counts as within range.
  6. Make a field whose answer is empty be judged unanswered, and a field that has an answer but no measured value be judged unverifiable. A field that is both is unanswered.
  7. Add --questions <경로> (the placeholder is the path) so that it writes a JSON array containing one question sentence to ask again for each field that did not match.
  8. Produce /root/intake/report.json and /root/intake/questions.json with the real answers and the real facts, and report in four sections in /root/intake/intake_report.md.

Notes

Get the answers and the system in hand

Create and run /root/intake/gen_site.py to produce /root/intake/answers.json (nine fields) and /root/intake/site/. site holds the version file, app.ini, 14 days of logs, and the production DB (settings, orders, backup_log).

What you hold on site is only one answers file and one system. Here we make those two ourselves. First create /root/intake, and inside it make the files and the sqlite DB with python3. The answers are values the customer wrote from memory, so they differ from the system in several fields.

Measure directly on the system

Write seven fields in /root/intake/facts.json: app_version, timezone, retention_days, log_days, daily_orders_max, integrations, and backup_interval_hours. All the values must be measured directly from site/, and do not put any field other than the seven.

The version is the single line of site/app/VERSION, the time zone and the retention period are in the settings table of site/data/app.db, and the integration list is in the integrations section of site/app/app.ini. log_days is the number of distinct dates in site/logs, daily_orders_max is the maximum of the orders table counted by date, and backup_interval_hours is the gap between adjacent times in backup_log.

Harden the verdict rules into a file

In /root/intake/rules.json, write all nine answer fields as a fields list. A field answered as a range gets kind range, a field answered as a list gets set, a field answered as an approximate figure (daily_orders_max) gets tolerance with tolerance_pct 20, and the rest are exact. For a field with no measured value, set source to null; for a field with one, write as a string where you measured it.

The rules must be decided before you reconcile. If you decide them while reconciling, you get the result you want to see. Which field has no place to measure depends on whether its name is in facts.json. In source, write evidence such as a path or a table name that you can later show to the customer.

Separate exact match from mismatch

Create /root/intake/reconcile.py so that it judges with the two rules exact and set. If the values are equal it is match_exact, and if they differ it is mismatch. set does not look at order. The report JSON holds summary and findings.

Sort findings by field name ascending and put the five keys field, rule, answer, fact, and verdict in each entry. summary counts each of the five verdict names and puts the total number of fields in fields. If you sort before comparing lists, they are not shaken by order.

Count within range separately

Add range and tolerance so that a field that falls within the range is judged match_in_range. range is a field whose answer is min and max, and it includes the boundary values. tolerance applies when the difference is within tolerance_pct percent of the answer, and if the measured value is exactly equal to the answer, it is match_exact.

If you write all the approximate-figure fields as mismatches, the real problem is buried in the same color. Whether to include or exclude the boundary values is decided by the rules, and this lab includes them. Do not hide the tolerance in the code; read and use tolerance_pct from the rules file.

Fields with no answer and unverifiable fields

Make a field whose answer is empty or entirely absent be judged unanswered, and a field that has an answer but no measured value be judged unverifiable. A field that is both is unanswered.

If you write a field you could not confirm as a match, nobody looks at that field again, and if you write it as a mismatch, you argue with the customer for nothing. Treat the case where the key is entirely absent from the answers the same as the case where the value is null. If you put the priority at the very front of the code, the remaining rules cannot touch these two.

Generate questions instead of accusations

Add --questions <경로> (the placeholder is the path) so that it writes a JSON array holding one entry for each field that did not match. Each entry contains field, verdict, answer, and fact, along with one ask sentence. ask contains the field name and must end with a question mark.

If you walk in holding a table saying twelve fields are wrong, the customer becomes defensive first. If you turn the same content into questions, it becomes a conversation. Keep a separate sentence template for each kind of verdict — for a mismatch, ask which side is right; for unverifiable, ask where to look; for a field with no answer, ask who decides the value.

Run it with the real answers and report

Produce /root/intake/report.json and /root/intake/questions.json with the real answers and the real facts, and write /root/intake/intake_report.md in four sections: ## 무엇을 대조했나 ## 어긋난 칸 ## 답이 없는 칸 ## 다시 물어볼 것 (the Korean headings mean "What was reconciled", "The mismatched fields", "The unanswered fields", and "What to ask again"). The names of the mismatched fields and the unanswered fields must all appear in the report.

Do not write the report by hand; generate it from report.json and questions.json. That way, when you rerun the same command two weeks later, the report is updated along with it. Take the numbers as they are from the summary, and take the field names from findings.