TT Lab
Get started
Learn Learning paths Courses

Insurance Domain Deep Dive

Five Rules, One Scorecard, and a Threshold You Can Actually Staff

Continue in TT Lab

Goal

You build a rule scoreboard on claims data and calculate precision, recall, and F1 yourself from the confusion matrix. You build a table moving the threshold, check group skew, and then pick the threshold to operate within the investigation staff capacity and produce it as a rulebook with a version number and data fingerprint attached.

Why it matters

Thinking up detection rules is not hard. What is hard is deciding where to cut the score and saying what cost that choice leaves on whom. In a low-base-rate problem, lowering the threshold even slightly collapses precision quickly. And it is rare for the point with the highest F1 to be the point you operate at. The number of cases the investigation team can look at is the real constraint. This lab deals only with the side that builds the detector. It does not cover real fraud methods or how to evade them. All the data is synthetic, the organization names are fictional, and the answer table (label_fraud) is planted to pick the threshold. The grader recomputes the expected values from the DB you built and checks against them. Write decimals to the fourth place, and the integer cells must match exactly.

Steps

  1. Create and run /root/fraud/make_claims.py to put 200 contracts and 400 claims (24 confirmed fraud) into /root/fraud/claims.db.
  2. Gather five rules and their weights in /root/fraud/rules.py, and write each rule's score to /root/fraud/rules.csv.
  3. Combine the scores and write the claims at 1 point or more to /root/fraud/scores.csv along with which rules fired.
  4. At the first proposed threshold of 4 points, write the confusion matrix and the three metrics to /root/fraud/confusion.json.
  5. Write the metrics for thresholds from 1 point to 11 points to /root/fraud/sweep.csv.
  6. At the threshold of 4 points, write the per-channel flag rate and base rate to /root/fraud/fairness.csv, and the most-flagged group and the lowest-precision group to /root/fraud/fairness.json.
  7. Pick the threshold within the investigation staff capacity of 10 cases and write it with its basis to /root/fraud/threshold_choice.json.
  8. Produce the rulebook in /root/fraud/rulebook.json and a summary for people to read in /root/fraud/fraud_report.md.

Notes

Generate the synthetic claims data

Create and run /root/fraud/make_claims.py to put 200 policy rows and 400 claim rows (24 with label_fraud of 1) into /root/fraud/claims.db.

Use random.Random with the seed value 20260917, and number contracts from P0001 and claims from C0001. Make the filing date by adding a delay in days to the accident date. The grader even checks the amount total.

Produce the scores of the five rules

Gather the rules and weights in /root/fraud/rules.py, and write each rule's hits, fraud_hits, and precision to /root/fraud/rules.csv.

The header is rule_id,weight,hits,fraud_hits,precision, written in the order R1 to R5. precision is fraud_hits divided by hits. A rule having low standalone precision does not mean the rule is wrong - it means it cannot serve as the basis alone.

Combine the scores to build the judgment list

Write the claims at 1 point or more to /root/fraud/scores.csv in descending score order, and in ascending claim_id order for the same score.

In the rules cell, write the rules that fired, in order from R1, joined with '+'. A list with only scores cannot be reviewed by an investigator - what caused the flag must be there too.

Calculate the confusion matrix and the three metrics

At the threshold of 4 points, write TP, FP, FN, TN and precision, recall, and F1 to /root/fraud/confusion.json. Also include total, positives, base_rate, and flagged.

TN is the total minus the number flagged and the number missed. Precision is the share of flagged cases that were right, and recall is the share of incidents caught. Look at what the base rate is first and then read the precision.

Build a table moving the threshold

Write the flagged, tp, fp, fn, tn, precision, recall, and f1 for each threshold from 1 point to 11 points to /root/fraud/sweep.csv.

The lower the threshold, the higher the recall and the lower the precision. In data with a base rate of 6 percent, see for yourself how far precision goes if you lower it to 1 point. If you calculate only one point, you cannot tell what you gave up.

See per-group flag rates and base rates side by side

At the threshold of 4 points, write the per-channel claims, fraud, base_rate, flagged, flag_rate, tp, and precision to /root/fraud/fairness.csv, and the most-flagged group and the lowest-precision group to /root/fraud/fairness.json.

With only the flag rate, you cannot tell which group is being wronged. Only when you place it side by side with that group's confirmed fraud ratio do "are there really more?" and "did the rules pick up an operational characteristic?" separate. Write channels in ascending group name order.

Pick the threshold within the investigation staff capacity

Take the 10 cases the investigation team can look at this quarter as the constraint, pick the threshold, and write it with its basis to /root/fraud/threshold_choice.json. Also leave the best choice you would have had with no capacity limit.

Among the thresholds that do not exceed the capacity, pick the one with the highest f1. You must write together where the best would have been with no capacity limit, so that you can use that difference as the basis when you ask for more staff.

Export the rulebook with a version number and fingerprint

Put version, threshold, capacity, rules, dataset (including the fingerprint), and metrics in /root/fraud/rulebook.json, and summarize in four sections in /root/fraud/fraud_report.md.

The report section titles are '## Rulebook and evidence', '## How the threshold was chosen', '## Per-group check', and '## Limits and next steps'. How to compute the fingerprint is written in the Notes of the lab instructions. In the limits section, write in numbers what happens to precision when the base rate is low.