TT Lab
Get started
Learn Learning paths Courses

Insurance Domain Deep Dive

Rules Are the Easy Part — the Threshold Is Where You Get Stuck

Continue in TT Lab

In one line

What is hard in fraud detection is not thinking up the rules but deciding where to cut the score. The threshold is a knob that trades precision against recall, and what actually turns that knob is not the model but the daily throughput of the investigation staff.

Why this was needed

Detection rule meetings usually go similarly. Remarks pour out like "an accident within a month of signing the contract is suspicious" and "a filing more than two months late is suspicious," and everyone nods. Moving the rules into code is not hard either. And when you run it, 190 of 400 claims get flagged. The investigation team can look at only ten this quarter.

From here is the real work. You have to measure in numbers how useful each rule is, decide where to cut the score, and be able to say what cost that choice leaves on whom. A rulebook that skips this process keeps only the words "it catches suspicious things," and in reality makes investigators spend their time on false positives.

Let us make one thing clear first. What this lab deals with is the side that builds the detector. It does not explain which methods get away and how, and it concentrates on picking a threshold with the answer table planted in synthetic data. The authority and procedures for investigating and reporting insurance fraud are set by law (Insurance Business Act), and what a data team makes is not a judgment but a list of investigation targets and their basis.

How it works

First you produce a score for each rule. If you count how many cases a rule flagged and how many of those were confirmed incidents, you get that rule's precision on its own. In practice this number is usually disappointing. A rule that exceeds half on its own is rare. So you do not judge with a single rule but add them up with weights attached. The weights reflect both the rule's standalone precision and the business staff's judgment.

Next is the confusion matrix. Once you set a threshold, the claims split into four cells. Flagged and actually an incident (TP), flagged but not (FP), not flagged but an incident (FN), not flagged and not (TN). Three metrics come out of this.

정밀도(precision) = TP / (TP + FP)     걸린 것 중 맞은 비율   -> 조사역의 헛수고
재현율(recall)    = TP / (TP + FN)     사고 중 잡은 비율     -> 새어 나간 손해
F1               = 2PR / (P + R)      둘의 조화평균

The definitions are well laid out in the model evaluation documentation of scikit-learn. The library is not in this lab image, but the calculation is a few divisions, so the statistics module and basic arithmetic are enough, and values that must be exact, like amounts, are handled with decimal. If you need per-group aggregation or cumulative calculations, SQLite's window functions are worth getting used to.

If you build a table raising the threshold from 1 point, you can see at a glance the two metrics moving in opposite directions. Here the base rate is decisive. If 24 of 400 are incidents, the base rate is 6 percent. If you lower the threshold to 1 point, recall gets close to 1, but precision collapses to around 10 percent. That means nine of ten flagged cases are wasted effort. In a low-base-rate problem, precision gets worse very fast, and this is not because the rules are bad but arithmetic.

The last is capacity. Far more often than not, the point with the highest F1 is not the point you operate at. If the investigation team can look at only ten this quarter, the only choices are thresholds that flag ten or fewer. You pick the best point within that, and write down together where the best would have been if there had been no capacity. That difference becomes the basis when you ask for more staff.

What it looks like in the field

First, group skew. If a particular channel or region gets flagged unusually often, there are two possibilities. The incident rate of that group is really higher, or the rules picked up an operational characteristic unrelated to fraud. Telephone intake desks have frequent late filings because documents arrive late, and branch customers have a practice of splitting small claims into several submissions. Both are flagged equally by the rules. Looking only at the flag rate becomes discrimination, and the judgment holds only when you place the flag rate next to that group's base rate.

Second, the cost of false positives is invisible. One false positive is the number 1 in a table, but to the customer it is a request for additional documents and a payment delay. If you read precision only as "accuracy," this cost disappears.

Third, reproducibility. The threshold moves with the data. If you do not attach a version number and a data fingerprint to the rulebook, two months later you cannot answer "by which rules did that judgment come out then?"

What really matters in practice

What you will do in the next lab

You build 400 synthetic claims yourself and produce each of the five rules' scores. You combine the scores to make a judgment list, and calculate the confusion matrix and the three metrics by hand at the first proposed threshold. You build a table raising the threshold from 1 point to the end, place the per-channel flag rates and base rates side by side to check for skew, and then pick the threshold to operate within the investigation staff capacity. Finally you produce a rulebook with a version number and data fingerprint attached and a report for people to read.