Clinic 08
Threshold Under Asymmetric Cost
A missed fraud costs 100 times as much as a false alarm. Compare operating points using their error costs.
Situation¶
A fraud system assigns scores that are not assumed calibrated. Missing a fraud costs $10,000; a false alarm costs $100. Correct decisions have zero additional cost in this exercise. Choose among the five tested thresholds using validation data, then evaluate the locked policy on untouched test data.
Artifact Packet¶
This fixed illustrative packet has a 5% fraud rate: 50 positives and 950 negatives per 1,000 transactions. Fractional counts represent rates scaled from a larger evaluation set. Every metric is calculated from the same TP, FN, FP, and TN counts.
| threshold | precision | recall | FPR | accuracy | cost per 1,000 |
|---|---|---|---|---|---|
| 0.50 | 0.706 | 0.610 | 0.013 | 0.9678 | $196,270 |
| 0.30 | 0.478 | 0.790 | 0.045 | 0.9463 | $109,320 |
| 0.15 | 0.265 | 0.910 | 0.133 | 0.8692 | $57,630 |
| 0.10 | 0.172 | 0.950 | 0.240 | 0.7694 | $47,810 |
| 0.05 | 0.089 | 0.980 | 0.527 | 0.4986 | $60,040 |
Underlying confusion counts:
| threshold | TP | FN | FP | TN |
|---|---|---|---|---|
| 0.50 | 30.5 | 19.5 | 12.7 | 937.3 |
| 0.30 | 39.5 | 10.5 | 43.2 | 906.8 |
| 0.15 | 45.5 | 4.5 | 126.3 | 823.7 |
| 0.10 | 47.5 | 2.5 | 228.1 | 721.9 |
| 0.05 | 49 | 1 | 500.4 | 449.6 |
cost = 10,000 × FN + 100 × FP. Accuracy is (TP + TN) / 1,000.
Decision Prompt¶
- Which tested threshold minimizes cost?
- Why does the accuracy winner lose on cost?
- Why does lowering the threshold from 0.10 to 0.05 increase cost?
- What changes if the scores become calibrated probabilities?
Strong Reasoning Looks Like¶
- compare the stated costs across the tested grid
- distinguish empirical score tuning from a population rule for calibrated probabilities
- check prevalence, calibration, review capacity, and cost changes before deployment
Run The Clinic In Browser¶
The runner prints this fixed illustrative packet and recalculates any derived columns. It does not run a new training experiment. Edit its PACKET values to explore the decision.
Validate Your Decision In Browser¶
The validator checks the original packet's grid choice; review the written explanation separately.
Reference Reveal¶
Open after writing your note
Choose **0.10**, costing **$47,810 per 1,000**: $25,000 from missed fraud and $22,810 from false alarms. At 0.50, accuracy is 0.9678 but cost is $196,270. At 0.05, the extra false alarms outweigh the savings from fewer misses. These are conclusions about the tested grid, not a proof of a global optimum. If `p` is a calibrated fraud probability for the deployment population, predicting fraud costs `100(1-p)` in expectation and predicting legitimate costs `10,000p`. Under these assumptions, choose fraud when `p ≥ 100 / 10,100 ≈ 0.00990`. That probability rule does not apply directly to arbitrary model scores. Capacity limits, additional costs, or a changed population require a revised decision rule. Lower false-negative costs or higher false-positive costs favor a higher calibrated-probability threshold. Verify changes empirically on validation data.What To Do Next¶
After this clinic:
- open Calibration and Thresholds
- run the matching threshold demo example
- use Imbalanced Triage and Review Budgets for the full cost-aware workflow