Clinic 09
Metric Choice Under An Unusual Task
The grader uses F4. Compute the score from consistent counts and choose using the stated rules.
Situation¶
The task grades binary positive-class F4. The split matches deployment, all four candidates use the same validation set, and there is no additional precision floor or review-capacity constraint. You have one submission slot left.
[ F_4 = rac{17PR}{16P+R} = rac{17TP}{17TP+16FN+FP}. ]
A weighted training loss and threshold tuning can both change this score; neither method is inherently more robust.
Artifact Packet¶
These are fixed illustrative confusion counts. The page and runner derive their metrics without rounding intermediate precision or recall.
| candidate | accuracy | precision | recall | F4 |
|---|---|---|---|---|
balanced_baseline |
0.9402 | 0.7098 | 0.6800 | 0.6817 |
recall_tuned_threshold |
0.9068 | 0.5201 | 0.8800 | 0.8456 |
recall_loss_retrain |
0.9234 | 0.5798 | 0.8500 | 0.8273 |
aggressive_recall |
0.8400 | 0.3800 | 0.9500 | 0.8730 |
Underlying confusion counts:
| candidate | TP | FN | FP | TN |
|---|---|---|---|---|
balanced_baseline |
680 | 320 | 278 | 8722 |
recall_tuned_threshold |
880 | 120 | 812 | 8188 |
recall_loss_retrain |
850 | 150 | 616 | 8384 |
aggressive_recall |
950 | 50 | 1550 | 7450 |
Decision Prompt¶
- Which candidate would you submit under the stated metric?
- Compute F4 directly from the counts; why does the accuracy winner lose?
- What additional evidence could justify selecting a lower-F4 candidate?
- How would an explicit precision floor change this decision?
Strong Reasoning Looks Like¶
- optimize the specified validation metric and state the deployment assumptions
- report precision and recall alongside F4 so the tradeoff is visible
- require repeated or shifted-set evidence before claiming one training method is more robust
- keep test labels out of model and threshold selection
Run The Clinic In Browser¶
The runner prints this fixed illustrative packet and recalculates any derived columns. It does not run a new training experiment. Edit its PACKET values to explore the decision.
Reference Reveal¶
Open after writing your note
Choose **`aggressive_recall`**: it has the highest validation F4 under the stated rules. Its lower precision is a real tradeoff, but it does not invalidate the result or prove fragility. The packet supplies no evidence that `recall_loss_retrain` is more robust. A documented deployment shift, repeat-run uncertainty, or a newly specified precision or capacity constraint could justify another choice. For example, a precision floor of 0.50 excludes `aggressive_recall`; among the remaining rows, `recall_tuned_threshold` has the highest F4. Evaluate that constraint on appropriate validation data before locking the policy. A recall-weighted loss is a surrogate, not a guarantee of maximizing F4. Compare the final predictions with the actual graded metric.What To Do Next¶
After this clinic:
- open Evaluation Metrics Deep Dive
- open Calibration and Thresholds for the threshold-side of this decision
- use IOAI Competition Surface for the full metric-reading drill under time