Skip to content

Clinic 09

Metric Choice Under An Unusual Task

The grader uses F4. Compute the score from consistent counts and choose using the stated rules.

Situation

The task grades binary positive-class F4. The split matches deployment, all four candidates use the same validation set, and there is no additional precision floor or review-capacity constraint. You have one submission slot left.

[ F_4 = rac{17PR}{16P+R} = rac{17TP}{17TP+16FN+FP}. ]

A weighted training loss and threshold tuning can both change this score; neither method is inherently more robust.

Artifact Packet

These are fixed illustrative confusion counts. The page and runner derive their metrics without rounding intermediate precision or recall.

candidate accuracy precision recall F4
balanced_baseline 0.9402 0.7098 0.6800 0.6817
recall_tuned_threshold 0.9068 0.5201 0.8800 0.8456
recall_loss_retrain 0.9234 0.5798 0.8500 0.8273
aggressive_recall 0.8400 0.3800 0.9500 0.8730

Underlying confusion counts:

candidate TP FN FP TN
balanced_baseline 680 320 278 8722
recall_tuned_threshold 880 120 812 8188
recall_loss_retrain 850 150 616 8384
aggressive_recall 950 50 1550 7450

Decision Prompt

  1. Which candidate would you submit under the stated metric?
  2. Compute F4 directly from the counts; why does the accuracy winner lose?
  3. What additional evidence could justify selecting a lower-F4 candidate?
  4. How would an explicit precision floor change this decision?

Strong Reasoning Looks Like

  • optimize the specified validation metric and state the deployment assumptions
  • report precision and recall alongside F4 so the tradeoff is visible
  • require repeated or shifted-set evidence before claiming one training method is more robust
  • keep test labels out of model and threshold selection

Run The Clinic In Browser

The runner prints this fixed illustrative packet and recalculates any derived columns. It does not run a new training experiment. Edit its PACKET values to explore the decision.

Reference Reveal

Open after writing your note Choose **`aggressive_recall`**: it has the highest validation F4 under the stated rules. Its lower precision is a real tradeoff, but it does not invalidate the result or prove fragility. The packet supplies no evidence that `recall_loss_retrain` is more robust. A documented deployment shift, repeat-run uncertainty, or a newly specified precision or capacity constraint could justify another choice. For example, a precision floor of 0.50 excludes `aggressive_recall`; among the remaining rows, `recall_tuned_threshold` has the highest F4. Evaluate that constraint on appropriate validation data before locking the policy. A recall-weighted loss is a surrogate, not a guarantee of maximizing F4. Compare the final predictions with the actual graded metric.

What To Do Next

After this clinic:

  1. open Evaluation Metrics Deep Dive
  2. open Calibration and Thresholds for the threshold-side of this decision
  3. use IOAI Competition Surface for the full metric-reading drill under time