Clinic 16
PEFT Depth Or Full Fine-Tune
Choose an adaptation experiment that fits the budget and leaves room to assess whether its gain repeats.
Situation¶
You have eight GPU-hours for the next development round. Small pilot runs produced the illustrative scores below. A candidate with validation score at least 0.740 is acceptable for the next comparison, and your current priority is to assess repeatability before spending more compute. Timings include a complete training/evaluation run but exclude one-time setup.
PEFT limits the trainable adaptation parameters. Full fine-tuning updates the base weights. Neither method guarantees better generalization, and actual memory depends on architecture, precision, activations, optimizer, and batch size.
Artifact Packet¶
Scores are single runs, not estimated means or confidence intervals. The runner calculates how many complete runs fit in the budget; it does not infer the score distribution or pick the maximum of invented repeats.
| method | GPU-hours per run | one-run validation score | complete runs in 8 GPU-hours |
|---|---|---|---|
rank4_last8 |
1.2 | 0.742 | 6 |
rank16_all |
3.5 | 0.751 | 2 |
full_finetune |
8.0 | 0.756 | 1 |
Decision Prompt¶
- How many complete runs fit for each option?
- Which experiment plan would you choose, and why?
- Does one best score establish a reliable winner?
- What must be saved to reproduce or roll back either adaptation method?
Strong Reasoning Looks Like¶
- distinguish the highest observed score from a reliably better method
- design repeats or a mixed allocation around the uncertainty that matters
- use the same validation protocol and inspect deployment-critical slices
- record the base checkpoint, adapter or weight delta, optimizer, data, seed, and configuration
Run The Clinic In Browser¶
The runner prints this fixed illustrative packet and recalculates any derived columns. It does not run a new training experiment. Edit its PACKET values to explore the decision.
Reference Reveal¶
Open after writing your note
Under the stated repeatability priority, **start with rank 4 on the last eight layers**: six full runs fit in 7.2 hours. It clears the pilot floor and allows more information about seed variability. Rank 16 can fit two runs, and full fine-tuning only one. These counts do not establish which method has the highest expected score. A mixed allocation or a higher-scoring candidate is defensible if you explain which comparison it enables and stay within budget. Predeclare how repeated results will be summarized; selecting only the highest seed exaggerates performance. Saving the original base checkpoint makes either method reversible. Adapters can simplify storage and swapping, but freezing the base does not guarantee unchanged behavior on old tasks. Test for regressions explicitly. No parameter-count estimate is valid without specifying the architecture and adapted matrices.What To Do Next¶
- open PEFT and LoRA for the math and the target-module priority
- open Transfer and Fine-Tuning for freezing schedules and layer-wise LR
- open Freeze Or Fine-Tune? for the adjacent classical clinic
- draft the compute spreadsheet for your own hardware and data size; if the math disagrees with this clinic's reference, the clinic's answer flips