Clinic 10
Splitter Choice Under Ambiguity
Choose a split that represents the users and time period the deployed model will face.
Situation¶
The table contains repeated observations per user over time. Deployment must predict outcomes for previously unseen users in a future time period. Features must be available at prediction time; exclude later events and allow for label delays.
Artifact Packet¶
These fixed illustrative results compare different validation populations. They are not paired estimates on identical examples, so differences cannot all be attributed to one cause. The runner calculates each AUC drop from the random-split score.
| split | AUC | average precision | AUC drop vs random |
|---|---|---|---|
random |
0.912 | 0.74 | 0.000 |
group_by_user |
0.681 | 0.38 | 0.231 |
chronological |
0.553 | 0.22 | 0.359 |
group_and_time |
0.541 | 0.21 | 0.371 |
Decision Prompt¶
- Which protocol matches new users in a future time period?
- What optimism can a random row split introduce here?
- Do these scores alone prove any split is correct?
- How would the answer change for future events from known users?
Strong Reasoning Looks Like¶
- derive the split from the deployment population and time horizon
- keep users disjoint and training observations earlier than validation for this scenario
- fit preprocessing and tune models inside the appropriate training folds
- inspect sample counts, prevalence, temporal drift, and label availability
Run The Clinic In Browser¶
The runner prints this fixed illustrative packet and recalculates any derived columns. It does not run a new training experiment. Edit its PACKET values to explore the decision.
Reference Reveal¶
Open after writing your note
Use **`group_and_time`** for this stated deployment: validation users are absent from training, and their validation events occur after the training cutoff. A prospective new-user cohort is a suitable design. Audit the actual indices and feature timestamps before trusting it. AUC 0.541 is not evidence of correctness by itself. The justification is the deployment match. Random splitting can share user-specific patterns and future information; group-only splitting can still train on future events, while time-only splitting can include previously seen users. If deployment instead serves future events from known users, a chronological split allowing past records from those users may be appropriate. The lowest score and the hardest possible split are not automatic goals. Use a protocol that answers the actual generalization question.What To Do Next¶
After this clinic:
- open Honest Splits and Baselines
- open Leakage Patterns for the adjacent failures
- use IOAI Competition Surface for the problem-reading habit that catches this earlier