Skip to content

Clinic 10

Splitter Choice Under Ambiguity

Choose a split that represents the users and time period the deployed model will face.

Situation

The table contains repeated observations per user over time. Deployment must predict outcomes for previously unseen users in a future time period. Features must be available at prediction time; exclude later events and allow for label delays.

Artifact Packet

These fixed illustrative results compare different validation populations. They are not paired estimates on identical examples, so differences cannot all be attributed to one cause. The runner calculates each AUC drop from the random-split score.

split AUC average precision AUC drop vs random
random 0.912 0.74 0.000
group_by_user 0.681 0.38 0.231
chronological 0.553 0.22 0.359
group_and_time 0.541 0.21 0.371

Decision Prompt

  1. Which protocol matches new users in a future time period?
  2. What optimism can a random row split introduce here?
  3. Do these scores alone prove any split is correct?
  4. How would the answer change for future events from known users?

Strong Reasoning Looks Like

  • derive the split from the deployment population and time horizon
  • keep users disjoint and training observations earlier than validation for this scenario
  • fit preprocessing and tune models inside the appropriate training folds
  • inspect sample counts, prevalence, temporal drift, and label availability

Run The Clinic In Browser

The runner prints this fixed illustrative packet and recalculates any derived columns. It does not run a new training experiment. Edit its PACKET values to explore the decision.

Reference Reveal

Open after writing your note Use **`group_and_time`** for this stated deployment: validation users are absent from training, and their validation events occur after the training cutoff. A prospective new-user cohort is a suitable design. Audit the actual indices and feature timestamps before trusting it. AUC 0.541 is not evidence of correctness by itself. The justification is the deployment match. Random splitting can share user-specific patterns and future information; group-only splitting can still train on future events, while time-only splitting can include previously seen users. If deployment instead serves future events from known users, a chronological split allowing past records from those users may be appropriate. The lowest score and the hardest possible split are not automatic goals. Use a protocol that answers the actual generalization question.

What To Do Next

After this clinic:

  1. open Honest Splits and Baselines
  2. open Leakage Patterns for the adjacent failures
  3. use IOAI Competition Surface for the problem-reading habit that catches this earlier