Clinic 23
Feature Selection Or Regularize
A feature selector saw validation labels. Repair the evaluation boundary before comparing it with a regularized model.
Situation¶
A binary dataset has 8,000 rows and 2,000 numeric features. One model uses the top 200 features ranked by mutual information with the label. Unfortunately, the ranking used all rows before the validation split. Another model fits a scaler and L1 logistic regression using training data only.
Artifact Packet¶
The table is illustrative. It explicitly records that the corrected MI experiment has not been run. No score can be inferred for that missing experiment.
| pipeline | selection fit rows | reported auc | usable for comparison |
|---|---|---|---|
mi_before_split |
all 8,000 rows | 0.87 | no: validation labels used |
mi_inside_folds |
training fold only | not yet run | pending |
l1_pipeline |
training fold only | 0.84 | yes: stated validation protocol |
Decision Prompt¶
- Which score is contaminated and how would you repair it?
- Can the honest MI score be predicted from this packet?
- Where should scaling, feature selection, and tuning happen?
- What evidence would justify choosing MI, L1, or their combination?
Strong Reasoning Looks Like¶
- reject the leaked 0.87 as an estimate of generalization
- fit each learned preprocessing and selection step inside the training fold
- tune MI feature count and classifier regularization using inner validation
- use an outer held-out evaluation or nested CV for a performance estimate after tuning
Run The Clinic In Browser¶
The runner prints this fixed illustrative packet and recalculates any derived columns. It does not run a new training experiment. Edit its PACKET values to explore the decision.
Reference Reveal¶
Open after writing your note
Use **the L1 pipeline as the current defensible baseline**, and rerun MI inside the same evaluation design before selecting a model family. The corrected MI score may fall, stay similar, or beat L1; this packet does not establish its value. A suitable starting pipeline is:from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GridSearchCV, StratifiedKFold
pipe = Pipeline([
("scale", StandardScaler()),
("clf", LogisticRegression(penalty="l1", solver="liblinear", max_iter=2000)),
])
grid = GridSearchCV(
pipe,
{"clf__C": [0.01, 0.1, 1.0, 10.0]},
cv=StratifiedKFold(n_splits=5, shuffle=True, random_state=0),
scoring="roc_auc",
)
grid.fit(X_train, y_train)
# Use X_test/y_test only after choosing the pipeline and its tuning procedure.
`grid.best_score_` helped select the hyperparameters; it is not an unbiased final performance estimate. Split by groups or time instead if deployment requires that structure. Fitting a scaler **after splitting, on training data only**, is correct.
MI, L1, tree-based selection, and combinations can all be evaluated honestly. Training a selector on all training-fold features is legitimate; consulting validation labels to select features is the problem. Held-out permutation importance is an inspection tool, but repeated selection using that same set also turns it into development data. Feature overlap alone does not establish selection stability or predictive validity.
What To Do Next¶
- open Feature Selection for filter/wrapper/embedded distinctions
- open Honest Splits and Baselines for pipeline leakage patterns
- open Data Cleaning Choice — the upstream clinic on column handling
- rerun your leaderboard with all feature-selection steps moved inside the CV loop; a ranking that survives is the honest ranking