Logistic Regression¶
What This Is¶
Logistic regression is the classification sibling of linear regression. It uses the same linear score z = Xw + b but passes it through a sigmoid (or softmax, for multi-class) to produce a probability: p(y=1 | x) = σ(z).
The practical lesson is that logistic regression is an honest baseline whose coefficients, thresholds, and calibration are all things you can read directly. If a complicated classifier cannot beat it meaningfully, the complicated classifier is adding risk without value.
When You Use It¶
- a classification baseline before you try trees or deep models
- any task where you actually need a probability, not just a label
- any task where the coefficients need to be read and defended by a human
- tasks with lots of features but limited training examples — a well-regularized linear classifier is hard to beat in that regime
- text-bag-of-words and other very sparse problems — linear classifiers dominate this regime
Do Not Use It When¶
- the classes are not linearly separable in the feature space you have and you are unwilling to add feature interactions
- you need nonlinear shape and tree ensembles are available — trees will usually handle it
- the probability calibration story is more important than the decision and you are not willing to calibrate — combine with Calibration and Thresholds
Tooling¶
LogisticRegressionSGDClassifier(loss="log_loss")for very large problemsPipeline+StandardScalerpredict_probadecision_functionclass_weightC— inverse regularization strengthpenalty="l1",penalty="l2",penalty="elasticnet"solver="liblinear","saga","lbfgs"multi_class="multinomial"for softmax
The Loss And The Gradient¶
Logistic regression minimizes log-loss (a.k.a. binary cross-entropy):
L(w) = -(1/n) Σ [ y_i log σ(z_i) + (1 - y_i) log (1 - σ(z_i)) ]
The gradient is beautifully clean:
∂L/∂w = (1/n) X^T (σ(z) - y)
That is the same broad structure as the linear-regression gradient—the residual-like error projected through the features. Scaling often improves numerical conditioning for gradient-based solvers, but convexity does not depend on standardized inputs.
Regularization Cheat Sheet¶
| Penalty | Effect | Use when |
|---|---|---|
l2 |
shrinks all coefficients smoothly | default for well-behaved features |
l1 |
drives some coefficients to zero | automatic feature selection, sparse inputs |
elasticnet |
mix of L1 and L2 | correlated features where sparsity is still wanted |
In scikit-learn, C is the inverse regularization-strength parameter, so smaller C means stronger regularization. Do not treat it as an exact 1/λ identity across libraries: objective normalization and penalty conventions differ.
Minimal Example¶
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
model = make_pipeline(StandardScaler(), LogisticRegression())
model.fit(X_train, y_train)
probs = model.predict_proba(X_valid)[:, 1]
Scaling is usually important when regularization should treat numeric features comparably and for solvers sensitive to conditioning. It may be unnecessary for already comparable features, binary indicators, or a deliberately unit-aware specification; coefficient interpretation must retain the units used.
Worked Pattern¶
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GridSearchCV
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
clf = make_pipeline(
StandardScaler(),
LogisticRegression(max_iter=1000, solver="liblinear"),
)
grid = GridSearchCV(
clf,
{"logisticregression__C": [0.01, 0.1, 1.0, 10.0, 100.0],
"logisticregression__penalty": ["l1", "l2"]},
cv=5,
scoring="roc_auc",
)
grid.fit(X_train, y_train)
A log-spaced C grid efficiently explores several orders of magnitude; the penalty effect is not itself exponential. Scoring by ROC-AUC separates ranking from threshold choice—pair with Calibration and Thresholds when probabilities or a transferable cutoff matter.
Multi-Class¶
For multi-class, logistic regression becomes softmax regression:
p(y=k | x) = exp(z_k) / Σ_j exp(z_j)
For three or more classes, current scikit-learn solvers other than liblinear optimize multinomial loss; multi_class is version-sensitive and deprecated in recent releases. If a solver supports only binary fits, wrap it explicitly with OneVsRestClassifier. Pin the scikit-learn version and compare calibration rather than assuming one formulation is better calibrated.
What To Inspect¶
- the ROC curve and AUC — the model's ranking ability independent of any threshold
- the calibration curve — is a score of 0.8 really about 80% positive?
- per-class precision/recall, not just overall accuracy
- coefficients on standardized features — which columns are pulling the decision
class_weight="balanced"on imbalanced problems — does it move the minority recall- the confusion matrix under two different operating points
Failure Pattern¶
Confusing C with regularization strength. Lowering scikit-learn's C increases regularization; a grid containing only large values explores the weakly regularized side.
A second failure is treating predict_proba as calibrated by default. For log-loss-trained linear models the calibration is usually close, but if you used class_weight="balanced" or heavy regularization the probabilities may be off — check the calibration curve and see Calibration and Thresholds.
A third failure is reading the coefficients before scaling. Without scaling, the coefficient on a dollar feature versus a percent feature is a unit artifact, not a signal.
Quick Checks¶
- Is the pipeline scaling the features before the classifier?
- Is
max_iterhigh enough that the solver converges cleanly? - Is
Cbeing searched on a log grid? - Is the threshold you are using the one
predictchose, or the one the cost structure demands? - Did you look at the calibration curve, not only the accuracy?
Practice¶
- Fit logistic regression on a toy 2D binary problem and plot the linear decision boundary.
- Compare
C = 0.01,C = 1, andC = 100on the same split. Note what the coefficients do. - Switch
penalty="l2"topenalty="l1"and count how many coefficients go to zero. - Add a second class and compare multinomial logistic regression against one-versus-rest.
- Measure the calibration curve with
CalibrationDisplayand explain what you see. - Move the decision threshold from 0.5 to 0.3 and explain which cost structure that implies.
- Explain why scaling matters for regularized logistic regression.
- Describe one signal that
class_weight="balanced"is helping and one signal that it is distorting the probabilities. - Explain why log-loss is the natural loss function for a probabilistic classifier.
- State one problem for which logistic regression will not be competitive no matter how well you regularize.
Runnable Example¶
Run the course-support logistic baseline from the repository root:
.venv/bin/python examples/course-support-baseline/logistic_baseline.py
Inspect held-out metrics and the largest coefficients. Then change only C or the decision threshold so you can attribute the effect to one decision at a time.
Longer Connection¶
Logistic regression sits next to:
- Calibration and Thresholds — how to turn scores into decisions
- Evaluation Metrics Deep Dive — the metrics to trust for classifiers
- Feature Selection — L1 regularization as a selection lens
- SVM Margins and Kernels — the margin-based alternative with different trade-offs
- Linear Regression — same gradient structure, continuous target
Logistic regression is the honest default. When it wins, you ship it. When it loses to a more complex model, you now know exactly what the complex model is buying you.