Skip to content

Logistic Regression

What This Is

Logistic regression is the classification sibling of linear regression. It uses the same linear score z = Xw + b but passes it through a sigmoid (or softmax, for multi-class) to produce a probability: p(y=1 | x) = σ(z).

The practical lesson is that logistic regression is an honest baseline whose coefficients, thresholds, and calibration are all things you can read directly. If a complicated classifier cannot beat it meaningfully, the complicated classifier is adding risk without value.

When You Use It

  • a classification baseline before you try trees or deep models
  • any task where you actually need a probability, not just a label
  • any task where the coefficients need to be read and defended by a human
  • tasks with lots of features but limited training examples — a well-regularized linear classifier is hard to beat in that regime
  • text-bag-of-words and other very sparse problems — linear classifiers dominate this regime

Do Not Use It When

  • the classes are not linearly separable in the feature space you have and you are unwilling to add feature interactions
  • you need nonlinear shape and tree ensembles are available — trees will usually handle it
  • the probability calibration story is more important than the decision and you are not willing to calibrate — combine with Calibration and Thresholds

Tooling

  • LogisticRegression
  • SGDClassifier(loss="log_loss") for very large problems
  • Pipeline + StandardScaler
  • predict_proba
  • decision_function
  • class_weight
  • C — inverse regularization strength
  • penalty="l1", penalty="l2", penalty="elasticnet"
  • solver="liblinear", "saga", "lbfgs"
  • multi_class="multinomial" for softmax

The Loss And The Gradient

Logistic regression minimizes log-loss (a.k.a. binary cross-entropy):

L(w) = -(1/n) Σ [ y_i log σ(z_i) + (1 - y_i) log (1 - σ(z_i)) ]

The gradient is beautifully clean:

∂L/∂w = (1/n) X^T (σ(z) - y)

That is the same broad structure as the linear-regression gradient—the residual-like error projected through the features. Scaling often improves numerical conditioning for gradient-based solvers, but convexity does not depend on standardized inputs.

Regularization Cheat Sheet

Penalty Effect Use when
l2 shrinks all coefficients smoothly default for well-behaved features
l1 drives some coefficients to zero automatic feature selection, sparse inputs
elasticnet mix of L1 and L2 correlated features where sparsity is still wanted

In scikit-learn, C is the inverse regularization-strength parameter, so smaller C means stronger regularization. Do not treat it as an exact 1/λ identity across libraries: objective normalization and penalty conventions differ.

Minimal Example

from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

model = make_pipeline(StandardScaler(), LogisticRegression())
model.fit(X_train, y_train)
probs = model.predict_proba(X_valid)[:, 1]

Scaling is usually important when regularization should treat numeric features comparably and for solvers sensitive to conditioning. It may be unnecessary for already comparable features, binary indicators, or a deliberately unit-aware specification; coefficient interpretation must retain the units used.

Worked Pattern

from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GridSearchCV
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

clf = make_pipeline(
    StandardScaler(),
    LogisticRegression(max_iter=1000, solver="liblinear"),
)
grid = GridSearchCV(
    clf,
    {"logisticregression__C": [0.01, 0.1, 1.0, 10.0, 100.0],
     "logisticregression__penalty": ["l1", "l2"]},
    cv=5,
    scoring="roc_auc",
)
grid.fit(X_train, y_train)

A log-spaced C grid efficiently explores several orders of magnitude; the penalty effect is not itself exponential. Scoring by ROC-AUC separates ranking from threshold choice—pair with Calibration and Thresholds when probabilities or a transferable cutoff matter.

Multi-Class

For multi-class, logistic regression becomes softmax regression:

p(y=k | x) = exp(z_k) / Σ_j exp(z_j)

For three or more classes, current scikit-learn solvers other than liblinear optimize multinomial loss; multi_class is version-sensitive and deprecated in recent releases. If a solver supports only binary fits, wrap it explicitly with OneVsRestClassifier. Pin the scikit-learn version and compare calibration rather than assuming one formulation is better calibrated.

What To Inspect

  • the ROC curve and AUC — the model's ranking ability independent of any threshold
  • the calibration curve — is a score of 0.8 really about 80% positive?
  • per-class precision/recall, not just overall accuracy
  • coefficients on standardized features — which columns are pulling the decision
  • class_weight="balanced" on imbalanced problems — does it move the minority recall
  • the confusion matrix under two different operating points

Failure Pattern

Confusing C with regularization strength. Lowering scikit-learn's C increases regularization; a grid containing only large values explores the weakly regularized side.

A second failure is treating predict_proba as calibrated by default. For log-loss-trained linear models the calibration is usually close, but if you used class_weight="balanced" or heavy regularization the probabilities may be off — check the calibration curve and see Calibration and Thresholds.

A third failure is reading the coefficients before scaling. Without scaling, the coefficient on a dollar feature versus a percent feature is a unit artifact, not a signal.

Quick Checks

  1. Is the pipeline scaling the features before the classifier?
  2. Is max_iter high enough that the solver converges cleanly?
  3. Is C being searched on a log grid?
  4. Is the threshold you are using the one predict chose, or the one the cost structure demands?
  5. Did you look at the calibration curve, not only the accuracy?

Practice

  1. Fit logistic regression on a toy 2D binary problem and plot the linear decision boundary.
  2. Compare C = 0.01, C = 1, and C = 100 on the same split. Note what the coefficients do.
  3. Switch penalty="l2" to penalty="l1" and count how many coefficients go to zero.
  4. Add a second class and compare multinomial logistic regression against one-versus-rest.
  5. Measure the calibration curve with CalibrationDisplay and explain what you see.
  6. Move the decision threshold from 0.5 to 0.3 and explain which cost structure that implies.
  7. Explain why scaling matters for regularized logistic regression.
  8. Describe one signal that class_weight="balanced" is helping and one signal that it is distorting the probabilities.
  9. Explain why log-loss is the natural loss function for a probabilistic classifier.
  10. State one problem for which logistic regression will not be competitive no matter how well you regularize.

Runnable Example

Run the course-support logistic baseline from the repository root:

.venv/bin/python examples/course-support-baseline/logistic_baseline.py

Inspect held-out metrics and the largest coefficients. Then change only C or the decision threshold so you can attribute the effect to one decision at a time.

Longer Connection

Logistic regression sits next to:

Logistic regression is the honest default. When it wins, you ship it. When it loses to a more complex model, you now know exactly what the complex model is buying you.