Skip to content

Learning Curves and Bias-Variance

What This Is

Learning curves answer one practical question:

  • should the next move be more data, a simpler model, a stronger model, or better features

The curves matter because they turn vague bias-variance language into a visible decision. Bias-variance is a decomposition of expected test error; the learning curve is how that decomposition becomes a chart you can read and act on.

When You Use It

  • after a first baseline, before deciding what to spend time on next
  • when train and validation behavior disagree
  • when you need to decide whether collecting more data is worth it
  • when a model feels too simple or too flexible and you want evidence
  • when a teammate claims "the model needs more data" and you want to check

Do Not Use It When

  • the split is unfair — learning curves on leaky splits are disinformation, not diagnosis
  • the metric is unstable at small sample sizes (e.g., AP at very low prevalence) — noise dominates the early part of the curve

Bias And Variance, Briefly

For a prediction ŷ(x) trained on a finite dataset:

E[(y - ŷ(x))²]  =  bias²(ŷ(x))  +  variance(ŷ(x))  +  irreducible noise
  • bias — the error from the model family being too restrictive (linear when truth is curved; shallow tree when truth is subtle). Adding data does not help.
  • variance — the error from the model being too sensitive to the specific training set. Adding data does help, up to a ceiling.
  • irreducible noise — the floor. Label noise, genuine ambiguity. No model reaches below it.

Learning curves let you see which term dominates — bias shows up as both curves plateauing together; variance shows up as a big and persistent train-validation gap.

Read The Curves As Decisions

Pattern Likely story Next move
both train and validation low underfitting (high bias) use a stronger model or better features
train high, validation low, large gap overfitting (high variance) simplify, regularize, or add data
both high and close healthy fit do not change much without a new reason
validation still climbing with more data data may still help collect or simulate more data if practical
train climbs, validation flat then dips feature quality crumbling on held-out check distribution shift or leakage

Minimal Pattern

from sklearn.model_selection import learning_curve
import numpy as np
import matplotlib.pyplot as plt

train_sizes, train_scores, valid_scores = learning_curve(
    estimator=model,
    X=X_train, y=y_train,
    train_sizes=np.linspace(0.1, 1.0, 8),
    cv=5,
    scoring="accuracy",
    shuffle=True, random_state=0,
)

train_mean = train_scores.mean(axis=1); train_std = train_scores.std(axis=1)
valid_mean = valid_scores.mean(axis=1); valid_std = valid_scores.std(axis=1)

plt.plot(train_sizes, train_mean, "o-", label="train")
plt.fill_between(train_sizes, train_mean - train_std, train_mean + train_std, alpha=0.2)
plt.plot(train_sizes, valid_mean, "o-", label="validation")
plt.fill_between(train_sizes, valid_mean - valid_std, valid_mean + valid_std, alpha=0.2)
plt.xlabel("training set size"); plt.ylabel("score"); plt.legend()

The point is not the plot alone. The point is the next action the plot justifies.

Variance-Decomposition Recipe

When bias-variance language is abstract, pin it down empirically on a small regression task:

# Repeatedly train the same model on bootstrap resamples; measure predictions on a fixed test set.
import numpy as np
from sklearn.utils import resample

N_BOOTS = 100
preds = np.zeros((N_BOOTS, len(X_test)))
for b in range(N_BOOTS):
    X_b, y_b = resample(X_train, y_train)
    m = clone(model).fit(X_b, y_b)
    preds[b] = m.predict(X_test)

mean_pred = preds.mean(axis=0)
bias2     = ((mean_pred - y_test) ** 2).mean()        # how far the average prediction is from truth
variance  = preds.var(axis=0).mean()                  # how much predictions wobble across resamples
# Total error ≈ bias² + variance + noise

Two models with similar total error can split very differently: one high-bias and steady, one low-bias and wobbly. The decomposition tells you which lever matters. For a narrated run-through see Ensemble Methods — bagging is the operational tool for reducing variance once the decomposition points there.

When More Data Actually Helps

"Just collect more data" is a reflexive answer. The learning curve tells you whether it is the right answer:

  • validation curve still climbing at the last train-size point — yes, more data will help. Estimate how much more by extrapolating the curve: if doubling the data moved the score 2 points, another doubling is likely to move it about 1–1.5 points. Diminishing returns are the rule.
  • validation curve flat for the last two points — no. The model is either bias-limited or at the irreducible-noise floor. More data spends time and money without moving the metric.
  • validation curve noisy, no clear trend — the CV configuration is not informative enough to decide. Increase folds, average across seeds, or widen the train-size range.

Before committing to a data-collection project, look at the slope of the validation curve at the right edge. A flat tangent is a definitive "no."

Double-Descent Caveat

The classical U-shape (simple → high bias, complex → high variance, sweet spot in the middle) is a real pattern. It is also not the whole story. Modern over-parameterized models — large neural nets, high-degree kernels — can keep improving past the "interpolation threshold" where they fit the training set exactly. Test error first rises (classical overfitting), then falls again as capacity keeps increasing. This is double descent.

What this changes in practice:

  • the learning-curve "too much capacity is bad" reading is correct for classical models and for under-regularized nets, but can mislead for well-regularized deep nets where bigger keeps winning
  • use a capacity sweep, not a single-point judgment — both a small and a large model may be better than the middle one
  • for deep learning, the honest curve is validation-vs-compute (or vs. model-size) rather than validation-vs-training-size; see Optimizers and Regularization and Learning Rate Schedulers

The classical shape is still the default read for classical ML. Double descent is the caveat to mention, not the headline.

What To Inspect

  • the train curve — does the model learn the training set at all? If not, the problem is bias, not data volume.
  • the validation curve — the score that matters for action
  • the gap — width tracks variance; closing gaps mean the added data is buying generalization
  • the slope of the validation curve at the right edge — still climbing means more data will help; flat means the ceiling is the model
  • CV variance bands — a wide band says the metric itself is unstable, and no single curve reads reliably
  • whether the curve is plotted across the same CV folds each point — mixing schemes gives misleading shapes

If you read only the validation curve, you lose the reason behind the shape.

Failure Pattern

The common failure is collecting more data when the model is clearly underfitting. Both curves sit at the same low level; adding rows moves neither. The fix is a stronger model or better features, not more data.

The opposite failure is adding model complexity when the train-validation gap is already wide. A deep model on a noisy small dataset widens the gap instead of closing it. The fix is regularization, simpler models, or data augmentation.

A third, quieter failure is reading a single noisy fold. The validation curve wobbles, the person squints, and declares a "trend" that CV bands would reveal as noise.

Common Mistakes

  • reading only one curve
  • ignoring the train-validation gap
  • assuming more data always helps
  • plotting a noisy single split and treating it as stable evidence
  • changing model capacity and data size at the same time — you lose the diagnostic
  • conflating "validation score" with "generalization" when the validation set itself is biased or leaky
  • using classical-ML intuitions on an over-parameterized deep net and concluding "this is overfitting" without checking whether the curve will re-descend

A Good Curve Note

After one curve read, the learner should be able to say:

  • whether the main problem is bias or variance
  • whether more data is likely to help
  • whether the next move should target model capacity or feature quality
  • how confident the call is — CV band width, and whether the validation curve is still moving

Decision: What To Spend The Next Week On

Diagnosis Evidence Spend the week on
high bias both curves low, small gap, both flat stronger model, new features, different feature space
high variance large persistent gap, validation still climbing more data, stronger regularization, simpler model, bagging
sweet spot both high, small gap, validation flat stop tuning capacity; find a better label signal or deployment check
instability huge CV bands fix the split, the metric, or the class prior before anything else
double descent (deep) validation U-shapes then improves at larger capacity scale up the model, tighten regularization, watch for the second descent

Practice

  1. Plot a learning curve on a small dataset and decide whether the model is underfitting or overfitting. Defend the call in two sentences.
  2. Explain one case where more data is useful and one where it is not. Use the confusion between bias and variance as the discriminator.
  3. Compare a simpler and a more complex model on the same learning-curve view. Show the gap shift.
  4. Run the bootstrap bias/variance decomposition above on a regression task for two model capacities. Report bias², variance, and total error side by side.
  5. State what action the curve justifies next. If you cannot state it, the curve was not informative enough — rerun with more folds or more train-size points.
  6. (Deep-learning reader) take a small CNN and a much larger version on a noisy dataset. Plot test error vs. model width. Identify the classical-descent region and, if present, the second-descent region.
  7. Take a model that is clearly under-performing on a task. Before training anything, predict whether the problem is bias or variance based on model capacity and data volume. Then run a learning curve and compare the prediction to the reality. A correct call is a signal the mental model is calibrated; a wrong call is a lesson.
  8. Write the two-sentence "next move" that the curve justifies. If the sentence would be the same regardless of what the curve showed, the curve was not a decision tool — find a different diagnostic.

Debugging A Suspicious Flat Curve

A validation curve that stays flat across all training sizes can mean three very different things; the right diagnosis matters because each has a different fix:

  • bias-dominated — train curve is also low and close to validation. Stronger model or better features.
  • noise floor — train curve is much higher than validation, and validation is roughly at irreducible error (estimate with label agreement between two annotators, if available). No amount of data helps; improve labels instead.
  • wrong metric — the metric is measuring the wrong thing. A rare-event AUC can be "flat" while precision-at-k keeps improving with data. Re-plot the curve on the metric that matches the downstream decision.

If the diagnosis is unclear, widen the train-size sweep (add 5% and 10% points) and increase CV folds — most "flat curves" turn out to be under-sampled curves.

Runnable Example

Longer Connection

Learning curves sit next to:

The curve is a decision tool, not a report. If a learning curve leaves the reader with the same question they had before, plot more points or try a different capacity — do not file it and move on.