Skip to content

Reinforcement Learning Foundations

What This Is

Reinforcement learning (RL) is the problem of learning a policy — a mapping from observations to actions — by interacting with an environment and receiving scalar rewards. The learner does not see the correct answer; it sees the reward signal and has to infer which of its actions were good.

Unlike supervised learning, RL has two sources of difficulty most beginners underestimate:

  • credit assignment across time: a reward received now may be the result of an action taken 100 steps ago
  • exploration vs. exploitation: the learner must try unknown actions to discover better ones, but trying unknown actions costs reward

This topic is a beginner-to-working treatment: the vocabulary, the two classical algorithm families, a readable PPO sketch, reward hacking as an inspection habit, and the decision "should I be using RL at all."

When You Use It

  • the task is sequential — decisions now affect data you see later
  • you have an environment you can simulate (or a safe way to run a policy in the real world)
  • the right action is not known in advance, but a scalar reward can be defined
  • RLHF: you want to align a large model with human preferences and you have preference data

Do Not Use It When

  • you have labeled (input, correct_action) pairs — supervised learning is strictly simpler and stronger
  • you cannot simulate, and real-world exploration is dangerous (robotics in the wild, medical decisions)
  • the reward is extremely sparse and you have no way to shape it — the student will spend months getting the agent off the floor

The Vocabulary

term meaning
state s_t what the agent observes at time t
action a_t what the agent does
reward r_t scalar signal the environment returns
policy π(a \| s) distribution over actions given a state
return G_t cumulative (optionally discounted) reward from t onward
value function V(s) expected return starting from s under π
action-value Q(s, a) expected return starting from s, taking a, then following π
discount factor γ 0 ≤ γ ≤ 1; how much future reward counts
trajectory / episode a sequence (s_0, a_0, r_0, s_1, a_1, r_1, ...)

An optimal policy maximizes expected return. Every RL algorithm is a way to estimate a value function, a policy, or both.

Two Families

RL algorithms split cleanly along one axis:

family what it learns examples
value-based a value function (Q or V); policy is implicit (greedy over Q) Q-learning, DQN, Rainbow
policy-based the policy directly; value function may be a helper REINFORCE, PPO, A3C, TRPO
actor-critic (hybrid) policy (actor) + value function (critic) together A2C, PPO, SAC, DDPG

For most deep-learning students approaching RL today, the path is: understand REINFORCE, then PPO. DQN and its descendants are important history and strong for discrete action spaces, but PPO is one common baseline for continuous actions and RLHF; algorithm choice depends on the environment, data access, and compute budget.

Policy Gradient — REINFORCE In One Equation

The objective is J(θ) = E_π[ G_0 ]. The policy-gradient theorem says:

∇_θ J(θ) = E_{trajectory ~ π_θ} [ Σ_t γ^t G_t ∇_θ log π_θ(a_t | s_t) ]

This is the likelihood-ratio gradient. In English: increase the log-probability of actions that led to high returns; decrease it for low returns. REINFORCE is the Monte Carlo estimate: roll out episodes, compute returns, multiply by log-probabilities, take a gradient step.

Readable REINFORCE:

import torch
from torch import nn

def reinforce_step(env, policy, optimizer, gamma=0.99):
    # Gymnasium API: reset() returns (obs, info); step() returns
    # (obs, reward, terminated, truncated, info). This example optimizes return
    # up to termination or truncation. Continuing tasks need appropriate
    # bootstrapping at a time limit; truncation is not necessarily terminal.
    log_probs, rewards = [], []
    s, _ = env.reset()
    done = False
    while not done:
        probs = policy(torch.as_tensor(s).float())
        dist = torch.distributions.Categorical(probs)
        a = dist.sample()
        s, r, terminated, truncated, _ = env.step(a.item())
        done = terminated or truncated
        log_probs.append(dist.log_prob(a))
        rewards.append(r)

    # Compute discounted returns
    returns, G = [], 0.0
    for r in reversed(rewards):
        G = r + gamma * G
        returns.insert(0, G)
    returns = torch.tensor(returns, dtype=torch.float32)
    # Raw returns use a zero baseline and retain the signal in one-step episodes.
    discounts = gamma ** torch.arange(len(rewards), dtype=torch.float32)
    loss = -(torch.stack(log_probs) * discounts * returns).sum()
    optimizer.zero_grad()
    loss.backward()
    optimizer.step()
    return sum(rewards)

Two points the snippet teaches:

  • subtracting a baseline b(s_t) that does not depend on the sampled action, conditional on state, preserves the expected policy gradient. A value estimate fitted on separate data or fixed before the rollout is one practical baseline. The snippet uses zero. Same-episode mean subtraction and division by a random sample standard deviation are separate normalization heuristics; they need not preserve the expected gradient and can erase the signal or produce NaN on singleton episodes. See the baseline derivation.
  • REINFORCE is high-variance and sample-inefficient; it is useful as a teaching object, not as a production algorithm

PPO — A Common Policy-Optimization Baseline

Proximal Policy Optimization (Schulman et al.) is the PPO-clip variant most widely used. It improves on REINFORCE in two ways:

  1. uses a critic (value estimate V(s)) to estimate advantages, for example A_t = G_t - V(s_t); a suitable baseline can reduce variance
  2. clips a surrogate objective to discourage large policy changes; clipping is not a strict bound on policy divergence and does not guarantee improvement

The clipped objective:

L_clip(θ) = E_t [ min( r_t(θ) · A_t ,  clip(r_t(θ), 1-ε, 1+ε) · A_t ) ]

where r_t(θ) = π_θ(a_t | s_t) / π_θ_old(a_t | s_t) is the importance-sampling ratio and ε is typically 0.2.

PPO in practice: collect a batch of trajectories with the current policy, compute advantages (usually via GAE), do a handful of SGD updates on the same batch with the clipped loss, then throw the batch away and repeat.

A working PPO also uses:

  • a value-function loss (MSE on returns or GAE targets)
  • an entropy bonus to keep exploration alive
  • careful normalization of observations, rewards, or advantages

Reward Design

Reward is the whole specification of the task. A poorly designed reward is the single largest source of RL failures.

Three patterns to watch:

  • sparse rewards (1 if goal reached, 0 otherwise) — easy to specify, extremely hard to learn from
  • shaped rewards (distance to goal, time penalty) — faster to learn, but prone to exploitation
  • curriculum — start with easier variants, increase difficulty as the agent improves

Reward Hacking

Reward hacking is when the agent finds a policy that maximizes the specified reward but not the task you wanted. Classic examples:

  • a boat-race agent that goes in circles collecting power-ups forever because power-ups give reward (OpenAI's CoastRunners bug)
  • a robot trained to "knock down a cup" that learns to hit the table so hard the cup falls
  • an LLM RLHF'd to maximize helpfulness reward that becomes sycophantic — saying what the rater wants

Every serious RL project should budget time for reward-hacking inspection:

  • held-out behavioral eval: run the learned policy against scenarios the reward did not explicitly score. Does it still do what you wanted?
  • adversarial probing: write one test case designed to exploit the reward. If the agent breaks in the predictable way, the reward is under-specified.
  • human spot-checks: the cheapest, most reliable eval; five humans labeling 100 rollouts catch what metrics do not

RLHF — Where RL Meets Modern LLMs

RLHF (Reinforcement Learning from Human Feedback) is the pipeline most modern chat assistants use for alignment:

  1. collect pairs of model outputs and a human preference label ("A is better than B")
  2. train a reward model r_φ(x, y) that scores y given x, via a cross-entropy loss on pairs
  3. fine-tune the language model with PPO, using the reward model in place of an environment reward

Two nuances that matter in practice:

  • a KL penalty keeps the policy close to the original supervised model; without it, the PPO policy drifts into reward-hacked text
  • the learned reward is a critical optimization target; misspecification can reward behavior that fails held-out human evaluation

DPO (Direct Preference Optimization) trains directly from preference pairs against a reference-policy objective, without fitting an explicit scalar reward model or running an online RL loop. It is simpler than PPO-based RLHF in some settings, but relative quality is dataset- and evaluation-dependent (Rafailov et al., 2023).

What To Inspect

  • mean episode return over training — should trend up; noisy plateaus are normal, sustained flats are a sign of bad exploration or credit assignment
  • policy entropy — collapses to zero means the policy has committed early; collapsing too fast is a symptom of PPO's objective plus too-high LR
  • value function loss — should track trajectory-return statistics; if it does not, the critic is unlearning the world
  • gradient norms — PPO in particular is sensitive; clip them
  • reward-hacking probes — run a small adversarial eval every N updates
  • KL between policy and reference in RLHF — plot it; sudden spikes are the policy abandoning the reference

Failure Pattern

The canonical RL failure is learned something, but not what you wanted. Metrics look healthy; reward is increasing; the agent's actual behavior is exploiting a quirk of the environment. Fix: separate training metrics from task metrics and measure both. The task metric is usually a human eval, a held-out behavioral test, or a counterfactual probe.

The second canonical failure is trained too long and collapsed. Policy entropy goes to zero, the agent always takes the same action, and the reward plateaus because exploration is dead. Fix: entropy bonus, earlier stopping, or a more diverse initialization.

Common Mistakes

  • choosing REINFORCE or PPO by reputation instead of matching the estimator and implementation complexity to the environment
  • PPO with the wrong entropy coefficient (too low kills exploration; too high prevents convergence)
  • shaping rewards aggressively without adversarial reward-hacking eval
  • judging RL success by training reward alone, not task-level behavior
  • shipping a policy without a held-out behavioral test set
  • in RLHF: training a policy longer than the reward model can credibly score
  • using supervised learning muscles (accuracy, loss) on an RL problem — they are the wrong primitives

Decision: RL vs. Supervised vs. Imitation vs. Behavior Cloning

option when it wins trade
supervised learning you have (input, correct_action) pairs simplest, strongest; needs labels
behavior cloning / imitation you have demonstrations but no reward no exploration needed; brittle off-distribution
offline RL you have logged trajectories with rewards but no environment powerful; data coverage often the bottleneck
online RL you can simulate and you need a policy that explores most flexible; most expensive
RLHF with PPO preference data converted into an explicit reward model reward-model and online-optimization complexity
DPO-style preference optimization preference pairs and a reference policy no explicit reward model; sensitive to preference data and objective settings

If the task can be cast as supervised or imitation, do that first. RL is the largest hammer in the toolbox; do not reach for it first.

Practice

  1. Implement REINFORCE on CartPole-v1 using the snippet above. Train to a solved threshold (average return ≥ 475 over 100 episodes).
  2. Add a critic (value network) and compute advantages; switch to an actor-critic update. Compare sample efficiency to REINFORCE.
  3. Switch to PPO using a library like Stable-Baselines3. Solve LunarLander-v2. Report training reward and a separate held-out behavioral test.
  4. Design a pathological reward for a simple environment (grid world): give +1 for every step the agent stands still near the start. Train and confirm the agent learns to do nothing. This is a one-hour reward-hacking demo.
  5. For an RLHF flavor: train a tiny reward model on a small preference dataset, PPO-fine-tune a tiny LM against it, and confirm KL to the reference stays bounded. Note what happens when you remove the KL penalty.
  6. Plot policy entropy across training for PPO. Note the relationship between entropy decay and training-reward plateau.

Runnable Example

Run the deterministic two-action environment from the repository root:

.venv/bin/python examples/deep-learning-recipes/advanced_mechanics_demo.py --demo rl

Inspect value estimates, action visits, and reward over time. The fixture isolates exploration versus exploitation without external environments; it is an orientation run, not evidence that an RL algorithm is deployment-ready.

Longer Connection

RL is adjacent to other academy topics along specific lines:

For the decision frame — "should I be using RL?" — re-read Baseline-First Task Solving. A supervised or imitation baseline that beats RL is a common and informative outcome.