Skip to content

Encoder-Decoder And Machine Translation

What This Is

An encoder-decoder model maps an input sequence to an output sequence through two modules: an encoder that reads the whole input into a set of contextual representations, and a decoder that generates the output token by token, conditioned on the encoder states and on what it has already produced.

The practical lesson is that encoder-decoder is not just a translation architecture. It is a natural fit for sequence-to-sequence tasks such as translation, summarization, generated question answering over context, and image captioning. Decoder-only and other architectures can solve some of the same tasks, so architecture remains an empirical choice.

When You Use It

  • machine translation
  • summarization — long article → short summary
  • question-answering — context + question → answer span or generated answer
  • image captioning — vision encoder → text decoder
  • speech recognition — audio encoder → text decoder
  • any task where the output length is not determined by the input length and must be generated one token at a time

Core Structure

source tokens → encoder → contextual representations (one per source position)
target tokens so far → decoder → next-token distribution
                                  attends to encoder states + its own past

Two forms:

  • RNN encoder-decoder — early systems used a fixed vector (Sutskever et al.); later attention-based systems aligned decoder states to encoder states (Bahdanau et al.)
  • Transformer encoder-decoder (Vaswani et al.) — self-attention within each side, cross-attention from decoder to encoder; widely used sequence-to-sequence architecture

See Attention and Transformers for the attention mechanics that make both work.

Teacher Forcing — How Training Works

During training the decoder is given the ground-truth previous tokens as input, not its own predictions. This is teacher forcing. The loss at each position is cross-entropy against the true next token.

input to decoder:    <BOS> y_1 y_2 ... y_{T-1}
target:              y_1   y_2 y_3 ... y_T

Teacher forcing lets a Transformer compute losses for all target positions in parallel; an RNN decoder still processes recurrent steps sequentially. Exclusive teacher forcing also creates a train/inference mismatch called exposure bias: at inference, the decoder conditions on its own earlier predictions. Scheduled sampling and sequence-level objectives are possible interventions, but each changes optimization and must be evaluated.

See Language Modeling Fundamentals for the underlying next-token-prediction framing.

Decoding At Inference Time

The decoder is autoregressive: it predicts one token, appends it to the context, and predicts the next. Strategies to pick each token:

Strategy Rule Good for
Greedy argmax at every step fast, deterministic baseline; can miss better sequence-level hypotheses
Beam search keep top-k partial hypotheses translation, summarization — the historical default
Sampling draw from the distribution open-ended generation
Top-k sampling sample from top k tokens, renormalized creative writing, chat
Top-p / nucleus sample from smallest set with total prob ≥ p modern default for open-ended text
Temperature scale logits by 1/T before softmax lower T = sharper, higher T = more random

Beam search remains a useful translation baseline, while sampling is appropriate when output diversity matters. Tune beam width, length penalty, temperature, and sampling filters on task-specific evaluation rather than treating one decoder as universally best. See Prompting and Tool Use.

Evaluation Metrics

Sequence generation has no single right answer, so evaluation is harder than classification.

Metric Measures Use for
BLEU n-gram precision against references translation — standard baseline
chrF character-level F-score translation for morphologically rich languages
ROUGE n-gram recall against references summarization
BERTScore embedding-space similarity to references semantic similarity beyond n-grams
COMET learned metric trained on human judgments translation — strong learned metric; pin the model/version (original paper)
perplexity exponentiated average token negative log-likelihood intrinsic fit when tokenization and evaluation setup match

No metric is a substitute for reading outputs. A BLEU of 42 says the n-grams overlap with the reference; it does not say the sentence is correct.

Length Control And /

Two special tokens hold the sequence-generation plumbing together:

  • <BOS> (beginning of sequence) — primes the decoder
  • <EOS> (end of sequence) — the decoder emits this to stop; without it the decoder generates forever
  • practical systems also impose a max_new_tokens safety cap

At inference, generation stops when the decoder emits <EOS> or the cap is hit. Rare <EOS> emission can come from data formatting, masking/label bugs, length bias, decoding settings, or a mismatch in training examples.

What To Inspect

  • decoder cross-attention patterns as a debugging signal for masks and indexing—not a requirement that every head align with human word correspondences
  • whether <EOS> is produced at a reasonable rate
  • output length distribution vs. reference length distribution — a systematic shortening or lengthening is a training-signal bug
  • whether the error pattern lives in translation quality (word choice) or in fluency (syntax)
  • BLEU and a small read of 20 examples — never ship on metric alone
  • teacher-forced token loss plus free-running sequence metrics; their difference is not a single directly comparable "loss gap"

Failure Pattern

Shipping a model that rarely emits <EOS> and relying on max_new_tokens to truncate. Inspect target construction, loss masking, decoding length bias, and training length distribution before choosing an intervention.

A second failure pattern is selecting greedy or beam decoding by convention without comparing task quality, length bias, and latency on held-out translations.

A third failure pattern is comparing two models on BLEU alone and picking the higher one. Estimate uncertainty with paired resampling or repeated evaluation where appropriate, and inspect a fixed output diff set.

A fourth failure pattern is evaluating generation quality on the training distribution. Exposure bias shows only at inference on fresh inputs, where the decoder must condition on its own output.

Quick Checks

  1. Is teacher forcing wired in training, and free-running wired at inference?
  2. Is the decoder producing <EOS> at a reasonable rate?
  3. Was the decoding strategy selected on held-out task quality, length behavior, and latency?
  4. Is the cross-attention alignment sensible on a sample?
  5. Is the output length distribution matched to the reference length distribution?

Practice

  1. Implement a tiny RNN encoder-decoder on a character-level reversal task (input → reversed).
  2. Replace the RNN decoder with a greedy transformer decoder and compare sample quality.
  3. Add beam search with k = 1, 4, 16 and compare BLEU on a small translation set.
  4. Lower decoding temperature from 1.0 to 0.2 on an open-ended generation task and observe the sample distribution.
  5. Compute teacher-forced token loss and free-running sequence metrics on the same validation set. Explain what each reveals and why they are not directly comparable.
  6. Explain why cross-attention is what makes the decoder "look at" the encoder.
  7. Describe one case where BLEU disagrees with human judgment.
  8. State why <EOS> is a trained decision and not a post-processing step.
  9. Explain the relationship between temperature and top-p and when they stack.
  10. Describe one systematic error pattern you would expect in a low-resource language pair.

Runnable Example

Run the matching lab from the repository root:

.venv/bin/python labs/encoder-decoder-translation/src/encoder_decoder_workflow.py

Inspect the teacher-forcing curve, decoded sequences, and EOS positions. A falling training loss is not enough if free-running decoding repeats or never terminates.

Longer Connection

Encoder-decoder sits next to:

Encoder-decoder is one powerful shape for sequence-in, sequence-out problems. The decisions—architecture, attention pattern, decoding strategy, tokenization, and metric—are where the real engineering lives.