Skip to content

Clinic 14

Tokenizer Choice

Measure coverage and sequence length across languages. Fragmentation raises costs, while unknown tokens and truncation can lose information.

Situation

You are building a multilingual language-ID classifier. An English-heavy tokenizer produces long sequences for Swahili and Telugu and unknown tokens in some code-switched text. You can train a new model, so tokenizer alternatives are available; replacing the tokenizer of an existing pretrained model would require compatible embedding/model adaptation rather than a drop-in swap.

Artifact Packet

The counts are fixed illustrative measurements; the runner does not invoke or train a tokenizer. Here an unknown-token rate is measured per whitespace-delimited word, so comparisons should preserve that definition. The BPE implementation in this packet lacks complete byte fallback; BPE as a family does not inherently require unknown tokens.

slice mean pieces per whitespace word fraction of whitespace words containing UNK
English 1.3 0.00
Swahili 6.4 0.00
Telugu 11.7 0.00
code-switched 7.2 0.08

Decision Prompt

  1. Which observations indicate inefficiency and which indicate lost information?
  2. What training-corpus and vocabulary alternatives would you compare?
  3. Why is 64k vocabulary not automatically best?
  4. What must be retrained or adapted if an existing model’s tokenizer changes?

Strong Reasoning Looks Like

  • distinguish lossless fragmentation from actual unknown-token replacement or truncation
  • train tokenizer candidates on the training corpus with an intentional language balance
  • measure tokens per input, truncation, coverage, latency, memory, and per-language F1
  • compare complete model/tokenizer systems under a stated compute or latency budget

Lossless byte sequences retain the original text; a model can learn from them. Their costs include longer sequences, context limits, and learning efficiency. Byte fragmentation is not proof that morphology is unrecoverable. ByT5 demonstrates useful language learning directly from bytes. A character vocabulary of about 1,000 symbols covers only its specified alphabet unless a fallback is provided.

Run The Clinic In Browser

The runner prints this fixed illustrative packet and recalculates any derived columns. It does not run a new training experiment. Edit its PACKET values to explore the decision.

Reference Reveal

Open after writing your note **Test a balanced multilingual tokenizer with explicit coverage handling**, alongside the existing tokenizer and a lossless-byte fallback baseline. Compare, for example, 32k and 64k vocabularies under a controlled training and evaluation design. This packet establishes a problem worth investigating, not the performance of those unrun alternatives. A larger vocabulary may shorten sequences but expands embedding/output tables and leaves less data per rare token. Measure the net cost and downstream language-ID quality. Report per-language performance and code-switched cases; do not promise a specific F1 increase from vocabulary size alone. For pure language ID, a character n-gram baseline is also a useful task-specific comparison. Keep vocabulary learning and model selection on the training/development side of the split, then evaluate the locked pipeline on held-out data.

What To Do Next

  1. open Text Representations and Order — the systematic treatment of subword choices
  2. open Text Generation and Language Models — the same tokenizer issues hit generation even harder
  3. open Reliability Slices — the slice discipline this clinic relies on
  4. rerun the product data through each tokenizer and plot subword-per-word histograms; one chart decides most of this clinic