Skip to content

Tokenizer training

Goal

Create a vocabulary that compresses text into tokens efficiently across languages, code, math, and markup.

Common tokenizer families:

  • BPE (Byte Pair Encoding) — the most common choice for modern LLMs.
  • Unigram / SentencePiece — alternative scoring; SentencePiece is the de-facto reference implementation.
  • WordPiece-like variants — used by older BERT-family models, less common for new LLMs.
  • Byte-level BPE — operates on bytes, so any input is representable; used by GPT-2 onward.

Modern LLM tokenizers usually support arbitrary bytes so the model can handle unseen characters.

Design choices

Vocabulary size

Typical modern vocab sizes range from roughly 32K to 200K+ tokens.

Model Vocab size
GPT-2 50,257
Llama 2 32,000
Llama 3 128,256
Gemma 2 256,000
Mistral 32,000

Larger vocabularies:

  • Improve compression
  • Help multilingual text
  • Help code and markup
  • Increase embedding/output matrix size
  • Can make rare-token learning harder

Smaller vocabularies:

  • Are simpler
  • Reduce embedding size
  • May be inefficient for non-English languages and code

Compression matters more than you'd think

A bigger vocabulary that better compresses your training corpus = more information per training step at fixed sequence length, and lower inference cost per output character. The Llama 3 paper explicitly attributes part of its multilingual gains to the 128K vocab.

Special tokens

Common special tokens:

  • Beginning-of-sequence (<|begin_of_text|>, <s>, <bos>)
  • End-of-sequence (<|end_of_text|>, </s>, <eos>)
  • Padding (<pad>)
  • Unknown, if used (<unk>)
  • System/user/assistant role markers (e.g., <|start_header_id|>user<|end_header_id|>)
  • Tool-call delimiters
  • Function-call delimiters
  • Code block tokens
  • Image/audio placeholders for multimodal models
  • Safety or control tokens, if used

Reserve special-token slots up front

Adding new special tokens after pretraining means re-initializing those embeddings. Reserve a generous block of unused IDs in the tokenizer (Llama 3 reserves several hundred) so you can add <tool_call>, <thinking>, vision placeholders, etc. without disturbing existing IDs.

Chat template compatibility

A bad chat template can hurt post-training. You need a stable serialization format for:

<system>
You are a helpful assistant.
</system>
<user>
Question...
</user>
<assistant>
Answer...
</assistant>

For tool models, you also need stable representations for:

  • Tool definitions
  • Tool calls
  • Tool results
  • Errors
  • Parallel calls
  • Invalid calls
  • Refusals
  • Natural-language answers after tool use

Pick the chat template before SFT, not during

Changing the chat template after SFT means retraining or carefully migrating. Decide on the role markers, tool-call format, and end-of-turn token before you generate or curate any SFT data.

Practical tips

  • Train the tokenizer on a representative sample, not just English web. If your corpus is 30% code and 20% non-English, the tokenizer training corpus should be too.
  • Validate compression per language and per language family. A vocab that wastes 3× tokens on Hindi or Tamil text will waste 3× compute and quality there.
  • Run a tokenizer on a real long document and inspect the tokens. Watch for over-fragmentation in code, URLs, or repeated whitespace.
  • Test special-token round-tripping. Ensure your role markers, tool delimiters, and EOS tokens encode and decode without leaking into normal text.
  • Test the chat template through the full training stack. A subtle off-by-one in BOS/EOS handling can cost percentage points on chat evals and is invisible from the loss curve.

Further reading