Tokenizer training¶
Goal¶
Create a vocabulary that compresses text into tokens efficiently across languages, code, math, and markup.
Common tokenizer families:
- BPE (Byte Pair Encoding) — the most common choice for modern LLMs.
- Unigram / SentencePiece — alternative scoring; SentencePiece is the de-facto reference implementation.
- WordPiece-like variants — used by older BERT-family models, less common for new LLMs.
- Byte-level BPE — operates on bytes, so any input is representable; used by GPT-2 onward.
Modern LLM tokenizers usually support arbitrary bytes so the model can handle unseen characters.
Design choices¶
Vocabulary size¶
Typical modern vocab sizes range from roughly 32K to 200K+ tokens.
| Model | Vocab size |
|---|---|
| GPT-2 | 50,257 |
| Llama 2 | 32,000 |
| Llama 3 | 128,256 |
| Gemma 2 | 256,000 |
| Mistral | 32,000 |
Larger vocabularies:
- Improve compression
- Help multilingual text
- Help code and markup
- Increase embedding/output matrix size
- Can make rare-token learning harder
Smaller vocabularies:
- Are simpler
- Reduce embedding size
- May be inefficient for non-English languages and code
Compression matters more than you'd think
A bigger vocabulary that better compresses your training corpus = more information per training step at fixed sequence length, and lower inference cost per output character. The Llama 3 paper explicitly attributes part of its multilingual gains to the 128K vocab.
Special tokens¶
Common special tokens:
- Beginning-of-sequence (
<|begin_of_text|>,<s>,<bos>) - End-of-sequence (
<|end_of_text|>,</s>,<eos>) - Padding (
<pad>) - Unknown, if used (
<unk>) - System/user/assistant role markers (e.g.,
<|start_header_id|>user<|end_header_id|>) - Tool-call delimiters
- Function-call delimiters
- Code block tokens
- Image/audio placeholders for multimodal models
- Safety or control tokens, if used
Reserve special-token slots up front
Adding new special tokens after pretraining means re-initializing those embeddings. Reserve a generous block of unused IDs in the tokenizer (Llama 3 reserves several hundred) so you can add <tool_call>, <thinking>, vision placeholders, etc. without disturbing existing IDs.
Chat template compatibility¶
A bad chat template can hurt post-training. You need a stable serialization format for:
<system>
You are a helpful assistant.
</system>
<user>
Question...
</user>
<assistant>
Answer...
</assistant>
For tool models, you also need stable representations for:
- Tool definitions
- Tool calls
- Tool results
- Errors
- Parallel calls
- Invalid calls
- Refusals
- Natural-language answers after tool use
Pick the chat template before SFT, not during
Changing the chat template after SFT means retraining or carefully migrating. Decide on the role markers, tool-call format, and end-of-turn token before you generate or curate any SFT data.
Practical tips¶
- Train the tokenizer on a representative sample, not just English web. If your corpus is 30% code and 20% non-English, the tokenizer training corpus should be too.
- Validate compression per language and per language family. A vocab that wastes 3× tokens on Hindi or Tamil text will waste 3× compute and quality there.
- Run a tokenizer on a real long document and inspect the tokens. Watch for over-fragmentation in code, URLs, or repeated whitespace.
- Test special-token round-tripping. Ensure your role markers, tool delimiters, and EOS tokens encode and decode without leaking into normal text.
- Test the chat template through the full training stack. A subtle off-by-one in BOS/EOS handling can cost percentage points on chat evals and is invisible from the loss curve.
Further reading¶
- Sennrich et al., "Neural Machine Translation of Rare Words with Subword Units", 2015. The original BPE paper. arxiv.org/abs/1508.07909
- Kudo, "Subword Regularization", 2018. The Unigram tokenizer. arxiv.org/abs/1804.10959
- Kudo & Richardson, "SentencePiece", 2018. The de-facto reference implementation. arxiv.org/abs/1808.06226
- HuggingFace
tokenizers— fast Rust BPE/Unigram/WordPiece training. github.com/huggingface/tokenizers - Andrej Karpathy — "Let's build the GPT Tokenizer" (2-hour video). The single best practical introduction. youtube.com/watch?v=zduSFxRajkE
minbpe— Karpathy's minimal BPE implementation, paired with the video. github.com/karpathy/minbpetiktoken— OpenAI's BPE tokenizer, used by GPT-3.5/4. github.com/openai/tiktoken