Build a tokenizer with byte-pair encoding
Start from single characters. Each step merges the most frequent neighboring pair into a new token.
Every character is its own token. The vocabulary is the 11 distinct characters in the text.
low␣lower␣lowest␣slow␣slower␣slowest␣new␣newer␣newest␣wide␣wider␣widest␣old␣older␣oldest
Vocabulary size11
Tokens in the text88
Characters per token1.00
Try this
- Predict the first three merges for Word family before pressing Next merge.
- Press Run all and read the chart: early merges save the most tokens.
- Type your own text and repeat one word several times. It becomes a single token.