Byte-pair encoding, step by step

Train a tiny tokenizer from single characters and watch merges build a vocabulary while the token count drops.

IntermediateExplained in Tokens

Build a tokenizer with byte-pair encoding

Start from single characters. Each step merges the most frequent neighboring pair into a new token.

Every character is its own token. The vocabulary is the 11 distinct characters in the text.

Training text
0 of 18
Vocabulary size11
Tokens in the text88
Characters per token1.00

Try this

  • Predict the first three merges for Word family before pressing Next merge.
  • Press Run all and read the chart: early merges save the most tokens.
  • Type your own text and repeat one word several times. It becomes a single token.

Try "embedding", "softmax", "overfitting", or "backpropagation".