Tokenizer playground

See any text split into tokens by GPT-4o's tokenizer or GPT-2's, with ids, counts, and the surprises in numbers, code, and other languages.

IntermediateExplained in Tokens

See text the way a model does

Type anything. Each colored chip is one token, the unit a language model reads and writes.

Example texts

Loading the tokenizer vocabulary...

Tokenizer
Tokens...
Characters58
Words11
Characters per tokenn/a

Try this

  • Compare the two tokenizers on Languages. The older one needs far more tokens for Japanese and Arabic.
  • Type a long number. How does each tokenizer cut it?
  • Paste a paragraph you wrote and read the characters-per-token readout.

Try "embedding", "softmax", "overfitting", or "backpropagation".