A language model is a machine for doing arithmetic on numbers. Before it can read a single sentence, the text has to become numbers. The first step is to cut the text into pieces called TokenThe unit of text a language model reads and writes. A token can be a whole word, part of a word, a single character, or punctuation, often including a leading space.Open in glossary and give each piece an id from a fixed list, the VocabularyThe fixed list of tokens a tokenizer and model know. Each token's position in the list is its id.Open in glossary. The program that does this is a TokenizerThe program that splits text into tokens and maps each one to an id in a vocabulary, and turns ids back into text.Open in glossary, and its choices shape everything the model sees.
Three ways to cut text
There are three obvious ways to cut “unbelievable cats”.
By character: u n b e l i e v a b l e ␣ c a t s. The vocabulary is tiny, and any text can be spelled. But sequences get very long, and the model has to rediscover from scratch that c a t s means something.
By word: unbelievable ␣cats. Sequences are short, but the vocabulary has to hold every word the model might ever meet. English alone has hundreds of thousands, before counting names, typos, code, and other languages. Any word outside the list becomes an “unknown” token, and its meaning is lost.
By subword: un believ able ␣cats. Common words stay whole, and rare words are spelled from common pieces. This middle path is what almost every modern language model uses. The question is how to choose the pieces, and the most widely used answer is surprisingly simple.
Byte-pair encoding
Byte-pair encodingA way to build a subword vocabulary by starting from single characters or bytes and repeatedly merging the most frequent neighboring pair into a new token.Open in glossary (BPE) started life as a data compression trick (Philip Gage, 1994) and was adapted for machine translation by Sennrich, Haddow, and Birch in 2016. It builds a vocabulary from the bottom up:
- Start with every character as its own token.
- Count every pair of neighboring tokens in the training text, within each word (merges never cross a space).
- Merge the most frequent pair into one new token and add it to the vocabulary.
- Repeat until the vocabulary is the size you want.
The list of merges, in order, is the tokenizer. To tokenize new text, you replay the same merges on it.
Build a tokenizer with byte-pair encoding
Start from single characters. Each step merges the most frequent neighboring pair into a new token.
Every character is its own token. The vocabulary is the 11 distinct characters in the text.
Try this
- With Word family, predict the first merge before you press Next merge. It is “lo”, then “low”, because “low” hides inside six of the words. Keep going and notice that the endings do not form the way grammar would split them: this text yields “st” and glues each “e” onto its stem (“lowe”, “newe”). BPE follows counts, not grammar.
- Press Run all and watch the chart. The first merges save many tokens each. Later merges save less and less, because they target rarer pairs.
- Switch to Alice. The ending “ing” forms within the first ten merges because it appears in five different words. Run all the merges: words that appear twice, like ” pictures”, end up as single tokens, while words that appear only once, like “beginning”, stay in pieces, because this demo stops once no pair appears more than once.
- Choose Your text and type the same word several times. It becomes one token quickly. A word you type once never does.
The chips mark each space with ␣. Like GPT-style tokenizers, this one keeps the space attached to the start of the following word, so ” cat” (with a space) and “cat” are different tokens. That is why the space shows up inside the chips instead of between them.
Looking through a real tokenizer
Here is a production tokenizer running in your browser. GPT-4o’s tokenizer has a vocabulary of about 200,000 tokens; GPT-2’s, from 2019, has 50,257. Switch between them and compare.
See text the way a model does
Type anything. Each colored chip is one token, the unit a language model reads and writes.
Loading the tokenizer vocabulary...
Try this
- Start with Plain English. Most common words are a single token, and an average token is about four characters. That ratio is a useful rule of thumb for English.
- Open Numbers. Long numbers are split into chunks of digits, so “12345” is not one token. This is one reason arithmetic is harder for language models than it looks.
- Try Spaces and case. “Hello”, ” hello”, and ” HELLO” become different tokens with unrelated ids, and the capitals even split into two. The model has to learn that they mean nearly the same thing.
- Open Languages and switch between GPT-4o and GPT-2. The same sentences take far more tokens in the older tokenizer, especially Japanese and Arabic, where a single character can need several byte tokens (shown with a dashed outline).
- Open Code and compare the two tokenizers on the indentation. Newer tokenizers learned that runs of spaces are common in code.
Why tokens matter
Tokens are the unit that everything else is measured in:
- Cost and limits. Language model APIs charge per token, and a model’s context window, the amount of text it can consider at once, is measured in tokens. A language that needs twice as many tokens costs twice as much to process and fits half as much text.
- What the model can see. A model that receives “strawberry” as a couple of tokens never directly sees its letters. Spelling, counting letters, and rhyming all have to be learned indirectly, which is one reason models stumble on tasks that seem trivial to people.
- The starting point for meaning. Each token id is just an index into the vocabulary. In the next lessons you will see how each id is turned into a list of numbers, an EmbeddingA list of numbers (a vector) that represents a token, word, sentence, or image, learned so that similar things end up with similar vectors.Open in glossary, and how those numbers come to capture meaning.
Key ideas
- Models read tokens: pieces of text, each with an id in a fixed vocabulary.
- Subword tokens balance short sequences against a manageable vocabulary and can spell any word.
- Byte-pair encoding builds the vocabulary by repeatedly merging the most frequent neighboring pair.
- Real tokenizers work on bytes, attach spaces to words, and split numbers and other languages in ways that affect cost and behavior.