When you ask a chatbot a question, the answer appears a few words at a time, as if it were typing. That is not a visual effect. It really is producing the text one small piece after another. At every step, it works out how likely each possible next piece is, picks one, adds it to the text, and does it all again.
A model that predicts what comes next in text is called a Language modelA model that assigns probabilities to what comes next in a piece of text. Writing text with one means repeatedly predicting the next word or token and picking one.Open in glossary. The ones behind chatbots are enormous. But you can build a tiny one that works on the same principle in a few lines of counting, and that is what the demo below does.
A model made of counts
The model below has read ten of Aesop’s fables, 959 words and punctuation marks in all. Its “training” was simply counting: for every word, it tallied which words came right after it. In the fables, the word “the” appears 107 times. Seven of those times it is followed by “lion”, six times by “town”, and so on through 42 different words.
Turn those counts into fractions and you have probabilities: after “the”, the next word is “lion” about 6.5% of the time. That list of probabilities is the model’s whole answer to the question “what comes next?” A model that works this way is called an n-gram modelA simple language model that predicts the next word by counting what followed the previous one or few words in its training text.Open in glossary.
Write with a word-counting model
This model learned by counting which words follow which in ten short fables. The bars show what it thinks comes next. Pick one, or let it choose at random.
The
What comes next?
and 34 other words with 60.7% between them
Try this
- Click a bar to add that word yourself. The highlighted words show what the model is looking at, and the bars update for the next step.
- Press Pick a word for me a few times. The model picks at random, but in proportion to the bars: a 30% word gets picked about 30% of the time.
- Drag Temperature to the far left and press Write 12 words. Then Start over, drag it to the far right, and write again. Compare the two.
- Switch to 2 words of context and write some more. You will see longer stretches copied straight from the fables.
Why pick at random?
If the model always took the single most likely word, it would write the same text every time and often get stuck in loops, because the most likely word after “the” leads to the most likely word after that, and so on around a circle. Picking at random, weighted by the probabilities, gives variety while still favoring sensible choices.
TemperatureA setting for how adventurous a model's choices are. Low temperature sticks to the likeliest options; high temperature spreads choices toward less likely ones. Technically, it divides the logits before softmax.Open in glossary controls how adventurous those picks are. At low temperature the probabilities are sharpened, so the top word gets nearly all the chance and the writing is safe but repetitive. At high temperature they are flattened, so unlikely words get picked more often and the writing becomes surprising, then strange, then nonsense. Chatbots typically use a moderate temperature, and many let developers adjust it.
More context: better sentences, less originality
With one word of context, the model only knows the previous word, so its sentences wander: each pair of words makes sense, but the whole rarely does. With two words of context it writes more convincing phrases, but there is a catch. Most two-word combinations appear only once or twice in such a short text, so the model often has just one option and copies the fables word for word.
That is the limit of counting. A counting model can only use contexts it has seen exactly, and the longer the context, the rarer each exact match becomes. It has no idea that “the fox” and “the wolf” are similar, so what it learned about one never helps with the other.
From counting to chatbots
Modern chatbots keep the same writing loop but replace almost everything else:
- A neural network instead of a table of counts. Large language models are neural networks, usually a design called a transformer. Instead of looking up exact matches, they learn patterns that carry over to text they have never seen.
- Much more context. Instead of one or two previous words, they consider thousands of earlier pieces of text at once, including the whole conversation.
- Pieces of words, not words. They read and write TokenThe unit of text a language model reads and writes. A token can be a whole word, part of a word, a single character, or punctuation, often including a leading space.Open in glossary, which are often whole words but can be parts of words or punctuation.
- Far more training text, then more training. They first learn to predict text from a huge collection of writing, then are trained further to follow instructions and answer helpfully.
The transformers track opens up that machinery. But at the moment of writing, a chatbot does what you just did: look at the text so far, get a probability for every possible next token, pick one, and repeat.
The math of temperatureOptional
At temperature 1, the model’s probability for each candidate word is its count divided by the total: . At temperature , each count is first raised to the power :
With below 1 the largest counts grow fastest, sharpening the distribution; above 1 the counts are pulled toward each other, flattening it. This is exactly the same as dividing the log-probabilities by before a softmax, which is how temperature is applied in large language models.
When the model has never seen the current two-word context, it falls back to one word of context, and if it has never seen that word either, to plain word frequencies. This fallback is called backoff.
Key ideas
- A language model predicts the next piece of text as a list of probabilities.
- Writing with it is a loop: predict, pick, append, repeat.
- Picking at random, in proportion to the probabilities, gives variety and avoids loops. Temperature sets how adventurous the picks are.
- Counting models only reuse exact contexts they have seen. Large language models use neural networks that generalize, and look back over thousands of tokens.