Every lesson so far has used small, hand-built examples. This one uses the real thing. SmolLM2-135M is a 135-million-parameter transformer released by Hugging Face in 2024. It is about the size of GPT-2 small, but it was trained on 2 trillion tokens of web text, educational pages, and code, nearly 15,000 tokens for each of its parameters. Below, it runs entirely inside your browser.
Its architecture is the one from the transformer block: 30 blocks with a residual stream 576 numbers wide, 9 query heads that share 3 key/value heads, a SwiGLU MLP, RMSNorm, rotary position embeddings, and a 49,152-token vocabulary whose embedding table is reused for the unembedding. That adds up to 134,515,008 parameters. The version your browser downloads stores the weights as 16-bit numbers, which we checked gives the same probabilities as the full 32-bit model (more on this at the end).
A real language model, running in your browser
SmolLM2-135M reads your text and scores every token in its vocabulary as the next one.
Runs a real model in your browser
Type any text and see what a real language model predicts next, with its actual probabilities. This uses SmolLM2-135M, a 135-million-parameter base language model from Hugging Face, released in 2024.
The first time, your browser downloads about 275 MB of model files from Hugging Face, plus a 14 MB runtime from this site. Both are cached for later visits. Everything runs on your device; nothing you type is sent anywhere.
Try this
- Leave the starting text and read the candidates. Several continuations get real probability, and none is certain. Select one to append it, and watch the whole distribution change.
- Replace the text with
Once upon aand compare the uncertainty readout with the default text. Then tryThe capital of France is. Paris is in the running, but so are tokens that begin other true sentences, such as ” located”. Then tryTwinkle, twinkle, little: ” star” leads, but with only about 29%. A model this small spreads its bets more than you might expect. - Set temperature to 0 and press Sample 12 tokens. The model always takes its top choice, so the continuation is fully determined by your text. Undo it, raise the temperature to 1.5, and generate again: the text wanders further from the likely path.
- Type
What is the capital of France?and generate. Then tryQ: What is the capital of Italy?on one line,A: Romeon the next, thenQ: What is the capital of France?andA:on the lines below.
What you are looking at
Tokens. The chips under your text show exactly what the model receives. A dot marks a token that starts with a space; most common words are one token that includes the space before them, and rarer words split into pieces. The model never sees letters, only these 49,152 possible chunks, each mapped to an ID.
Probabilities. For the last position, the model outputs one LogitA raw, unnormalized score a model outputs for one option, such as one token in the vocabulary. Softmax turns a vector of logits into probabilities.Open in glossary per vocabulary entry, and SoftmaxA function that turns a list of real-valued scores into probabilities that are positive and sum to 1, by exponentiating each score and dividing by the total.Open in glossary turns them into probabilities. The bars show those probabilities at temperature 1, computed over the full vocabulary, so they are the model’s own numbers. The readout “probability held by the top 20” tells you how much of the distribution you are seeing. When it is near 100%, the model has narrowed down to a few options. When it is low, probability is spread thin across thousands of tokens.
Uncertainty. The first readout is the EntropyA measure of how uncertain a probability distribution is, in bits. Zero means one outcome is certain; k bits is as uncertain as a fair choice among 2 to the power k outcomes.Open in glossary of the full distribution, in bits:
It is 0 when one token has all the probability and , about 15.6 bits here, when all 49,152 tokens are equally likely. A useful reading: bits is as uncertain as choosing evenly among tokens. Even a near-certain continuation is not 0 bits for this model: after “Once upon a”, ” time” gets about 88%, but thousands of unlikely tokens share the rest, and the entropy comes out near 1.4 bits. After “The capital of France is” it is about 6.4 bits. That is not a failure; many continuations really are reasonable there.
Sampling. Your temperature, top-k, and top-p settings work exactly as in predicting the next token, applied to the logits of the 20 most likely tokens. The last column shows each token’s chance of being picked with your settings, and “cut” marks tokens your settings remove.
A base model, not an assistant
Ask this model a question and it may answer, ramble, or ask three more questions. That is not a bug. SmolLM2-135M is a Base modelA language model after pretraining only. It continues text in whatever style the input suggests, rather than following instructions or answering like an assistant.Open in glossary: it has only been pretrained, so its single skill is continuing text the way documents on the web continue. A question followed by more questions is a common pattern, in quizzes, forums, and FAQ pages.
You can still get answers by making the answer the natural continuation. The question-and-answer layout in the last item above works because, in a document that already shows one question answered after “A:”, the most likely thing after the next “A:” is an answer. Putting worked examples in the prompt like this is called few-shot prompting, and its power in large models was one of the central findings of the GPT-3 paper (Brown et al., 2020). The instruction-following you are used to from chatbots comes from the fine-tuning stages in the next lesson.
What this demo simplifies
The probabilities are real, but a few things differ from how production systems generate text:
- Sampling is limited to the top 20 tokens. The model scores all 49,152 tokens, but only the 20 most likely are sent to the page. A real sampler with top-k and top-p turned off could occasionally pick a token outside them; here it cannot, so top-k at its maximum of 20 is not quite the same as top-k off. The coverage readout tells you how much probability that leaves out.
- Text is re-read from scratch each step. The demo appends a chosen token’s text and sends the whole text again, so the model re-tokenizes it and reruns every position. Real generation keeps the token IDs and reuses earlier computation with a KV cache. The cache saves work but never changes the results; very occasionally, though, re-tokenizing joined text can split it differently than the tokens that were chosen.
- Weights are 16-bit. We compared this version with the full 32-bit model on 24 prompts, and the probabilities matched. Smaller 8-bit and 4-bit versions exist and would cut the download to half or less, but they changed the top prediction on a third to nearly half of those prompts (after “Twinkle, twinkle, little”, the 8-bit version gives ” star” about 3% instead of 29%). A lab about a model’s probabilities needs the faithful version, so this one uses the larger download.
Key ideas
- A real language model does exactly what the earlier lessons described: tokens in, one logit per vocabulary entry out, softmax, then a sampling rule.
- The next-token distribution can be nearly certain or spread across thousands of tokens. Entropy in bits measures which.
- A base model continues text. Formatting the prompt so the answer is the natural continuation, as in few-shot prompting, gets answers out of it.
- The model has no lookup table of facts. Right and wrong continuations come from the same learned probabilities.