A live language model

Run SmolLM2-135M in your browser, see how it tokenizes your text, read its real next-token probabilities, and generate one token at a time.

AdvancedExplained in Inside a real language model

A real language model, running in your browser

SmolLM2-135M reads your text and scores every token in its vocabulary as the next one.

Runs a real model in your browser

Type any text and see what a real language model predicts next, with its actual probabilities. This uses SmolLM2-135M, a 135-million-parameter base language model from Hugging Face, released in 2024.

The first time, your browser downloads about 275 MB of model files from Hugging Face, plus a 14 MB runtime from this site. Both are cached for later visits. Everything runs on your device; nothing you type is sent anywhere.

Try this

  • Compare the uncertainty readout for Once upon a with The capital of France is.
  • Set temperature to 0 and generate twice, then set it to 1.5.
  • Ask a question, then rewrite it as a “Q: … A:” document and compare.

This is a base model: it continues text rather than following instructions. Probabilities are the model’s own, over its full 49,152-token vocabulary; sampling here draws from the 20 most likely tokens.

Try "embedding", "softmax", "overfitting", or "backpropagation".