A language model produces its answer one TokenThe unit of text a language model reads and writes. A token can be a whole word, part of a word, a single character, or punctuation, often including a leading space.Open in glossary at a time, and each new token passes through the same fixed stack of layers, so the computation available per token is roughly fixed (attention grows somewhat with the length of the text, but the depth does not). Ask a hard multi-step question and demand the answer immediately, and the model has to do all the work in the few tokens before it commits. Researchers found several ways to give models more room, all of which spend more computation at answer time. This family of ideas is called Test-time computeComputation a model spends while answering, rather than during training, such as writing out reasoning, sampling several answers, or searching over candidates. More of it often improves accuracy on hard problems.Open in glossary.
Writing out the steps
The first idea is to let the model write its reasoning before the answer. Wei et al. (2022) showed that including a few worked examples with step-by-step solutions in the prompt, which they called Chain of thoughtIntermediate reasoning steps a language model writes before its final answer. Prompting for them, or training models to produce them, often improves results on multi-step problems.Open in glossary prompting, sharply improved large models on math word problems and other multi-step tasks. Kojima et al. (2022) found that even appending “Let’s think step by step” to a question helped.
One common explanation is that the written steps act as working memory. Every token the model writes becomes part of its input for the next token, so intermediate results do not have to be held inside a single forward pass. This explanation fits the evidence but is not the whole story, and why chain of thought helps on some tasks and not others is still an active research question.
Asking several times and voting
If you sample with a TemperatureA setting for how adventurous a model's choices are. Low temperature sticks to the likeliest options; high temperature spreads choices toward less likely ones. Technically, it divides the logits before softmax.Open in glossary above zero, the same question can produce different reasoning paths and different final answers. Self-consistencySampling several reasoning paths for the same question and returning the most common final answer. It helps when the correct answer is the model's most likely answer.Open in glossary (Wang et al., 2022) samples several chains of thought and returns the most common final answer.
How much does that help? If you treat each sample as an independent draw that is right with probability , the answer can be computed exactly. With two possible answers and an odd number of samples , the vote is right when more than half of the samples are:
With more possible answers the same idea holds, summing over every possible tally of answers (a multinomial), and splitting ties evenly. The demo does exactly that.
Majority vote over sampled answers
If each sampled chain of thought is right with probability p, how often is the most common answer among n samples right?
Try this
- With 2 answers, set p to 60% and slide n upward. The curve climbs toward 100%. Now set p to 40%: the same voting drives accuracy toward zero.
- Switch to 4, spread with p at 40%. Even though one sample is wrong most of the time, the vote improves, because each wrong answer is only 20% likely and the correct answer is the most common one.
- Switch to 4, one tempting and set p to 30%. One wrong answer (49%) is now more likely than the right one, and more samples drive the vote toward it: accuracy falls to about 8% at 41 samples. Then try p = 40%. The race is close (40% against 42%), so the vote first improves slightly as the rare wrong answers drop out, and only then slowly declines.
- Press Run 1,000 votes and compare the simulated accuracy with the exact value. They should usually agree within about three percentage points; with 1,000 votes, random sampling alone moves the result by about 1.6 points on a typical run.
The rule behind every curve: voting converges on the most likely answer. With enough samples, it helps exactly when the correct answer is the single most common output, whether or not that answer is right more than half the time. It cannot rescue a model whose most common answer is wrong.
Models trained to reason
Prompting and voting work on any model. A newer approach trains the model itself to produce long reasoning before answering. OpenAI’s o1 (previewed in September 2024, released in full that December) and DeepSeek-R1 (January 2025) are prominent examples. DeepSeek’s report describes reinforcement learning in which the model is rewarded for reaching correct answers on problems that can be checked automatically, such as math with known answers and code with tests. A variant trained with reinforcement learning alone (R1-Zero) already learned long reasoning; the released R1 added supervised fine-tuning stages before and between rounds of reinforcement learning to make its output more readable. Trained this way, models learned to write long chains that include checking their work and backtracking.
OpenAI also reported that accuracy kept rising as the model was allowed to reason for longer, roughly in proportion to the logarithm of the extra computation over the range they tested. This is described as scaling with test-time compute, alongside the familiar scaling with training compute. The cost is real: long reasoning can mean thousands of extra tokens per answer, which costs time and money.
Choosing among samples without votingOptional
Majority voting needs no extra model, but it throws away information about which reasoning was sound. Alternatives include best-of-n, where a separate scoring model (often called a verifier or reward model) rates each candidate and the top one is returned, and search, where partial reasoning is scored step by step and promising branches are extended. These can beat voting when a good verifier exists, and can be fooled when the verifier has blind spots the generator learns to exploit.
Key ideas
- Each token passes through a fixed number of layers, so writing intermediate steps gives a model more computation and a place to keep intermediate results.
- Self-consistency samples several reasoning paths and takes the most common answer.
- Voting converges on the most likely answer: it helps when that answer is correct and hurts when it is not.
- Real samples are correlated, so voting gains are smaller than the independent-sample math predicts.
- Reasoning models are trained, often with rewards for verifiably correct answers, to think at length, trading extra tokens for accuracy.