How LLMs Pick the Next Word
When you chat with an AI, the answer often shows up piece by piece, as if someone is typing. What happens behind it is simpler than it looks. A large language model, or LLM, does one small thing again and again: it looks at the text so far and guesses what comes next. In this content, we will see how that guess is made.
First, models do not read letters or whole words. They read tokens. A token is a small piece of text. It is often a short word with the space in front of it, sometimes a part of a longer word, and sometimes a single symbol. Type anything below to see how it is split.
Real tokens from o200k_base, the tokenizer used by GPT-4o. Other models use other tokenizers, so their counts can be quite different. A dot (·) is a space, and ×2 means one character was split across two tokens.
Common English words are usually one token each. Long or rare words are split into pieces, numbers are cut into small groups of digits, and many emoji need more than one token. In this example, the Turkish sentence needs more tokens than the English one. Different tokenizers split the same text differently, so counts vary across languages and models. This matters, because a model can only read a limited number of tokens at once, and many AI services charge per token.
Now the main idea. At every step, the model gives a chance to every token it knows. Not just one answer, but a whole list of possibilities, each with a percentage. Pick a sentence below and look at the most likely next tokens.
- mat31%
- floor14%
- couch9.0%
- bed8.0%
- edge5.0%
- table4.0%
- chair3.5%
- ground3.0%
- roof2.5%
- sofa2.0%
Example probabilities, chosen by hand to show the idea. A real model gives a number to every token in its vocabulary.
In the France example, one token gets almost all of the chance. In the food example, the chances are spread out. A real model learns patterns from training text and uses the text so far to predict what comes next. The model itself is not looking anything up, though a chatbot can also use tools to search for information.
If the model always took the most likely token, it would almost always give the same answer to the same question, and its writing would often feel flat and repetitive. So many apps let it pick a token at random, following the chances. A setting called temperature controls how adventurous this choice is. A low temperature makes likely tokens even more likely. A high temperature flattens the chances, so unusual tokens get picked more often. Move the slider and sample a few times.
- mat38%
- floor17%
- couch11%
- bed9.8%
- edge6.1%
- table4.9%
- chair4.3%
- ground3.7%
- roof3.0%
- sofa2.4%
Press Sample ×20 to let the model pick the next token 20 times.
Example probabilities. To keep it simple, this demo only uses the 10 tokens shown. For T > 0, each probability p becomes p^(1/T), and then they are scaled to add up to 100%. At T = 0, this demo simply picks the most likely token.
This is why the same question can get a different answer each time. A low temperature makes the choices less random; a higher one gives more variety. But less randomness does not guarantee correct code or true facts.
Random picking has a small risk. Tokens with a tiny chance can still be picked now and then, and one odd token can send the whole sentence in a strange direction. Two common settings simply cut off the unlikely end of the list. Top-k keeps only the k most likely tokens. Top-p keeps the smallest set of the most likely tokens whose chances add up to at least p.
- pizza56%
- pasta22%
- sushi22%
- chicken–
- definitely–
- probably–
- steak–
- Mexican–
- ice–
- tacos–
Example probabilities, using only the 10 tokens shown.
Top-p adapts to the situation. When the model is sure, one or two tokens are enough. When it is unsure, more of them stay in the race. Many AI tools let you set temperature, top-k and top-p, and some of them let you use more than one at the same time.
Now let's put it all together. To write a full answer, the model repeats the same steps: read the text so far, give every token a chance, pick one, add it to the text and start again. Every token in a long answer is made this way, one after another.
- there58%
- in21%
- ,9.0%
- a4.0%
Example probabilities. Here the model always takes the most likely token (this is called greedy decoding). With sampling, it could take another branch, and the rest of the sentence would change.
That is how an LLM writes. It picks a likely next token, again and again. This also explains one of its best-known problems. At its core, the model is built to continue text in a likely way, not to check facts. A high chance for a token does not mean the answer is true. A confident but wrong answer can still sound like a likely continuation. This is what people call a hallucination.
If this content helped you, you can buy me a coffee.
You can join the newsletter to be notified of awesome interactive articles and courses about software, design and AI. You will receive at most a few emails per month.