AILLMInference1 min read

Prefill and Decode: The Two Speeds of an LLM

When you send a message to a chat model, there is often a short pause. Then the answer starts to stream in, a few words at a time. That pause and that stream are two different jobs, and the model does them in two different ways.

First comes prefill: the model reads your whole prompt. Then comes decode: it writes the answer, one token at a time. (A token is a small piece of text, often a word or part of a word.) It is the same model in both phases. Press Run and count how many times the text goes through the model.

Interactive
Prompt
Why·is·the·sky·blue?
Memory
Weights
KV cache
↓ read
↓ read
↑ write
GPU
idle
Answer
Sunlight·scatters·off·the·air,·and·blue·scatters·most.
PassesPrefill: 0Decode: 0

Press Run and count the passes.

Example tokens, chosen by hand, and slowed down a lot. Each pass block shows how many tokens went into the model. The first answer token comes out of the prefill pass, so it has a blue border. The KV cache is explained further down.

In prefill, every prompt token is already known, so the model can work on all of them at the same time. One pass through the model reads the whole prompt and also gives the first token of the answer. It also saves a few numbers for every prompt token in the KV cache, which we will look at further down. (Very long prompts are often split into a few large chunks, but the idea is the same.)

Decode cannot do that. The next token depends on the one before it, and that one does not exist yet. So the model makes one token, adds it to the text and runs again for the next one, until the answer is finished. Every new token after the first needs its own pass. (Some tricks, like speculative decoding, can check a few guessed tokens in one pass, but the basic loop is one token per pass.)

Why does that matter? In every pass, the model has to do two things. It has to read its weights from memory, all of them, and it has to do the math for each token in the pass. Reading the weights costs about the same no matter how many tokens there are. The math grows with every token.

Think of a cook whose ingredients are kept in a storeroom down the hall. Before every round of cooking, the cook has to carry all of them in. The walk takes the same time whether the cook makes one plate or a hundred, but a hundred plates take much longer to cook. Change the number of tokens in one pass and watch which job has to wait for the other.

Interactive
Tokens handled in one pass

Like decode: one new token for one conversation.

Reading all the weightsworking…
Doing the mathworking…
Waiting for the other job
One pass takes
16 ms
Tokens per second
63

Memory-bound: the math is done early and waits for the weights.

A simplified model, not real measurements: an example graphics card that reads 1,000 GB per second and does 100 trillion operations per second, running a model with 8 billion weights at 16 bits (16 GB). Each token needs about 2 operations per weight. The demo assumes reading and math fully overlap, and leaves out attention and the KV cache. The bars are slowed down 40 times.

With a single token, the math is tiny, and the chip spends most of its time waiting for the weights to arrive from memory. We say decode is usually memory-bound: its speed depends mostly on how fast the memory is (its bandwidth), and on how big the model is. That is also why a quantized model, with smaller weights, often writes faster.

With a long prompt, prefill puts hundreds or thousands of tokens into one pass. The weights are read once and used for all of them, so now the math is the slow part. Prefill is usually compute-bound: its speed depends mostly on how much math the chip can do per second.

Servers use the same trick for decode. With batching, they put the next token of many conversations into one pass. The weights are still read once, so the server makes many more tokens per second in total, even though each single conversation does not get faster, and can even get a bit slower.

There is one more detail. To write a new token, the model looks back at all the tokens before it, using attention. For that, it needs some numbers for every earlier token, called keys and values. Those numbers do not change, so instead of computing them again in every pass, the model saves them in the KV cache. Press Play and watch the rows grow, then turn the cache on and play it again.

Interactive
#1#2#3#4#5#6
Run through the modelTaken from the cache
Tokens run through the model, in total
Without the cache39
With the cache9

Each row is one pass that makes one answer token. Each column is one token: blue for the 4 prompt tokens, amber for answer tokens fed back in. A simplified picture: real models do this in every layer.

Without the cache, every pass runs the whole text through the model again, and the work grows very quickly as the answer gets longer. With the cache, each pass only runs the newest token. That is why real systems keep a cache: the extra memory is a good trade.

But that memory has a cost, and it grows with every token in the conversation. Long conversations cost you twice: the cache takes up more memory, and every new token has to read more of it, so writing gets slower too. Drag the slider and watch the cache grow next to the weights.

Interactive
Memory16 GB + 1 GB
24 GB
WeightsKV cacheExample card with 24 GB

The weights and the cache fit on the 24 GB card.

Time per output token
17 ms
Tokens per second
59

A model shaped like Llama 3.1 8B: 32 layers, 8 key/value heads of 128 numbers each. Per token it stores 2 (keys and values) × 32 × 8 × 128 numbers, which is 131,072 bytes at 16 bits. Weights at 16 bits. Real 8-bit cache formats also store a few scale numbers, so they save a little less than half. The time is a simplified estimate on the same example card as above (reading 1,000 GB per second); it counts only memory reads. GB here means a billion bytes.

For a model like this, a conversation of 128,000 tokens needs about as much memory for the cache as for all of its weights. Some tools, like llama.cpp, reserve the cache for the maximum length up front, so the memory is used even before the conversation gets long.

When the runtime supports it, there are a few ways to help: start a fresh conversation, summarize the older messages, set a smaller maximum length, or store the cache with fewer bits, the same idea as quantization for weights.

The two phases also explain the numbers people use to describe how fast a model feels. Time to first token is how long you wait before anything appears, and it mostly comes from prefill. Time per output token is the gap between two tokens of the answer, and it comes from decode. Together they roughly give the whole wait: time to first token plus the number of answer tokens times the time per token. Pick an example and press Send to feel the wait.

Interactive
Chat0.0 s
Press Send to feel the wait.
Prefill: waiting for the first tokenDecode: the answer streams in
Time to first token
0.23 s
Time per output token
8.7 ms
Whole answer
2.84 s

Here most of the wait is decode: writing the answer.

Plays in real time. Speeds from the llama bench example in the llama.app guide: a small 1-billion-weight model (Gemma 3 1B, quantized to about 4 bits) reading about 2,184 prompt tokens and writing about 115 answer tokens per second. Real speeds also drop as the conversation gets longer, so long prompts and long answers take somewhat longer than this straight-line estimate.

A long prompt, like a pasted document, makes you wait longer before the answer starts. A long answer makes the stream take longer. That is why one number like "tokens per second" is not enough to describe a model: always check whether it means reading the prompt or writing the answer. In the example from the guide, the model read prompt tokens about 19 times faster than it wrote answer tokens.

So the same model runs at two very different speeds. Prefill reads the whole prompt at once and is usually limited by math. Decode writes one token at a time and is usually limited by memory. The KV cache keeps decode from redoing old work, at the cost of memory that grows with the conversation. Once you know which phase you are waiting for, the pause before the first word and the speed of the stream both make sense.

Source we learned fromPrefill vs. Decode · llama.appThis content could be supported by youClaim this spot
author@aykutkardas

0/50 claps

If this content helped you, you can buy me a coffee.


You can join the newsletter to be notified of awesome interactive articles and courses about AI, software and design. You will receive at most a few emails per month.

Join 800+ curious readers.


Join our Supporters