Attention, Visually
Some words change their meaning with the words around them. In "I sat on the river bank", a bank is the side of a river. In "I put money in the bank", it is a place for money. A language model has to work this out for every word, every time. The tool it uses for this is called attention.
The idea is simple. When the model looks at a word, it also looks at the other words in the text and decides how much each one matters. We use whole words here to keep it simple; real models do the same with tokens. Pick a sentence and see how much attention "bank" pays to each word.
Example weights, chosen by hand. The purple box is the word that is looking; darker words get more of its attention, and all its weights add up to 100%.
In the river sentence, "bank" looks mostly at "river". In the money sentence, it looks mostly at "money". Notice the words after "bank": they get nothing. Models that write text one token at a time, like the ones in the earlier articles, can only let each word look at itself and the words before it.
How does the model decide these numbers? Every word makes three lists of numbers. A query says what the word is looking for. A key says what the word has to offer. A value is the information the word passes on when another word looks at it. To decide how much "bank" looks at another word, the model compares the query of "bank" with the key of that word.
The comparison is the arrow idea from the embeddings article: when two arrows point the same way, the score is high. Then a function called softmax turns all the scores into weights that add up to 100%. Turn the query and watch where the attention goes.
| Key | Score | Weight |
|---|---|---|
| river | 0.99 | 81% |
| money | 0.48 | 18% |
| the | -0.45 | 1% |
Example vectors with two numbers each. The score is the dot product of the query and each key. Softmax then turns the scores into weights that add up to 100%. Here the scores are multiplied by 3 first, so the differences are easier to see; real models also scale them by a fixed number.
So nobody writes these weights by hand. They come from how well each query fits each key, and during training the model learns how to make good queries and keys for every word. That is what lets "bank" find "river" or "money" on its own.
The last step is the mix. With the weights ready, the model takes the values of the words, weighs them and adds them up. This mix is added to the word's own numbers, so "bank" now carries a bit of "river" or a bit of "money". Switch between the sentences and watch "bank" move.
I sat on the river bank and read
Example vectors with two numbers each, using the weights from the first demo. Real models have hundreds or thousands of numbers per word.
That is attention: look around, decide what matters and mix it in. Real models do this many times side by side, each time with different queries and keys. Each of these is called an attention head. They also repeat the whole step in many layers, one on top of the other. The transformer, the design behind most of today's language models, is built around this idea.
One word of caution. Pictures of attention like the ones above are fun to look at, but in a real model they do not always explain why it gave an answer. Many heads and layers work together, and one set of weights shows only a small part of the story.
If this content helped you, you can buy me a coffee.
You can join the newsletter to be notified of awesome interactive articles and courses about software, design and AI. You will receive at most a few emails per month.