AIMathGPU1 min read

Why AI Is Mostly Matrix Multiplication

Language models write essays and image models draw pictures, and both can feel like magic. But look at what the computer is actually doing, and most of its arithmetic is one simple thing, repeated over and over: multiply numbers, then add them up. Arranged in grids, this is called matrix multiplication. Let's see why, starting from the smallest piece.

The smallest piece of a neural network is a neuron. It takes a few numbers as inputs, multiplies each one by its own weight and adds up the results. The total is its output. A weight says how much an input matters: a positive weight pushes the output up, a negative one pulls it down. Here is a neuron that scores how good a day is for a picnic. Move the inputs.

Interactive
0.9 × 2 = 1.8
0.4 × 1 = 0.4
0.7 × −2 = −1.4
Output: the sum0.8
Bad day for a picnicGood day

A toy example with weights chosen by hand. In a real network, the weights are learned during training, and the inputs rarely have names like these. Real neurons also add one extra learned number, called a bias.

That is all a neuron computes: multiply, then add. This is called a weighted sum. In a real network, nobody picks the weights by hand. During training, the network adjusts them little by little until its outputs become useful. The parameters of a model, the billions you hear about, are mostly these weights.

One neuron cannot do much on its own. A layer has many neurons, and they all look at the same inputs. Write the weights of each neuron as a row, stack the rows, and you get a grid of numbers: a matrix. Running the whole layer means doing a weighted sum for every row. That is a matrix multiplied by a list of numbers. Step through it.

Interactive
WeightsInputsOutputs
2
1
−2
0.9Sunny
Picnic
1
2
−1
×
0.4Warm
=
Beach
1
0
3
0.7Windy
Kite
Each row holds the weights of one neuron. Press Next step to run the first one.

Toy weights chosen by hand. The first row is the picnic neuron from the demo above.

Every output came from the same recipe: walk along a row, multiply each weight by its input, add everything up. Real layers are just much bigger. As we saw in the embeddings article, a language model turns each token into a list of hundreds or thousands of numbers, so a single matrix in one layer can hold millions of weights.

A model also works on many tokens at once. Put their lists side by side and they form a matrix too, so the work becomes a matrix times a matrix. Even attention, from the previous article, is built from matrix multiplications: the queries, keys and values are made by multiplying with weight matrices, and comparing every query with every key and mixing the values are matrix multiplications as well.

Between the multiplications there are a few other steps: softmax, small functions that bend the numbers (one popular choice turns every negative number into zero) and a step that keeps the numbers in a healthy range. They matter: without steps like these, many layers of matrix multiplication stacked on top of each other would be no more powerful than a single one. But they need very little arithmetic. Change the width of the model and compare.

Interactive
Share of the arithmetic that is matrix multiplication95.3%
Matrix multiplication
354,300
Everything else
17,500

A rough count of operations for one layer of a GPT-style transformer, for one new token with 1,000 tokens of text before it. Each multiply and each add counts as one operation. "Everything else" is softmax, activation functions, normalization and a few additions, counted generously. This counts arithmetic, not time.

The matrices grow with the square of the width, while the other steps grow much more slowly. So the bigger the model, the larger the share of matrix multiplication. In today's large models, with widths in the thousands, it is more than 99% of the arithmetic in this count. That is arithmetic, not time: moving all those weights in and out of memory takes time too.

This is also why AI runs on GPUs. In a matrix multiplication, every output is its own weighted sum, and none of them has to wait for another. So the work can be split across thousands of small cores that all calculate at the same time, which is exactly what a GPU is built for. Many AI chips go further and have parts made for nothing but matrix multiplication.

How much arithmetic is it in total? A handy rule of thumb: to produce one token, a language model does about two operations, one multiply and one add, for every parameter. Pick a model size and count.

Interactive
Model size (parameters)
Multiplications and additions
14 billion
By hand, one per second, without a break
444 years
A chip doing 100 trillion per second
140 microseconds

A rule of thumb, not a measurement: about two operations per parameter for each token. It leaves out the attention part, which grows with the length of the text. Modern AI chips can do 100 trillion operations per second or more, but real speed is often limited by how fast the weights can be read from memory, so real times are longer.

So under the hood, a large part of AI is one simple operation done an enormous number of times: multiply, add, repeat. What makes a model useful is not a special kind of math. It is the values of its weights, the billions of numbers that training found.

That does not make AI simple, though. People sometimes say "it is just matrix multiplication" as if that explains everything. The operation is simple, but what billions of learned weights do together is hard to predict, even for the people who build these models.


If this content helped you, you can buy me a coffee.
RelatedIntroduction to CUDA

You can join the newsletter to be notified of awesome interactive articles and courses about software, design and AI. You will receive at most a few emails per month.


Join our Supporters