AITrainingMath1 min read

How a Neural Network Learns

In the matrix article we said that training finds the weights. But how? Nobody tells the network the right numbers. It starts with random weights, makes guesses, measures how wrong they are and adjusts the weights a tiny bit. Then it does it again, and again. Let's take this loop apart with the smallest possible network: one with a single weight.

Our network takes a number and multiplies it by its weight. That is its guess. We want it to learn from five examples, and for each one we know the correct answer. Move the weight and try to make the line pass through the dots. The amber lines show how far off each guess is.

Interactive
InputAnswer
Loss: the average squared miss68.97

Made-up examples whose answers are close to 3 × the input. White dots are the correct answers, the line is the network's guess, and the amber lines are the misses.

To learn, the network needs one number that says how wrong it is. This number is called the loss. Here we take each miss, square it and average the results. Squaring makes misses above and below the line count the same, and makes big misses count much more. The lower the loss, the better the weight. This is one common way to measure loss; language models use a different one that checks how much probability they gave to the correct next token.

You found a good weight by looking at the picture. The network cannot do that: it only knows the loss at its current weight. But it can calculate one more thing, the slope: whether the loss goes up or down when the weight grows a little. If the slope is negative, a bigger weight means a lower loss. If it is positive, a smaller one does. So the network takes a small step downhill. Press Step.

Interactive
WeightLoss
Weight
0.50
Loss
68.97
Slope
−55.1
The slope is −55.1: the loss goes down to the right, so the weight grows (+1.65).

Each step changes the weight by −0.03 × the slope. The amber curve is the loss for every possible weight; the network itself never sees this curve, only the loss and slope where it stands.

Every step goes downhill, and the steps get smaller as the curve flattens near the bottom. This is called gradient descent. A gradient is simply the slope when there are many weights instead of one. Variations of this idea are how nearly all of today's neural networks are trained.

How big should a step be? Each step is the slope multiplied by a small number called the learning rate. Pick one and press Run.

Interactive
Learning rate
WeightLoss
Press Run to take 20 steps from the same start.

Each step changes the weight by −learning rate × the slope. All runs start at weight 0.5. The best learning rate depends on the problem; these values only fit this example.

Too small, and training takes forever. Too big, and the weight jumps over the bottom, back and forth. Much too big, and every jump lands higher than the last, until the numbers blow up. Choosing the learning rate is one of the most important decisions when training a network.

A real network does the same with billions of weights at once. Every weight needs its own slope: how the loss would change if only that weight changed a little. A method called backpropagation calculates all of these slopes in one pass backwards through the network, again mostly with matrix multiplications. Then every weight takes its small step.

Real training also does not look at all the examples for every step. It takes a small batch of them, measures the loss and the slopes on that batch, takes a step and moves on to the next batch. Training a large language model repeats this loop hundreds of thousands of times or more, over huge amounts of text.

So for a neural network, learning is this loop: guess, measure how wrong the guess is, nudge every weight a little downhill, repeat. Everything the network ends up knowing is stored in where its weights settled.

One common misunderstanding is that the lowest possible loss is the goal. A network can learn its training examples too well, including their random noise, and then do worse on examples it has never seen. This is called overfitting, and it is why networks are tested on examples that were kept out of training.


If this content helped you, you can buy me a coffee.
RelatedWhy AI Is Mostly Matrix Multiplication
RelatedSeeds and Image Editing, Explained

You can join the newsletter to be notified of awesome interactive articles and courses about AI, software and design. You will receive at most a few emails per month.

Join 800+ curious readers.


Join our Supporters