AIGenerative AIDiffusion1 min read

How AI Turns Noise into Images

Type a sentence, wait a few seconds, and a picture appears that nobody has drawn. Many of today's image generators make it in a surprising way: they start from pure random noise, like the static on an old TV, and remove the noise step by step until a picture is left. This method is called diffusion.

To learn this, the model first studies the opposite direction: real pictures getting noisier. Every pixel gets a small random nudge, and with more noise the picture fades away, until only noise is left. Move the noise level.

Interactive
Picture × 1.00
+
Noise × 0.00
=
Noisy picture

Every pixel gets its own random number from a bell curve. The two shares follow a common noise schedule and always satisfy a² + b² = 1. At 100%, nothing of the picture is left. The fox was made with Stable Diffusion XL.

This direction is easy. It is just adding random numbers, and a computer can do it to any picture without learning anything. The hard direction is the way back.

Training goes like this: take a real picture, add a random amount of noise and ask the model to guess the noise that was added, or, which comes to the same thing, what the clean picture looked like. Compare the guess with the truth and nudge the weights, exactly as in the article on how a neural network learns. After millions of pictures, the model is good at one thing: looking at a noisy picture and guessing what is underneath.

To see what such a guess looks like, let's shrink the world. Here a "picture" is just a single dot, and the training pictures are dots that together draw a heart. In a world this small we can calculate the best possible guess exactly, without training anything. Drag the noisy dot and change the noise level.

Interactive
Some noise: the nearer part of the heart becomes more likely, and the guess moves toward it.
  • The noisy dot
  • The model's best guess of the clean dot
  • Training dots; brighter ones are more likely the source

In this tiny world, an "image" is a single dot, and the training images are 120 dots that draw a heart. Here the best guess can be calculated exactly: a weighted average of the training dots. Real models have to learn their guesses.

With a lot of noise, the guess lands near the middle. The noisy dot could have come from almost anywhere, so the best guess is a blend of everything. With little noise, the guess snaps to the nearest part of the heart. Real image models behave the same way: at high noise they can only guess the rough layout and colors, and the details come in later, when there is less noise.

Making a new picture is now a loop. Start with pure noise. Guess the clean result, move a little toward it and keep a bit less noise than before. Repeat until the noise is gone. Choose the number of steps and press Run.

Interactive
Number of steps
Step 0 / 10
Press Run to turn the noise into a picture.

400 dots start as pure noise. At every step each dot moves toward the model's best guess and keeps a little less noise (the DDIM method). The faint dots are the training dots.

With one step, every dot jumps to the blend in the middle: a single guess from pure noise can only give the average. With more steps, each guess is made with a little less noise, and the dots find the heart.

A real image model behaves the same way. Here is Stable Diffusion XL with the same prompt and the same starting noise, but a different number of steps. Move the slider.

Interactive
Prompt: a lighthouse on a rocky cliff at sunset, oil painting
seed 5
seed 6
4 steps: only a foggy layout. Where the lighthouse and the rocks go is already decided, but there are no details yet.

Real pictures made with Stable Diffusion XL (fast-sdxl on fal.ai), guidance scale 7. Each picture is its own run with that many steps, not a snapshot from the middle of a longer run. With fewer than four steps, this setup gave broken pictures, so they are left out.

With few steps you get exactly what the dots showed: a blurry blend, with the rough layout already decided. More steps add the details, until extra steps barely change anything. This is why image generators usually take a few dozen steps, although some newer models are trained to get by with only a few.

Our tiny model has one flaw on purpose: it knows its 120 training dots exactly, so every dot ends up on one of them. A real model cannot store billions of pictures like that. It learns a smoothed sense of what pictures look like, and that is what lets it make pictures that never existed. Even so, a picture that appeared many times in the training data can sometimes come out almost unchanged.

Real models add two more ideas. They do not work on every single pixel but on a compressed version of the picture, and the text you type steers every guess. Those are the topics of the next articles. Some newer models, like Stable Diffusion 3 and FLUX, use a close relative of diffusion called flow matching, but the idea of turning noise into a picture step by step stays the same.

A common belief is that image generators cut out pieces of stored pictures and glue them together. That is not how diffusion works: the model keeps weights, not a folder of pictures, and builds every new picture out of noise, one guess at a time.


If this content helped you, you can buy me a coffee.
RelatedStable Diffusion Prompt Guide
RelatedSeeds and Image Editing, Explained

You can join the newsletter to be notified of awesome interactive articles and courses about AI, software and design. You will receive at most a few emails per month.

Join 800+ curious readers.


Join our Supporters