Why Image Models Work in Latent Space
In the previous articles a model turned noise into a picture, one guess at a time. There is a catch we skipped. A 512 × 512 color picture is 786,432 numbers, and the model has to make a guess for all of them at every step. Stable Diffusion found a way around this: it does not work on the pixels at all, but on a much smaller version of the picture called the latent.
A separate model makes this small version. Its first half, the encoder, squeezes, roughly speaking, every 8 × 8 patch of pixels into just 4 numbers. Move the square over the picture to see what one latent position covers.

The whole picture: 512 × 512 × 3 = 786,432 numbers. Its latent: 64 × 64 × 4 = 16,384 numbers, 48 times fewer.
Sizes as in Stable Diffusion 1.5, which makes 512 × 512 pictures. Each latent position mostly describes its own patch, but the encoder also looks at the surroundings. The real four numbers come out of the encoder model, which does not run in this page, so we show question marks instead of made-up values.
Those 4 numbers are not a tiny picture with one color per patch. They are learned: during training, the compressor figured out which 4 numbers best describe each patch so that the picture can be rebuilt from them. That is the job of the second half, the decoder, which turns the latent back into pixels. This pair is called an autoencoder, and it is trained before the diffusion model ever sees a picture.
So when Stable Diffusion makes a picture, all the steps from the previous articles happen in latent space: the noise is latent noise, and every guess is a guess of a latent. Only at the very end does the decoder turn the result into the pixels you see. Why go through all this? Pick a picture size and compare.
Latent 64 × 64 × 4: 48 times fewer numbers and 4,096 times fewer pairs.
Exact counts for a latent that is 8 times smaller on each side with 4 numbers per position, as in Stable Diffusion 1.5 and XL. Real models do not compare every position with every other at full size; the last row shows why that would be out of reach for pixels.
48 times fewer numbers already saves a lot of work. But it gets better. Attention, from the earlier article, lets positions look at each other, and the number of pairs grows with the square of the number of positions. For pixels, that quickly becomes impossible; for a latent it is merely big. That difference is what made it possible to run image generators on an ordinary graphics card.
But is the latent not just a tiny, blurry picture? Here is the test: the same picture shrunk to a 128 × 128 grid of pixels, and the same picture sent through a 128 × 128 latent and back.

Detail at full sizeBoth small versions use the same 128 × 128 grid: the latent keeps 4 numbers per position, the small picture 3. The latent version was made with Stable Diffusion XL on fal.ai with the smallest change the service allows (strength 0.05: one denoising step at very low noise), so it is almost, but not exactly, a pure trip through the encoder and decoder.
The shrunk picture is blurry everywhere. The latent version is sharp, because the decoder learned how real pictures look and fills in the details the 4 numbers point to. But the tiniest details do not survive the trip: the smallest text came back garbled. This is one reason why early Stable Diffusion models were bad at small text and tiny faces. Newer models like Stable Diffusion 3 and FLUX keep 16 numbers per position instead of 4 to hold on to more detail.
So an image generator is really two models working together: an autoencoder that moves pictures in and out of a small, clever space, and a diffusion model that does all of its work inside that space. The method is called latent diffusion, and it is what the "Stable Diffusion" family is built on.
If this content helped you, you can buy me a coffee.
You can join the newsletter to be notified of awesome interactive articles and courses about AI, software and design. You will receive at most a few emails per month.
Join 800+ curious readers.