How Text Steers an AI Image
In the previous article, a model turned noise into a picture, one guess at a time. But which picture? On its own, the model can only make something that looks like its training pictures in general. To get "a red fox in the snow", your prompt has to take part in every single guess.
First the prompt is turned into numbers. A separate model, the text encoder, turns the words into lists of numbers that capture their meaning, much like the embeddings from the earlier article. The image model reads these numbers at every step using attention: each part of the picture looks at the words and takes what it needs.
Let's go back to our tiny world, where a picture is a single dot. This time the training pictures come with captions: some dots draw a heart, some a ring and some a star. A model that knows the prompt only has to consider training pictures with a matching caption. Pick a prompt.
The tiny world from the previous article, now with captions: the training dots draw a heart, a ring and a star. With a prompt, the best guess only uses dots with that caption. It is calculated exactly; nothing is trained.
Without a prompt, the dots end up on all three shapes. With one, every guess only blends dots with that caption, so the very same noise turns into a heart, a ring or a star. That is all the text does: it changes the guess at every step.
Real models learn this from pictures together with their captions. During training the caption is sometimes left out on purpose, so the same model also learns to guess without any text. That turns out to be very useful.
The difference between the guess with the prompt and the guess without it points in the direction of the prompt. Image generators take that difference and multiply it by a number, the guidance scale. This trick is called classifier-free guidance. Here are four scales side by side.
The same 300 noise dots in every panel, 30 steps each. A dot counts as "on the shape" when the training dot closest to it has the prompt's caption; "different spots" counts how many of the shape's 60 training dots were reached.
At 0 the prompt is ignored, and the dots spread over all three shapes. At 1 the model simply follows its guess with the prompt, and in our exact toy world that is already enough. Above 1 the dots still land on the right shape, but they crowd onto fewer and fewer spots: the result follows the prompt more strictly and loses variety.
Real models learn their guesses imperfectly, so at a scale of 1 they often follow the prompt only loosely. Here are real pictures from Stable Diffusion XL: the same prompt and four different seeds, each with its own starting noise. Move the guidance scale.
seed 11
seed 22
seed 33
seed 44Real pictures made with Stable Diffusion XL (fast-sdxl on fal.ai), 25 steps. The slider only changes the guidance scale; each seed keeps its own starting noise.
It is the same pattern as with the dots. Too low, and the prompt is half ignored. Too high, and the pictures get harsh and start to look alike. That is why values around 5 to 8 are a common default in tools like Stable Diffusion.
The same trick also gives you negative prompts. Instead of starting from the guess without any text, the generator starts from the guess for what you do not want, and the guidance pushes every step away from it. Pick something you do not want in the bowl.

Real pictures made with Stable Diffusion XL (fast-sdxl on fal.ai): the same prompt, seed 9, guidance scale 7 and 25 steps. Only the negative prompt changes. The prompts are in English because the model was mostly trained on English captions.
Notice that more than one fruit changes. The negative prompt changes every guess along the way, so the whole picture shifts a little. In a real generator the prompt and the negative prompt work together: every guess is pulled toward one and pushed away from the other.
So the text never draws anything by itself. It changes the model's guess at every step, and the guidance scale sets how hard that change pushes. Keep in mind that the model does not understand your sentence the way a person does. It follows patterns it learned from captions and pictures, which is why image generators can still struggle with things like counting objects or putting them on the left or right.
If this content helped you, you can buy me a coffee.
You can join the newsletter to be notified of awesome interactive articles and courses about AI, software and design. You will receive at most a few emails per month.
Join 800+ curious readers.