How AI Image Generation Actually Works

Text generation and image generation look like the same technology from the outside and they work almost nothing alike. A language model predicts the next chunk of text, one piece at a time, left to right. An image model does something stranger: it starts with a rectangle of pure random noise and gradually removes the noise until a picture is sitting there.

That description sounds like a metaphor. It is close to literally what happens.

Training by destroying images #

The insight behind diffusion models is backwards from what you would expect. Instead of teaching a model to draw, you teach it to undo damage.

Take a training image, a photograph of a cat. Add a small amount of random noise to it. Add a bit more. Keep going for hundreds of steps until the original image is completely gone and you have nothing but static. This process is called the forward diffusion process, and it is trivial, because adding noise requires no intelligence at all.

Now train a neural network on the reverse. Show it a noisy image and the step number, and ask it to predict the noise that was added. That is a well-posed supervised learning problem, because you know exactly what noise you added, so you have a perfect answer to check against.

Do that across billions of images at every noise level, and you end up with a model that, given anything noisy, can estimate what part of it is noise.

Generating by running it backwards #

Once you have a model that can spot noise, generation falls out of it.

Start with a rectangle of pure random static. Ask the model what part of that is noise. Subtract a fraction of its answer. The result is very slightly less noisy, which is to say very slightly more like a real image. Feed that back in and repeat.

After twenty to fifty steps, something coherent has emerged. The model was never told what to draw. It was told to make the thing in front of it look less like noise, over and over, and the only way to do that is to move it toward the kinds of images it saw during training.

The randomness in the starting static is what makes every generation different. Same prompt, different seed, different image. That is also why image tools let you fix a seed: reuse the same starting noise and you get the same picture back.

Where the text comes in #

So far nothing has said what the image should be of. Left alone, the process produces a plausible image of nothing in particular.

The steering comes from conditioning. Your text prompt is run through a text encoder, which turns it into a vector representation of its meaning, the same underlying idea as embeddings. That representation is injected into the denoising network at every step through cross-attention, so at each step the model is not answering “what is the noise here” but “what is the noise here, given that this is supposed to be a photograph of a cat wearing a hat.”

The link between text and images comes from training on hundreds of millions of image and caption pairs scraped from the web. The model never learned what a cat is. It learned which visual patterns tend to co-occur with the word “cat” in captions, which is a different thing that mostly produces the same result.

Classifier-free guidance is the knob that controls how hard the text pulls. Internally the model runs the denoising step twice, once with your prompt and once without, and then exaggerates the difference between them. A higher guidance scale pushes harder toward the prompt, producing images that follow the text more literally and often look oversaturated and stiff. A lower scale produces more natural, more varied images that wander further from what you asked for. That single number explains most of the difference between two tools that otherwise use the same architecture.

Latent diffusion, which is why it fits on a laptop #

Running this process on full-resolution pixels is enormously expensive. A 1024 by 1024 image is over a million pixels, and denoising all of them fifty times over is a lot of arithmetic.

The trick that made image generation practical on consumer hardware is doing the whole thing in a compressed space instead. A separate autoencoder learns to squash an image down to a much smaller representation, maybe 64 by 64 with more channels, and to expand it back out again with acceptable fidelity. The diffusion process runs entirely in that compressed latent space, and only at the very end does the decoder expand the result into actual pixels.

That is what the “latent” in latent diffusion means, and it cut the compute cost by more than an order of magnitude. It is the same category of move as quantization for language models: not a better model, a cheaper representation, and the thing that moved the technology from a data center onto hardware people own.

Why hands were hard #

The famous failure was hands, and the reason is instructive about what these models are actually doing.

The model has no concept of a hand as an object with a fixed structure. It has learned statistical associations between visual patterns. Hands appear in training images at wildly varying angles, in partial occlusion, holding things, overlapping each other, and captions almost never mention them. So the model learned “roughly this texture and shape appears in roughly this region” without ever learning “exactly five fingers, always.”

Text in images failed for the same reason for a long time. Letters are precise symbolic shapes where being approximately right is completely wrong, and a model optimizing for plausible local texture will happily produce something that has the visual character of writing without being writing.

Both improved substantially, through better training data, better captioning, higher resolution training, and architectural changes. Neither is fully solved, and the underlying reason is the same: these models are extremely good at texture and style and weaker at things that require exact discrete structure.

It is the visual equivalent of hallucination in a language model. The system generates something plausible, and plausible and correct are not the same requirement.

What the other knobs do #

Steps. How many denoising iterations. More steps means finer refinement up to a point, and past maybe 30 to 50 the returns are minimal and you are just spending time.

Sampler or scheduler. The specific algorithm for stepping from noise toward an image. Different samplers reach different results in different numbers of steps, and the differences are real but much smaller than people arguing about them online suggest.

Seed. The starting random noise. Fix it for reproducibility, change it for variation.

Negative prompt. A second conditioning signal describing what to push away from. It works by running the unconditioned pass with your negative prompt instead of with nothing, so guidance pushes away from it. This is why “blurry, low quality, extra fingers” in a negative prompt does something, though less than people credit it for.

img2img and inpainting. Instead of starting from pure noise, start from an existing image with a moderate amount of noise added. The model denoises from there, so the output keeps the structure of the original while changing its content. Inpainting is the same idea restricted to a masked region.

ControlNet and similar. Additional conditioning that constrains the output to follow a pose skeleton, a depth map, or an edge outline. This is what turned image generation from a slot machine into something you can direct.

Diffusion is not only for images #

The same machinery generalizes, which is a good sign that the idea is fundamental rather than a trick.

Video generation is diffusion with the time axis included, which is why it is so much more expensive and why temporal consistency, meaning a character’s shirt staying the same color between frames, is the hard part. Audio and music generation use diffusion over spectrograms, which are the same image-like representation of sound described in how speech recognition works. There is even work on diffusion for text, though autoregressive prediction still dominates there.

The honest caveats #

The training data question is real and unresolved. These models were trained on images scraped from the internet, including copyrighted work, mostly without permission or compensation. There is active litigation, and the legal position differs by jurisdiction. Anyone telling you the answer is obvious in either direction is telling you their opinion.

The models reproduce the biases of their training data, reliably and visibly. Ask for a generic professional in a given field and observe what you get. This is well documented and it is a property of the data rather than a bug in the code.

Memorization happens. Models can and occasionally do reproduce near-copies of training images, particularly for images that appeared many times in the dataset. It is rare and it is not zero.

And running these locally is entirely possible, since open-weight image models exist and run on a consumer graphics card. The tradeoffs are the same as for running language models locally: free per use, private, offline, and behind the frontier hosted models on quality.

The one-paragraph version #

An image model is trained to remove noise from noisy pictures. To generate, you hand it pure noise and have it remove noise repeatedly until an image appears, with your text prompt steering every step. The work happens in a compressed latent space to make it affordable. It is not drawing, and it is not retrieving. It is running a destruction process backwards, guided by what you asked for.

More on the models underneath in how LLMs work, and on what those systems do with your input in what happens to your data when you use AI.