/Step 3 — The Picture Is Refined Out of Static

Step 3 — The Picture Is Refined Out of Static

Here is the part that surprises people. The model does not draw your image. It does not start with a blank canvas and add a horizon, then a boat, then a sky. It starts with a rectangle of pure random static — visual noise, like an untuned television — and repeatedly *removes* the parts that don't belong.

The sculpture analogy is the accurate one. The model is not painting; it is carving. It looks at the noise and asks "if this static were a slightly blurry version of the image described in the brief, what would I need to take away?" It removes a bit. It looks again. It removes a bit more.

Do that twenty to fifty times, checking against your brief every time, and a coherent image emerges from the static. Early steps decide the big things — composition, masses, where the light comes from. Late steps decide texture, edges, fine detail.

Why this is worth knowing:

  • The composition is locked in early. By roughly a quarter of the way through, the layout is decided. This is why prompt words about composition and framing carry more weight than words about small details — and why adding "tiny brass buttons" rarely gets you tiny brass buttons.
  • More steps is not more quality. Past a point the image stops changing meaningfully and you're just spending money. Defaults are almost always fine; doubling the steps is not a fix for a weak prompt.
  • The starting static is a real ingredient. That random rectangle is generated from a number called the seed. Same prompt + same seed + same settings = the same image, every time. Same prompt + different seed = a genuinely different image. When you get something almost right, locking the seed and changing one word is how you make a controlled adjustment instead of rolling the dice again.