Controlling the Result
Prompting alone tops out fast. Everything that separates a lucky generation from a deliverable happens after the first image exists — showing the model references, fixing regions instead of rerolling, editing by conversation, and finishing at proper resolution.
The mental shift that matters: stop treating generation as a slot machine and start treating it as a first draft. Professionals generate once to establish the composition, then spend the remaining time steering that specific image. Every technique below is a way of steering rather than rerolling.
The most underused capability in every modern tool. A picture communicates a look in a way no sentence can, and essentially every 2026 platform will take one.
The three things a reference can carry, usually as separate controls:
- Style reference — take the palette, light, texture and mark-making from this image, but not its content. This is how you get a consistent look across a whole set, and how you apply your own style rather than borrowing someone's name.
- Character or subject reference — keep this face, this person, this product across many different images.
- Composition or structure reference — keep this layout, this pose, this perspective, and change everything else.
Most tools let you weight how strongly the reference applies. Low weight is a hint; high weight is nearly a copy. Start low and increase — the failure mode of a strong reference is a near-duplicate of the input, which is both useless and, if the reference isn't yours, a problem.
Use your own images. Your photographs, your sketches, your palettes, your previous work. This is the single clearest route to output that looks like you made it, and it strengthens your position on authorship considerably.
The most common real-world requirement — a children's book, a comic, a storyboard, a brand mascot, a product across a campaign — and historically the hardest thing to do. It got dramatically easier in 2026.
In rough order of reliability:
- Dedicated character features. Most major platforms now have an explicit character-reference control: upload two or three images of your subject and reuse them. Use this first if it exists.
- Conversational editing. Generate the character once, then ask for new scenes with "the same person, now sitting on the steps." Chat-based tools hold onto the subject well within a session.
- A very specific written description, reused verbatim. Not "a girl" but "a nine-year-old girl with short curly auburn hair, round tortoiseshell glasses, freckles, a mustard-yellow raincoat." Paste the identical block into every prompt.
- A locked seed plus small prompt changes. Cheap, works only for modest variations, but useful for a set of near-identical shots.
- A character sheet. Generate one image containing the same character from several angles and expressions, then use that single image as your reference for everything afterwards. This is the trick that professional illustrators land on.
Set expectations with clients. Perfect consistency across dozens of images still is not solved. Plan for a retouching pass, and price for it.
The biggest usability change of the last two years, and the one that most closes the gap between AI tools and normal creative software. Instead of writing a new prompt and hoping, you keep the image you have and describe the change.
"Make the jacket dark green." "Remove the car on the left." "Same shot, later in the evening." "Pull back so we see the whole building." "Keep everything, just open her eyes."
The model edits the image rather than regenerating it, so everything you didn't mention mostly survives. Available in GPT Image, Gemini / Nano Banana, and increasingly everywhere else.
How to get the most from it:
- One change per turn. Bundled requests get partially applied and you won't know which part failed.
- Say what must not change: "keeping the lighting and her expression exactly as they are."
- Save versions as you go. Long edit chains drift — colours shift, faces soften — and you will want to go back three steps.
- Know when to stop. After many turns, quality degrades. Export, fix the rest in your normal editor.
Painting over a selected region, and extending the canvas past its edges. Between them they solve most of the problems you'd otherwise reroll for — and they are available in every serious tool, including Photoshop.
Inpainting — mask an area, describe what should be there, regenerate only that area.
- Fixing a bad hand, a wrong expression, a duplicated object
- Removing something you don't want, or adding something you do
- Changing one element — a garment, a sign, a product — while everything else stays identical
- Practical tip: feather your mask generously and include some surrounding context in the selection. Tight masks produce visible seams.
Outpainting — extend the image beyond its original frame.
- Turning a square into a banner, or a portrait crop into a landscape
- Recovering headroom the original composition didn't leave you
- Building a wide environment from a single strong central image
- Practical tip: extend in modest steps rather than one huge jump, and re-anchor with a prompt each time. Big extensions drift into a different scene.
Feed in an existing image and have the model produce a new version of it, keeping the composition and structure. The one control that matters is *strength* — how far it is allowed to depart from what you gave it.
- Low strength — a light pass. Colour and texture shift, composition intact. Good for grading, unifying a set, or adding grain and material.
- Medium strength — the useful middle. Your sketch keeps its layout but becomes a rendered piece. This is where most of the value lives.
- High strength — barely related to the input. At this point you may as well have prompted from scratch.
What artists actually use it for:
- Sketch to finish. Rough out the composition yourself — badly is fine, it only needs to read — and let the model render it. You keep authorship of the composition, which is both the creatively interesting part and the legally significant one.
- Style unification across an inconsistent set of images.
- Photo transformation — turning your own reference photography into illustration, and staying entirely within material you own.
A step beyond image-to-image: instead of loosely following a reference, the model is held to a specific structural property of it — the pose of a figure, the depth of a scene, the lines of a drawing — while everything else is free.
You will meet these as named options in the interface rather than as anything you have to configure:
- Pose — extract a skeleton from a reference photo and generate a completely different character in exactly that pose. Solves the hardest part of figure work.
- Depth — preserve the spatial layout of a scene and restyle everything in it. Useful for keeping an architectural or environmental composition while changing its world.
- Edge or line art — follow the outlines of a drawing exactly. This is the one that matters most to illustrators: draw your own linework, let the model handle colour and rendering, and the drawing remains unambiguously yours.
- Scribble — the loose version of the same idea, for rough thumbnails.
Why this is the most important section here for working artists. Guided generation is what turns these tools from a slot machine into an instrument. You supply the drawing, the pose, the composition — the parts that require judgement and skill — and the model does the rendering labour. That is a defensible creative process, a repeatable one, and the one that survives contact with a client who asks how the image was made.
The most complete implementations are in the open-source tooling covered in the Deep Learning track, but usable versions now ship in OpenArt, Krea, Leonardo, Freepik and Photoshop — no installation required.
Models generate at a fixed native size — often around one to two megapixels. That is fine on screen and nowhere near enough for print, large format, or anything a client will inspect closely. Upscaling is the last step of essentially every professional workflow.
- Standard upscale — two to four times the resolution, inventing plausible detail as it goes. Built into most platforms.
- Creative upscale — adds detail that wasn't there. Spectacular results, but it will change faces and small features, so check identity afterwards.
- Face restoration — a specialised pass for portraits, and the fix for the mushy-distant-face problem from earlier.
- Do it in stages. Two 2× passes usually beat one 4× pass.
And then finish properly. Colour grade, crop, retouch, composite, add grain, set type. Almost nothing professional ships straight out of a generator. The final ten percent — the part that makes it look authored rather than generated — still happens in Photoshop, Affinity, Lightroom, Procreate or After Effects. Your existing craft is not obsolete here; it is the differentiator.