TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 17—19 · FACTUAL · DURABLE

Text-to-image vs image-to-image

Choosing between the two modes is really choosing how much of your source to keep.

The idea

In a diffusion or flow pipeline, text-to-image starts the denoising loop from pure noise, so layout, massing and viewpoint are invented fresh every run. Image-to-image starts from a noised-up version of a picture you supply. The less noise you add, the more of your picture's structure survives. The more noise, the freer the model gets. The setting that controls this is usually called denoising strength. Choosing between the two modes is really choosing how much of your source to keep. Keep the massing and change the mood, or keep almost nothing and take a fresh look.

Why it matters

Most "it ignored my sketch" complaints come from a strength setting that is too high. Most "it just traced my sketch" complaints come from the same setting, too low.

See it in the studio

You have a massing screenshot and want it at dusk, in brick, with people. Text-to-image gives you a different building. Image-to-image at moderate strength keeps your masses and repaints the rest.

Watch for this

Using image-to-image on a weak drawing and expecting it to be redesigned. It keeps what you gave it. A wrong plan, lit beautifully, is still a wrong plan.

Try it

Take one sketch. Run image-to-image at low, middle and high strength. Lay the three beside the sketch. Mark, on each, what survived and what was invented.

Prove it

Explain why image-to-image keeps the layout that text-to-image makes up fresh each time, and describe what you give up as you turn denoising strength up.

How it works

Image-to-image adds noise to your picture part of the way along the same path training used, then denoises from there. The idea was set out in a method called SDEdit. It is the simplest form of grounding. CONTROL INPUT on this strand is the stronger form, where edges or depth are fed alongside the prompt at every step rather than only as a starting point. Other generator families — autoregressive image models, for one — have no denoising loop and do "start from my picture" differently; this card describes the diffusion and flow family.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON