TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 18—20 · FACTUAL · DURABLE
Conditioning and control inputs
Conditioning separates what you have decided from what you are still exploring.
The idea
Generation does not have to run on text alone. You can condition it on structural inputs as well: an edge map, a depth map, a pose, a reference image, a line drawing of your plan. In ControlNet-style diffusion systems the control input is read at every denoising step and constrains the output towards its geometry. Adherence rises a great deal; it is not guaranteed, and it varies with control strength, model and prompt. The prompt stays free to decide material, light and mood. The walls mostly stay where the edge map put them when "brick" becomes "glass". Check.
Why it matters
Conditioning separates what you have decided from what you are still exploring. That separation is the core skill in using these tools for design rather than for pictures.
See it in the studio
You have a sketch elevation and want to see it in three materials. Edge-map control keeps the openings and proportions fixed across all three; only the prompt changes. Without control, you would get three different buildings.
Watch for this
Conditioning on a drawing that carries no decisions — a box with holes — and reading the result as a design. The control holds the box. It was still a box. And writing "preserve the geometry" in the prompt instead of supplying a control: a written instruction asks; the control input and its strength are what enforce — and even they enforce adherence, not a lock.
Try it
Extract an edge map from one of your own elevations; most tools do it in a click. Generate three materials with the same control and seed. Then run the same prompts with no control. Lay the six out and describe what the control gave you.
Prove it
Explain what a structural control input holds close and what it leaves free, and why a control image keeps the walls near their place when you change the prompt — and why "near", not "fixed", is the honest word.
How it works
ControlNet, the best-known method, attaches a trained copy of part of the diffusion network. That copy reads the control image and feeds its influence into each step, scaled by a conditioning strength you can set. The base model is untouched. Other methods bring reference images in as extra tokens, and non-diffusion generators condition in their own ways. All of them are conditioning — an input that shapes every step. Image-to-image is different: it only sets the starting point (TWO STARTS). PIXELS ARE NOT GEOMETRY on this strand explains why even a well-held edge map is still a picture of lines, not lines.
What this idea builds on
What this idea opens up
- Nothing yet names this as a foundation.
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.