TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · The Brief · AGE 16—18 · FACTUAL · EVOLVING
What a reference image says
The model will copy the wrong thing well, because it reads a reference with no ranking.
The idea
A reference image is a dense brief. One picture carries light, material, proportion, mood and a dozen things you never noticed. But a model reads what the image shows; it cannot know why you pinned it unless you say. You chose that photograph for the way the stair turns. The model sees the stair, the railing, the floor tile, the window, the colour grade and the cat, all at once and with no ranking. Unless you say which quality you want carried over, it borrows whatever is loudest.
Why it matters
In most reference failures the model did not copy badly. It copied the wrong thing, and copied it well.
See it in the studio
You attach a photo of a Goan house for its deep verandah shadow. The output arrives with the verandah, and the laterite, the red-oxide floor, the Portuguese balusters and the tiled roof. You wanted one quality. You received the whole house. The brief never said which part was the point.
Watch for this
Two references that look alike to a model and mean opposite things to you. A monastery corridor and a hotel corridor can share proportion, light and material and carry opposite intent. A model can see proportion, and it can guess at intent when your words give it a lead. It cannot know an intent you never stated.
Try it
Take one reference you have actually used. Make two columns: what a model would read from it (everything visible), and why you picked it (usually one or two things). The gap between the columns is what your brief must say out loud.
Prove it
Show two references a model would read as similar, and explain, in words a model could act on, which quality each is there to contribute.
How it works
Modern multimodal systems put images and text in one shared space. The technique most people first met as CLIP, in 2021, showed how such a space is learnt by matching millions of photographs to their captions. Today's image-reasoning systems are built differently and do more with a reference than CLIP could, including inferring a plausible reason for it when your text gives them a lead. When you upload a reference, it lands in the window beside your words, and attention looks for what satisfies both. The image brings everything in it at once. Only your words can rank one quality above another. EVOLVING: image-reasoning models are getting better at taking instructions about a reference ("keep the proportion, not the material"), so how much you need to spell out is changing. Measured today; reviewed each edition.
What this idea builds on
What this idea opens up
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.