TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · The Brief · AGE 16—18 · FACTUAL · EVOLVING

What a reference image says

The model will copy the wrong thing well, because it reads a reference with no ranking.

The idea

A reference image is a dense brief. One picture carries light, material, proportion, mood and a dozen things you never noticed. But a model reads what the image shows; it cannot know why you pinned it unless you say. You chose that photograph for the way the stair turns. The model sees the stair, the railing, the floor tile, the window, the colour grade and the cat, all at once and with no ranking. Unless you say which quality you want carried over, it borrows whatever is loudest.

Why it matters

In most reference failures the model did not copy badly. It copied the wrong thing, and copied it well.

See it in the studio

You attach a photo of a Goan house for its deep verandah shadow. The output arrives with the verandah, and the laterite, the red-oxide floor, the Portuguese balusters and the tiled roof. You wanted one quality. You received the whole house. The brief never said which part was the point.

Watch for this

Two references that look alike to a model and mean opposite things to you. A monastery corridor and a hotel corridor can share proportion, light and material and carry opposite intent. A model can see proportion, and it can guess at intent when your words give it a lead. It cannot know an intent you never stated.

Try it

Take one reference you have actually used. Make two columns: what a model would read from it (everything visible), and why you picked it (usually one or two things). The gap between the columns is what your brief must say out loud.

Prove it

Show two references a model would read as similar, and explain, in words a model could act on, which quality each is there to contribute.

How it works

Modern multimodal systems put images and text in one shared space. The technique most people first met as CLIP, in 2021, showed how such a space is learnt by matching millions of photographs to their captions. Today's image-reasoning systems are built differently and do more with a reference than CLIP could, including inferring a plausible reason for it when your text gives them a lead. When you upload a reference, it lands in the window beside your words, and attention looks for what satisfies both. The image brings everything in it at once. Only your words can rank one quality above another. EVOLVING: image-reasoning models are getting better at taking instructions about a reference ("keep the proportion, not the material"), so how much you need to spell out is changing. Measured today; reviewed each edition.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON