TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Foundations · AGE 15—17 · FACTUAL · EVOLVING
Trained on photos, not drawings
Image models learnt what architecture looks like, not how it is written down.
The idea
The big image models learned from image–text pairs scraped from the web, and in the open pools we can inspect, drawings are a small share next to photographs. Most makers do not publish their full diet, so this is an inference from the pools we can see and from how the tools behave. The behaviour is consistent: fluent in how buildings look, uneven in how buildings are drawn. They learnt what architecture looks like, not how it is written down.
Why it matters
If you expect a photo-trained model to read drawings, you will end up distrusting a tool that is good at something else. Use it where its diet was rich. Check it hard where its diet was thin.
See it in the studio
Ask for a moody courtyard photograph: strong. Ask for a dimensioned section of the same courtyard: confident lines, wrong conventions, invented dimensions. Same tool, opposite reliability, because one task is inside its diet and the other is outside.
Watch for this
A generated "plan" that has plan-flavour — walls, hatching, labels — but breaks the notation everywhere a plan carries information. It is a picture of a plan, not a plan.
Try it
Generate "a floor plan of a two-bedroom house." Red-pen it like a junior's drawing: doors that cannot swing, walls that carry nothing, stairs to nowhere. Count the violations.
Prove it
Explain why the same model is strong at building photographs and weak at building drawings, in terms of diet, not talent.
How it works
Recent benchmarks give this a shape. AECV-Bench (2026) tested multimodal models on real drawings: strong at reading the text on a drawing (up to 0.95 accuracy), moderate at spatial reasoning, weak at counting doors and windows from their symbols (often 0.40–0.55). Blueprint-Bench asked models to rebuild a flat's room layout from photographs; most scored at or below a random baseline, far under people. That is uneven drawing literacy, not blindness, and it is why this card is marked EVOLVING. Models trained on drawings may close the gap, and this card will be reviewed when they do.
What this idea builds on
What this idea opens up
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.