TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 18—20 · FACTUAL · EVOLVING
Can the model read a plan?
Models read the labels on a plan well and do worst at its symbols.
The idea
Reading a floor plan means holding three things at once: the geometry of the rooms, the connections between them, and the symbols — door swings, window breaks, stair arrows, the drawing's own notation. Current multimodal models have been tested on exactly this. On AECV-Bench (2026) they read the text on a drawing well, reason about space unevenly, and do worst at the symbols — counting doors, windows and rooms — with wide differences between models and tasks. "It saw my plan" is true. "It understood my plan" is, on that evidence and today, mostly not true.
Why it matters
A model that reads your title block and room names will sound as if it has read your plan. Its fluent summary can hide that it never counted the doors.
See it in the studio
You upload a hostel floor plan and ask for a fire-exit check. The model names the rooms from their labels, then reasons about travel distance from rooms it has not located, past doors it has not counted. The reply reads like a checker's report. It is a guess built from the labels.
Watch for this
A plan summary that mentions every labelled room and no unlabelled one. That is the sign that it read the words, not the drawing.
Try it
Give a model one of your plans and test three things in turn: read the labels; judge two adjacencies; count the doors. Score each against the drawing. If your sheet carries Kannada labels, run the label test twice — once on the Kannada sheet, once on an English copy — and score the difference. Keep the score sheet, dated, for the next model.
Prove it
Show one thing a model reads on a plan and one thing it cannot, from your own test — and explain the gap between reading text on a drawing and understanding the drawing.
How it works
AECV-Bench tests how well current models read built-environment drawings: it finds them strong on document-style questions and weak on counting architectural objects, with door and window counts the worst, and the gap varying by model. That fits PHOTOS, NOT PLANS on Foundations: the training data was photographs, not drawing notation. PIXELS ARE NOT GEOMETRY on this strand gives the deeper reason — the model is reading a raster, not lines. EVOLVING because the benchmark exists to be re-run and drawing-trained models are an active line of work. This card carries its review date for that reason.
What this idea builds on
What this idea opens up
- Nothing yet names this as a foundation.
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.