TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 17—19 · FACTUAL · DURABLE
The machine that reads images
Reading images is a different job from making them, and often the more useful one.
The idea
Reading images is a different job from making them, and often the more useful one. A vision model can find objects, sort photographs, count windows, classify a building type from a street view, or flag where shade appears in each frame of a time-lapse. That is a large part of visual site documentation, at a scale no hand can match. It is not site analysis. Site analysis also needs time of day, access, where water collects, who uses an edge, regulation, and what a photograph leaves out. You choose what the tool looks for. You answer for what it misses.
Why it matters
Making images gets the attention. Reading them does the everyday work: surveys, audits, mapping. A student who can only generate has half the toolbox.
See it in the studio
Two hundred photographs from a site walk along a bazaar street. A vision tool sorts them by shopfront type, counts awnings, flags every hand-painted sign. It also labels a chai stall's tarpaulin as "tent" and misses the temple gopuram behind the cables entirely.
Watch for this
Trusting counts you did not spot-check. A vision model's error rate on your street is unknown until you measure it on your street.
Try it
Run fifty of your own site photos through any image-tagging tool. Check twenty tags by hand. Write down the error rate and the kind of thing it got wrong.
Prove it
Name two site-analysis jobs an image-reading tool could speed up, and say what such a tool is likely to miss on an unfamiliar Indian street.
How it works
Vision models learn from labelled images, or from image–text pairs in the CLIP family. NAME OR MAKE on Foundations sets out why a reader can be checked against truth while a maker cannot. Coverage follows the data. The well-photographed world reads well and the under-photographed world reads badly. DENSE, SPARSE applies here too.
Lineage
For decades, machine vision ran on features people designed by hand — edge detectors, corner finders, colour histograms. It worked only where the hand-written feature happened to fit. Convolutional networks changed that. The network learns its own features from labelled images, layer by layer, from edges up to parts and then objects. The 2012 ImageNet result, where a convolutional net beat hand-built methods by a wide margin, is usually taken as the moment the field turned. The current form is the vision-language model — CLIP and its descendants. It learns from image–text pairs rather than hand-assigned labels, so it can name things it was never explicitly taught. Each step widened what the machine could read, and each inherited the data problem from the step before. (lineage as of Aug 2026)
What this idea builds on
What this idea opens up
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.