TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 17—19 · FACTUAL · DURABLE

The machine that reads images

Reading images is a different job from making them, and often the more useful one.

The idea

Reading images is a different job from making them, and often the more useful one. A vision model can find objects, sort photographs, count windows, classify a building type from a street view, or flag where shade appears in each frame of a time-lapse. That is a large part of visual site documentation, at a scale no hand can match. It is not site analysis. Site analysis also needs time of day, access, where water collects, who uses an edge, regulation, and what a photograph leaves out. You choose what the tool looks for. You answer for what it misses.

Why it matters

Making images gets the attention. Reading them does the everyday work: surveys, audits, mapping. A student who can only generate has half the toolbox.

See it in the studio

Two hundred photographs from a site walk along a bazaar street. A vision tool sorts them by shopfront type, counts awnings, flags every hand-painted sign. It also labels a chai stall's tarpaulin as "tent" and misses the temple gopuram behind the cables entirely.

Watch for this

Trusting counts you did not spot-check. A vision model's error rate on your street is unknown until you measure it on your street.

Try it

Run fifty of your own site photos through any image-tagging tool. Check twenty tags by hand. Write down the error rate and the kind of thing it got wrong.

Prove it

Name two site-analysis jobs an image-reading tool could speed up, and say what such a tool is likely to miss on an unfamiliar Indian street.

How it works

Vision models learn from labelled images, or from image–text pairs in the CLIP family. NAME OR MAKE on Foundations sets out why a reader can be checked against truth while a maker cannot. Coverage follows the data. The well-photographed world reads well and the under-photographed world reads badly. DENSE, SPARSE applies here too.

Lineage

For decades, machine vision ran on features people designed by hand — edge detectors, corner finders, colour histograms. It worked only where the hand-written feature happened to fit. Convolutional networks changed that. The network learns its own features from labelled images, layer by layer, from edges up to parts and then objects. The 2012 ImageNet result, where a convolutional net beat hand-built methods by a wide margin, is usually taken as the moment the field turned. The current form is the vision-language model — CLIP and its descendants. It learns from image–text pairs rather than hand-assigned labels, so it can name things it was never explicitly taught. Each step widened what the machine could read, and each inherited the data problem from the step before. (lineage as of Aug 2026)

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON