TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Foundations · AGE 16—18 · FACTUAL · EVOLVING

Datasets of the built environment

A named dataset turns vague trust into checkable questions.

The idea

When a purpose-built model seems to know buildings, a specific dataset taught it: floor-plan collections, street-view corpora, scanned archives. Each has a name, a size, a geography, and gaps. A general model is different: its architecture knowledge came from a vast mixture of sources that its maker may not disclose. So "the AI understands architecture" means one of two things — a model trained on particular collections you can look up, or a model trained on a mixture you cannot — and you need to know which.

Why it matters

Named datasets turn vague trust into checkable questions: does the collection cover my building type, my region, my era? If not, the model's confidence about them is an act.

See it in the studio

A plan-generation tool trained on apartment layouts from one country will quietly impose that country's way of arranging rooms on yours. The dataset's habits become your defaults, unless you know the dataset exists.

Watch for this

Tools that will not say what they were trained on. Sometimes trade secret, sometimes embarrassment. Either way: test before trusting, because you cannot look up what is hidden.

Try it

Pick one AI tool that claims building knowledge. Spend ten minutes finding its dataset's name and coverage. Write down what you found, including if you found nothing.

Prove it

Name one real built-environment dataset, what it contains, and one gap a designer should know about.

How it works

Research datasets get papers describing their contents: CubiCasa5K, for example, is 5,000 floor plans annotated into over 80 object categories, mostly from Finland. Commercial training mixtures often stay opaque. Benchmarks such as AECV-Bench exist because claims about coverage needed independent testing; a benchmark is an exam, not the diet. Model documentation — cards, terms, stated limits — is where the trail starts, and a later idea shows you how to read them.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON