TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Judgment · AGE 18—20 · FACTUAL · DURABLE

The score is not your work

Marketing quotes scores, and your liability is your work.

The idea

A benchmark score says how a model did on a fixed, shared test: a set task, a set of inputs, and a scoring rule everyone uses — sometimes known answers, sometimes human raters, sometimes a hidden test set. That is useful. It is also narrow. Your studio has a particular site, drawings in a particular convention, a client who changes their mind, and a submission date. A high score proves the tool can do the test. Whether it can do your work is a separate question, and the only way to answer it is the five-input test on this map.

Why it matters

Marketing quotes scores. Your liability is your work. The two are not measured in the same place.

See it in the studio

A model scores well on a document-reading benchmark built from clean PDFs. Your byelaw is a 1998 photocopy, scanned crooked, with a stamp across clause 7. The score did not include clause 7.

Watch for this

The phrase "state of the art". It means "top of a particular table on a particular day". Ask which table, and what was on it.

Try it

Find the benchmark a tool you use quotes. Read what the test actually contains: the inputs, the answers, who judged. List three ways your work differs from it.

Prove it

Explain what a benchmark can tell you and what it cannot, with one concrete example of each.

How it works

Benchmarks are built from test sets with known answers, and the known answers are the limit. Whatever the set did not contain was not measured. The two benchmarks closest to your work show their scale honestly: fifty apartments in one, a hundred and twenty floor plans in the other. They also show how a serious test is made, which is why they appear on the previous cards. Two further cautions are standard in the field. A model may have seen data resembling the test during training, which inflates the score. And a score is a point in time. The next version moves it. Neither caution makes benchmarks useless. Both make them a starting point for your own test, not a substitute for it.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON