TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Foundations · AGE 15—17 · FACTUAL · DURABLE

What a score measures

A score tells you how a model did on that paper, on that day.

The idea

A benchmark is a fixed exam: a set of questions, a marking scheme, and a number at the end. Benchmarks make progress visible and comparable, and the field runs on them. But an exam only measures what its setters thought to ask. A model can be trained on material close to the exam, tuned until its score rises, and still fail on work the exam never contained. A score tells you how a model did on that paper, on that day. It does not tell you how it will do on your brief.

Why it matters

Scores travel in headlines and sales decks. Knowing what a score is, and what it is not, stops a leaderboard from making a studio decision for you.

See it in the studio

A model "scores 90 per cent on a reasoning benchmark" and is sold to your office on that basis. Then it reads a floor plan and mistakes a staircase for a corridor. The exam had no floor plans in it. Researchers who built a plan-reconstruction benchmark found leading models far below people, several at or below a random baseline; the exam that asked the studio's question returned a very different number.

Watch for this

Goodhart's trap: once a score becomes the target, it can stop measuring what it was meant to measure. And watch for the benchmark you never see: the one that went unpublished because the model did badly on it.

Try it

Take any published benchmark score for a model you use. Find the benchmark's own description: what tasks, how many, who wrote them. Then write three questions from this week's studio work and check whether anything like them is in the exam.

Prove it

Explain what a benchmark score does measure, what it cannot, and how you would decide whether a given score applies to your work.

How it works

Benchmarks are datasets with answers. They can leak into training data, become saturated, or reward the exam's style rather than the underlying skill; the idea train · validate · test, beside this one, shows how a score can lie before the exam is even sat. Raji and colleagues have argued that treating a handful of benchmarks as measures of "general" capability has no sound basis. Well-made task-specific benchmarks are still the best public evidence we have. Blueprint-Bench and AECV-Bench are two built for the built environment, and they are why several cards on this map can say "measured" instead of "probably". The Judgment strand takes the next step: your own test, on your own work.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON