TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Foundations · AGE 15—17 · FACTUAL · DURABLE
What a score measures
A score tells you how a model did on that paper, on that day.
The idea
A benchmark is a fixed exam: a set of questions, a marking scheme, and a number at the end. Benchmarks make progress visible and comparable, and the field runs on them. But an exam only measures what its setters thought to ask. A model can be trained on material close to the exam, tuned until its score rises, and still fail on work the exam never contained. A score tells you how a model did on that paper, on that day. It does not tell you how it will do on your brief.
Why it matters
Scores travel in headlines and sales decks. Knowing what a score is, and what it is not, stops a leaderboard from making a studio decision for you.
See it in the studio
A model "scores 90 per cent on a reasoning benchmark" and is sold to your office on that basis. Then it reads a floor plan and mistakes a staircase for a corridor. The exam had no floor plans in it. Researchers who built a plan-reconstruction benchmark found leading models far below people, several at or below a random baseline; the exam that asked the studio's question returned a very different number.
Watch for this
Goodhart's trap: once a score becomes the target, it can stop measuring what it was meant to measure. And watch for the benchmark you never see: the one that went unpublished because the model did badly on it.
Try it
Take any published benchmark score for a model you use. Find the benchmark's own description: what tasks, how many, who wrote them. Then write three questions from this week's studio work and check whether anything like them is in the exam.
Prove it
Explain what a benchmark score does measure, what it cannot, and how you would decide whether a given score applies to your work.
How it works
Benchmarks are datasets with answers. They can leak into training data, become saturated, or reward the exam's style rather than the underlying skill; the idea train · validate · test, beside this one, shows how a score can lie before the exam is even sat. Raji and colleagues have argued that treating a handful of benchmarks as measures of "general" capability has no sound basis. Well-made task-specific benchmarks are still the best public evidence we have. Blueprint-Bench and AECV-Bench are two built for the built environment, and they are why several cards on this map can say "measured" instead of "probably". The Judgment strand takes the next step: your own test, on your own work.
What this idea builds on
What this idea opens up
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.