TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 19—21 · FACTUAL · DURABLE

Evaluating a model for a task

Twenty of your own survey reports tell you more about a model than any leaderboard.

The dilemma

Two models, one decision. One tops the leaderboards this month. The other is cheaper, runs locally, and nobody is talking about it. The studio needs one for reading survey reports and drafting design statements. How do you choose?

The choices

Pick the leader, because everyone does. Pick the cheap one, because cost is real. Or build a small test set from your own work — twenty survey reports, ten briefs, the kind of mistake that actually hurts you — run both models on it, and read where each holds and where each breaks.

The consequence

Pick by leaderboard and you inherit a ranking built on tasks that are not yours, possibly gamed, possibly out of date. Pick by cost and you may save money on a tool that invents survey figures. Build the test and you spend a day. In return you have a reusable instrument that tells you, for every future model, whether it fits your work.

The case

A practice adopts the top-ranked model for reading geotechnical reports. It is excellent at summarising and, twice in a month, confidently misreads a bearing-capacity table. The unfashionable model, tested later on the same twenty reports, misreads none — and is worse at prose. The leaderboard had measured prose.

Take it to crit

When you picked a model, can you show evidence from your own task behind the choice — or only point to what is popular? Bring the test set.

How it works

Public benchmarks measure general tasks under conditions you do not control, and the rankings themselves have been shown to be distorted by how labs test and what they disclose. A studio test set is small, specific and honest: your inputs, your error costs, your judges. THE PIPELINE on Foundations names the trap. A model can pass its own exam and fail in your studio.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON