TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 19—21 · FACTUAL · DURABLE
Evaluating a model for a task
Twenty of your own survey reports tell you more about a model than any leaderboard.
The dilemma
Two models, one decision. One tops the leaderboards this month. The other is cheaper, runs locally, and nobody is talking about it. The studio needs one for reading survey reports and drafting design statements. How do you choose?
The choices
Pick the leader, because everyone does. Pick the cheap one, because cost is real. Or build a small test set from your own work — twenty survey reports, ten briefs, the kind of mistake that actually hurts you — run both models on it, and read where each holds and where each breaks.
The consequence
Pick by leaderboard and you inherit a ranking built on tasks that are not yours, possibly gamed, possibly out of date. Pick by cost and you may save money on a tool that invents survey figures. Build the test and you spend a day. In return you have a reusable instrument that tells you, for every future model, whether it fits your work.
The case
A practice adopts the top-ranked model for reading geotechnical reports. It is excellent at summarising and, twice in a month, confidently misreads a bearing-capacity table. The unfashionable model, tested later on the same twenty reports, misreads none — and is worse at prose. The leaderboard had measured prose.
Take it to crit
When you picked a model, can you show evidence from your own task behind the choice — or only point to what is popular? Bring the test set.
How it works
Public benchmarks measure general tasks under conditions you do not control, and the rankings themselves have been shown to be distorted by how labs test and what they disclose. A studio test set is small, specific and honest: your inputs, your error costs, your judges. THE PIPELINE on Foundations names the trap. A model can pass its own exam and fail in your studio.
What this idea builds on
What this idea opens up
Sources
- DES.A
- understanding-ai
- ch08
- Singh et al., The Leaderboard Illusion
- Google, Machine Learning Crash Course
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.