TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Judgment · AGE 18—20 · POSITIONAL · METHOD
Build your own test
Run five inputs you know cold through every new tool and every update.
When to use
Every time a new tool, model or version enters your workflow. Again after every update. Always before the deadline, never during it.
The method
Choose five inputs you know cold: a plan you drew, a site photograph you took, a specification you have checked, a question about a building you know well, and one deliberately awkward case. For each, write the expected answer before you run anything. Run all five. Score each: pass, partial, fail. Keep the five and the expected answers in one folder. Re-run the same five on every new version, and compare scores. Twenty minutes, once; five minutes, every time after.
Watch for this
Testing with the vendor's sample. Their sample was chosen to pass. Your five were chosen because you know the right answer.
The Lab's note
The five-input test is the Lab's recommended method for meeting a new tool, not the only one. Larger evaluations exist and the next cards describe them; this one is sized to fit a desk and a Monday morning, and to be re-run without ceremony every time the product changes under you.
Try it
Build the five today, from what is on your desk. Make two of them the same plan twice — once annotated in English, once in Kannada script (ಅಡುಗೆಮನೆ, ಮಲಗುವ ಕೋಣೆ, ಹಜಾರ) — and see whether the tool reads the rooms in both. Run all five on two tools that claim the same job. Write the two score sheets.
Prove it
Show the five inputs, the five expected answers, and the score sheets for two tools, and say which you would trust for what.
How it works
This is what researchers do at scale. A benchmark is your five, multiplied. Blueprint-Bench asks models to turn photographs of fifty apartments into floor plans and scores them against the real plans. AECV-Bench asks them to count doors and windows on a hundred and twenty plans and to answer questions grounded in the drawings. The principle is identical: known inputs, known answers, a score. Your private benchmark has one advantage over theirs. It is made of your work, so a pass means something for your work. Products change under you. The five stay. That is why the test, not the product name, is the thing worth knowing.
What this idea builds on
What this idea opens up
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.