TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Judgment · AGE 18—20 · POSITIONAL · METHOD

Build your own test

Run five inputs you know cold through every new tool and every update.

When to use

Every time a new tool, model or version enters your workflow. Again after every update. Always before the deadline, never during it.

The method

Choose five inputs you know cold: a plan you drew, a site photograph you took, a specification you have checked, a question about a building you know well, and one deliberately awkward case. For each, write the expected answer before you run anything. Run all five. Score each: pass, partial, fail. Keep the five and the expected answers in one folder. Re-run the same five on every new version, and compare scores. Twenty minutes, once; five minutes, every time after.

Watch for this

Testing with the vendor's sample. Their sample was chosen to pass. Your five were chosen because you know the right answer.

The Lab's note

The five-input test is the Lab's recommended method for meeting a new tool, not the only one. Larger evaluations exist and the next cards describe them; this one is sized to fit a desk and a Monday morning, and to be re-run without ceremony every time the product changes under you.

Try it

Build the five today, from what is on your desk. Make two of them the same plan twice — once annotated in English, once in Kannada script (ಅಡುಗೆಮನೆ, ಮಲಗುವ ಕೋಣೆ, ಹಜಾರ) — and see whether the tool reads the rooms in both. Run all five on two tools that claim the same job. Write the two score sheets.

Prove it

Show the five inputs, the five expected answers, and the score sheets for two tools, and say which you would trust for what.

How it works

This is what researchers do at scale. A benchmark is your five, multiplied. Blueprint-Bench asks models to turn photographs of fifty apartments into floor plans and scores them against the real plans. AECV-Bench asks them to count doors and windows on a hundred and twenty plans and to answer questions grounded in the drawings. The principle is identical: known inputs, known answers, a score. Your private benchmark has one advantage over theirs. It is made of your work, so a pass means something for your work. Products change under you. The five stay. That is why the test, not the product name, is the thing worth knowing.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON