TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Judgment · AGE 17—19 · POSITIONAL · METHOD
A real test hunts failure
A test that only confirms what you hoped is a second demo.
The idea
A demo is chosen by the person selling the tool. It shows the inputs the tool handles best. A test is chosen by you, and it must include the cases the tool is likely to fail, not only the ones it is likely to pass. Feed it the awkward case, the torn scan, the brief that contradicts itself, the building it has probably never seen — alongside ordinary cases, so you learn where the edge is and not only that one exists. If you only confirm what you hoped, you have not tested. You have watched a second demo.
Why it matters
You will meet the worst case eventually. The only question is whether you meet it on purpose, in a test, or by accident, on a deadline.
See it in the studio
A plan-reading tool counts every door on the vendor's sample plan. Your survey is hand-drawn, photographed at an angle, with room names in Kannada. Run that one first. If it holds, the clean plans will hold too.
Watch for this
Testing with the same kind of input the demo used. Clean, typical, well-lit. That is not a test of your work. It is a test of their marketing.
The Lab's note
Hunting for failure is the Lab's recommended way to start testing a tool, not the whole of testing. A full evaluation also runs ordinary, representative cases and counts the errors by kind — MISS OR ALARM on this strand — and, where it matters, checks how confident the tool was when it was wrong. The awkward five find the edge. The ordinary cases tell you how often you will meet it.
Try it
Pick one tool you use. Write five inputs designed to break it: a site photograph at night, a plan with labels in an Indian language, a façade with no windows, a brief that asks for two things that cannot both be true, a question about a building that does not exist. Run all five. Write down what happened.
Prove it
Show the input that broke the tool, and say what that tells you about where the tool is safe to use.
How it works
Researchers formalised this in 2014 with adversarial examples: tiny, deliberate changes to an image that make a classifier confidently wrong, while the image looks unchanged to a person. The lesson spread. Modern systems are tested by teams whose whole job is to break them (the practice is called red-teaming), because the failures that matter are the ones nobody thought to look for. Your five awkward inputs are a small red team. The failure catalogue elsewhere on this map is where their findings go.
What this idea builds on
What this idea opens up
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.