TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Judgment · AGE 17—19 · POSITIONAL · METHOD

A real test hunts failure

A test that only confirms what you hoped is a second demo.

The idea

A demo is chosen by the person selling the tool. It shows the inputs the tool handles best. A test is chosen by you, and it must include the cases the tool is likely to fail, not only the ones it is likely to pass. Feed it the awkward case, the torn scan, the brief that contradicts itself, the building it has probably never seen — alongside ordinary cases, so you learn where the edge is and not only that one exists. If you only confirm what you hoped, you have not tested. You have watched a second demo.

Why it matters

You will meet the worst case eventually. The only question is whether you meet it on purpose, in a test, or by accident, on a deadline.

See it in the studio

A plan-reading tool counts every door on the vendor's sample plan. Your survey is hand-drawn, photographed at an angle, with room names in Kannada. Run that one first. If it holds, the clean plans will hold too.

Watch for this

Testing with the same kind of input the demo used. Clean, typical, well-lit. That is not a test of your work. It is a test of their marketing.

The Lab's note

Hunting for failure is the Lab's recommended way to start testing a tool, not the whole of testing. A full evaluation also runs ordinary, representative cases and counts the errors by kind — MISS OR ALARM on this strand — and, where it matters, checks how confident the tool was when it was wrong. The awkward five find the edge. The ordinary cases tell you how often you will meet it.

Try it

Pick one tool you use. Write five inputs designed to break it: a site photograph at night, a plan with labels in an Indian language, a façade with no windows, a brief that asks for two things that cannot both be true, a question about a building that does not exist. Run all five. Write down what happened.

Prove it

Show the input that broke the tool, and say what that tells you about where the tool is safe to use.

How it works

Researchers formalised this in 2014 with adversarial examples: tiny, deliberate changes to an image that make a classifier confidently wrong, while the image looks unchanged to a person. The lesson spread. Modern systems are tested by teams whose whole job is to break them (the practice is called red-teaming), because the failures that matter are the ones nobody thought to look for. Your five awkward inputs are a small red team. The failure catalogue elsewhere on this map is where their findings go.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON