TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Judgment · AGE 18—21 · FACTUAL · DURABLE

Is this tool any good?

A tool fails two ways, and the costlier failure is the one to test for.

The dilemma

A tool claims to flag structural problems in a plan, or to check a drawing against a byelaw, or to spot generated images. It is right most of the time. Is "most of the time" good enough? That depends on a question the advert never asks. Which way does it fail, and which failure hurts you?

The choices

Trust the accuracy number. Or ask two sharper questions before relying on it. When it says "problem" and there is none (a false alarm), what does that cost me? When it says "fine" and there is a problem (a miss), what does that cost me? Then run a small test built around the costlier one.

The consequence

Trust the number and you can end up with a tool that is 95 per cent right and wrong in the one way you cannot afford: a byelaw checker that misses the clause that fails your approval, or a fake-detector that flags every real photograph. Ask which error costs more and you know what to test, how much to trust it, and where you must still look yourself.

The case

A student wants to use a tool that flags "structurally implausible" elements in generated images before they go into a presentation. False alarm: a sound cantilever gets flagged, which costs a minute to check. Miss: an impossible span gets waved through, which costs the client's trust. The miss is the expensive error. So the test is not "how often is it right?" but "of ten images with a real structural fault, how many does it catch?" Six of ten. The tool is useful as a first pass and useless as a last one. The student now knows which.

Try it

For a task you actually do, say whether a false alarm or a miss would cost more, and why. Then describe a simple check, with no maths, that would show whether a tool is good enough for that job: ten cases you already know the answer to, and a count.

Take it to crit

Ask the student whether a false alarm or a miss would hurt more for this job, and how they would check. If they quote an accuracy number, ask what it is an accuracy at.

How it works

Testing a classifier means counting its two kinds of error separately. "How many of its alarms were real?" and "how many of the real problems did it find?" are different questions with different answers. A tool can score well on one while failing the other. A detector that flags everything finds every problem and is useless. Which question matters depends on the job. A check before a client meeting wants few misses. A check that interrupts your drawing wants few false alarms. Benchmarks for image models work the same way. They set controlled tasks (put the red cube left of the blue sphere; draw a plan that could be walked through) and count the failures by kind, which is why the benchmark cards elsewhere on this map report "at or near random" for spatial tasks rather than an overall score.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON