TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 16—18 · FACTUAL · DURABLE

Uncertainty has more than one meaning

An answer carries four kinds of confidence, and each is checked in a different place.

The idea

When people say a model is "confident", they may mean four different things. First, the probability the model gives each next token: a number inside the machine, about words. Second, calibration: whether a stated or internal probability matches how often the answer is actually right. Third, tone: how sure the prose sounds. Fourth, real uncertainty about the world: whether anyone could know the answer at all. These do not track each other. A sentence can sound sure while its tokens were a close call, and a token can be near-certain about a fact that is simply unknown.

Why it matters

"Confident is not correct" is true, and it is only the start. The useful question is which kind of confidence you are looking at, because each one is checked in a different place.

See it in the studio

You ask for the setback rule on a plot in Hubballi. The answer is fluent: that is tone. Asked, the model might put its own chance of being right at 70 per cent: that is a stated probability, which may or may not be calibrated. The bye-law exists and can be looked up, so there is no real uncertainty about the world here; your doubt is settled by opening the bye-law, not by reading the tone. Ask instead what the monsoon will do to that site in 2040, and the doubt is real. No tone and no probability removes it.

Watch for this

Reading a probability the tool shows you as a fact about the world. A token probability is about which word comes next. A "confidence score" in a product may be calibrated or may be decoration. Ask what it was measured against.

Try it

Ask a model twenty questions. Ten you can check: a code clause, a date, a span. Ten nobody can settle: what a client will want next year. For each answer, mark how sure the tone sounds (1–5) and ask the model to state its own chance of being right. Then check the ten checkable ones. Put the three columns side by side. They will not agree with each other, and that disagreement is the lesson.

Prove it

Take one AI answer and separate its four confidences: the probability inside the machine, the model's stated chance of being right, the tone of the sentence, and how uncertain the matter really is. Say how you would check each one.

How it works

The four kinds are measured in different places. Token probability comes from the model's output spread at each step; the sampling idea on this strand picks one word from it. Calibration is measured by asking a model for a probability many times and comparing it with the hit rate. Guo and colleagues showed in 2017 that modern neural networks are often over-confident in this measured sense. Kadavath and colleagues found in 2022 that large language models can be reasonably well calibrated when asked whether their own answer is true, and less so when asked to predict whether they know the answer to a new kind of task. Tone is learned from text and from reward tuning, and it carries no measurement at all. Real uncertainty about the world is what statisticians split into two kinds: noise that no amount of data removes, and ignorance that more data could fix. Only that fourth kind is about the world. The other three are about the machine and its writing.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON