TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 16—18 · FACTUAL · DURABLE

The mechanics of hallucination

Pretraining never separates a true continuation from a merely plausible one.

The idea

Two layers explain it. In pretraining, a model learns to continue text in the most plausible way, and nothing in that objective separates a true continuation from a merely plausible one. Post-training then rewards answers people prefer — truthful, sourced, hedged — and that reduces invention without removing it: its knowledge is incomplete, its sense of its own uncertainty poorly calibrated, and the tests models are scored on have mostly rewarded a confident guess over "I don't know". Whether an invention is a failure depends on the task. In ideation it may be the point. In a bye-law it is an error.

Why it matters

Once you know where invention comes from, you stop waiting for a fix that removes it. You decide whether this task can tolerate any, and supply what the task needs: the facts, and the check.

See it in the studio

You ask for the setback rule for a plot in your city. You get a clause number, a distance, and a confident tone. The clause may not exist. The model produced what a setback clause tends to look like. The real one is in the development control regulations, a document with a page number.

Watch for this

Asking the model whether it is sure. It answers that question the same way, with the most plausible reply. Certainty is generated text too.

Try it

Ask a chatbot for three papers on courtyard cooling in hot-dry climates, with authors and years. Search for each. Count how many exist exactly as given. Then ask again with one real paper pasted in, and see how the answer changes.

Prove it

Explain the two layers — why pretraining produces fluent false statements and why post-training reduces that without ending it — name one task where invention is welcome and one where it is an error, and name the two things you must add from outside to get checkable answers.

How it works

Pretraining rewards the model for predicting the next token well across its data; a confident, specific answer scores well on that target even when the specifics are invented, because its shape matches the data. Post-training (PRETRAINING / POST-TRAINING on this strand) rewards preferred behaviour and does shift models towards saying "I don't know" — but OpenAI's own analysis (Why language models hallucinate, September 2025) argues that most evaluations still score a guess above an admission of uncertainty, so developers are pushed to build models that guess. Many current products add repairs on top — retrieval of real documents, citations, tools that look things up. These move the facts from inside the model to inside your prompt, which is what the Brief strand calls "bring your own facts". The pretraining layer underneath does not change.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON