TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Ethics & Provenance · AGE 18—21 · FACTUAL · EVOLVING

Trained rarely, used constantly

Inference is a small cost per call and an enormous one in total.

The idea

A model carries two bills. Training is a large cost concentrated in model-building runs, weeks of thousands of chips, and the runs repeat: models are retrained, fine-tuned, distilled and replaced. Inference is what happens every time anyone uses the model: a small cost per call, and "small" varies a hundredfold between a text reply and a minute of video. The small cost runs billions of times a day, and current analysis, including the IEA's 2025 report, points to inference as where most ongoing energy use now builds up. Keep this at rough scale. Nobody has a clean number per prompt.

Why it matters

Your prompt is small and the total is huge, and both are true at once. Seeing how the two fit together lets you judge the cost instead of just feeling bad about it.

See it in the studio

"My one render is nothing" and "AI's footprint is enormous" said across the same table. Neither speaker is wrong. One is describing a single call; the other is describing the sum. The studio's share of the sum is its habits — how many calls, how loosely briefed.

Watch for this

Inventing a number to win the argument. Per-prompt figures circulate widely, differ by a hundredfold, and mostly lack a source. Quote the IEA or quote nothing.

Try it

Explain to a classmate, with a water bill as the analogy, how a connection charge paid each time the line is re-laid and a per-litre rate can both be large in different ways. Then map training and inference onto the two, and say why "per litre" is not one number when some calls are a sip and some are a tank.

Prove it

Tell the concentrated cost of training runs apart from the repeating cost of everyday use, in plain words, and explain how "my one prompt is tiny" and "inference is the bigger footprint" can both be true at once.

How it works

Training adjusts the weights in each run; inference reads them on every call. Training's cost sits in one place and the company that paid it can measure it. Inference's cost is spread across every data centre that serves the model, and from outside it can only be estimated, not measured. The balance between the two shifts with model size, efficiency and usage, and the IEA's projections will be revised. That is why the card is EVOLVING. The Two clocks card on the Foundations strand is the same split seen from the learning side.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON