TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Ethics & Provenance · AGE 18—21 · FACTUAL · EVOLVING
Trained rarely, used constantly
Inference is a small cost per call and an enormous one in total.
The idea
A model carries two bills. Training is a large cost concentrated in model-building runs, weeks of thousands of chips, and the runs repeat: models are retrained, fine-tuned, distilled and replaced. Inference is what happens every time anyone uses the model: a small cost per call, and "small" varies a hundredfold between a text reply and a minute of video. The small cost runs billions of times a day, and current analysis, including the IEA's 2025 report, points to inference as where most ongoing energy use now builds up. Keep this at rough scale. Nobody has a clean number per prompt.
Why it matters
Your prompt is small and the total is huge, and both are true at once. Seeing how the two fit together lets you judge the cost instead of just feeling bad about it.
See it in the studio
"My one render is nothing" and "AI's footprint is enormous" said across the same table. Neither speaker is wrong. One is describing a single call; the other is describing the sum. The studio's share of the sum is its habits — how many calls, how loosely briefed.
Watch for this
Inventing a number to win the argument. Per-prompt figures circulate widely, differ by a hundredfold, and mostly lack a source. Quote the IEA or quote nothing.
Try it
Explain to a classmate, with a water bill as the analogy, how a connection charge paid each time the line is re-laid and a per-litre rate can both be large in different ways. Then map training and inference onto the two, and say why "per litre" is not one number when some calls are a sip and some are a tank.
Prove it
Tell the concentrated cost of training runs apart from the repeating cost of everyday use, in plain words, and explain how "my one prompt is tiny" and "inference is the bigger footprint" can both be true at once.
How it works
Training adjusts the weights in each run; inference reads them on every call. Training's cost sits in one place and the company that paid it can measure it. Inference's cost is spread across every data centre that serves the model, and from outside it can only be estimated, not measured. The balance between the two shifts with model size, efficiency and usage, and the IEA's projections will be revised. That is why the card is EVOLVING. The Two clocks card on the Foundations strand is the same split seen from the learning side.
What this idea builds on
What this idea opens up
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.