TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Studio Practice · AGE 17—19 · POSITIONAL · METHOD
Choosing the smaller model is a design decision
Set model, count, resolution and re-rolls before you run, and write down what you spent.
When to use
Every time you set up an AI step in studio work: which model, how many generations, at what resolution, and how many re-rolls before you stop and read what you have.
The method
Four dials, set before you run. Model: the smallest model that does this job — a small model for a summary, a caption, a first sort; a large one only where the task has shown it needs one. Count: decide the batch before you generate, and read the whole contact sheet before asking for more. Resolution: draft low, and render once at the size the deliverable needs. Re-rolls: decide how many you will allow before you go back and fix the brief. A re-roll usually means the brief is loose; it is not a free retry. Write the four settings at the top of the log beside the brief. At the end of the task, write what you actually spent: calls, images, hours.
Watch for this
The "just one more" batch. If it takes fifty images to find one, you are not exploring; you have a loose brief, and you are paying for it in energy and time. Fix the brief, not the count. And read the model dial as a heuristic, not a carbon calculation: a small model that fails three times can cost more than a large one that succeeds once, hardware and provider efficiency move the figure more than size does, and providers publish almost nothing per task.
The Lab's note
Treating generation as a budget line is the Lab's method. Many studios treat it as free, because the invoice is small and the energy is spent somewhere else. We hold that an architect who prices a detail before drawing five hundred of them should price a batch before running fifty. Choosing the smaller model, when it does the job, is a design decision, because it is a decision about consequences.
Try it
Take a task you did this week with a large model. Do it again with the smallest model you can access. Judge the two outputs blind. Then count the generations you ran last week on one project and write the number where you can see it.
Prove it
Show one project log where the four dials were set before running and the actual spend recorded after — and say which dial you would move next time.
How it works
Every call runs in a data centre that draws electricity and, usually, cooling water; the footprint and not-free cards on this map hold the scale. The energy per call is argued over and depends on model size, hardware and load, which is why this card carries no figure and is reviewed each edition. What does not change is the direction: larger models cost more per call than smaller ones, more generations cost more than fewer, and higher resolution costs more than lower. The International Energy Agency's 2025 report on energy and AI is the reference for the scale at which this adds up across grids. The method here is the studio-sized version of the cost-and-latency budget card — the same discipline, with energy counted beside rupees.
What this idea builds on
What this idea opens up
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.