TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 18—21 · POSITIONAL · HELD
Models have temperaments
Every model listens differently, so brief for its temperament.
The dilemma
You have a brief that worked beautifully in one model. You paste it into another and get a flatter, more literal, or oddly cautious result. Do you conclude the second model is worse, rewrite the brief for it, or stop using one brief for all?
The choices
Rank the models and keep the "best". Keep one brief and blame the model that did not suit it. Or treat each model as a collaborator with a disposition — one wants a crisp task, another wants to know why first, a third pushes back on ambiguity — and adjust the brief to the disposition before judging the output.
The consequence
Rank the models, and you mistake fit for quality and lose the model that would have been right for the next task. Blame the model, and you learn nothing. Brief for disposition, and you pay a little time per model and get each one's strengths. You also risk reading personality into what is really a mixture of training data, system prompt and sampling.
The case
A studio asks two models for a design-intent statement for a school in a hot-dry town. One returns a crisp five-point list, on target but thin. The other asks two questions about the site before it writes, then produces a paragraph that reads like a colleague. The intern reports that the second "is better". The senior reruns the first with the two questions answered in the brief. The list becomes a paragraph too.
Our position
Different models show recognisably different behavioural profiles under a given version and configuration, and you get more by briefing for that profile than by assuming every model listens the same way. "Temperament" is our shorthand for the profile. It is real enough to plan for, even though it comes from training choices, system prompts and sampling rather than from a character — and it is measured on your tasks, never inferred as personality.
Why we hold it
In teaching, the students who adjusted for the profile got better work from every model. The students who ranked models got good work from one and gave up on the rest. The habit carries over when the models change.
The strongest objection
"Temperament" may not be a stable thing at all. What the word bundles together — system prompt, sampling settings, refusal training, routing between models, context length, verbosity defaults, which tools are available, the interface — can each vary independently, and often do, release by release. If they vary independently there is no single underlying disposition to brief for; there is a list of settings, and the student has learned to read character into a configuration. Briefing for "the model's temperament" may then be less useful than briefing for the specific settings, and less honest.
What would make us revise it
If the profile differences between the major models shrink to noise on studio tasks, or if the components above are shown to vary so independently that a profile measured in March predicts nothing in June, or if wrappers expose the configuration directly so that "brief for the model" becomes "set the settings", the card retires into CHATBOT on Foundations. Reviewed every edition, with the same paired-brief test.
Take it to crit
Did you treat the model's temperament as something to brief for, or expect every model to respond the same way? Show the same brief in two models, and the adjustment you made.
How it works
Labs now deliberately shape a model's behaviour during post-training, and some describe how they do it. Add a system prompt and a sampler on top and you have a behavioural profile. It is stable within a release and configuration and unstable across them, so record the model version and settings beside any profile you measure. The review date in the status strip is doing real work here.
What this idea builds on
What this idea opens up
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.