TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · The Brief · AGE 16—18 · FACTUAL · EVOLVING
Briefing in your own language
Your own language carries design ideas that English flattens.
The idea
When you write in Kannada, Hindi, Marathi or Tamil, two things can happen. The tokeniser — the part that chops text into pieces before the model reads it — may cut your script into far more pieces than English, so the same sentence costs more and fills the window faster. And the model may have seen less writing in your language, so its habits there are thinner. Both gaps vary by language and tokeniser, so test the model you are using. You can still brief in your own language, and sometimes you should: a jagali has no clean English name.
Why it matters
Your own language carries design ideas that English flattens. Where a system supports English better, that is a limit of the system, not of your language. A good brief often uses both, on purpose.
See it in the studio
A student briefs a house "with a jagali" — the raised plinth-seat along the front of a Karnataka house. Translated as "a porch," it returns a Western porch with a swing. The word was right. The model simply knew very little about it. Name it, then describe it: "a raised masonry seat along the front wall, under the eave, where visitors sit before entering."
Watch for this
Assuming that a fluent answer in your language is a correct one. A model sounds fluent before it is right, in every language. Less training data makes that gap wider.
Try it
Write one short brief in English and the same brief in your own language. Run both. Compare how faithful each output is to what you meant, and where each one drifted. Then try a third version: your language for the idea, English for the constraints. If your tool shows a token count, write the same Kannada sentence once in Kannada script and once in Roman letters, and compare what each one costs.
Prove it
Name one design idea from your own region that a model handles poorly, and explain, using tokens and training data, why it handles it poorly — on the model you tested, not on models in general.
How it works
Researchers have taken the same text in different languages and counted the tokens. The counts differ several times over, and the languages the tokeniser saw least pay the most. The gap is tokeniser-specific: a newer tokeniser can cut the same Kannada sentence into far fewer pieces than an older one did. How much of any one model's training was in your language is usually not disclosed, so "trained mostly on English" is a reasonable guess about many systems, not a fact you can check for yours. This card is EVOLVING because both things are changing: newer models train on more multilingual data, and tokenisers are being rebuilt. The gap is what we measure today, not a fixed fact.
What this idea builds on
What this idea opens up
- Nothing yet names this as a foundation.
Sources
- prompt-theory-primer
- social-impact-ethics
- Petrov et al., Language Model Tokenizers Introduce Unfairness Between Languages
- Tokens, Context, and Why AI Forgets
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.