TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · The Brief · AGE 16—18 · FACTUAL · EVOLVING

Briefing in your own language

Your own language carries design ideas that English flattens.

The idea

When you write in Kannada, Hindi, Marathi or Tamil, two things can happen. The tokeniser — the part that chops text into pieces before the model reads it — may cut your script into far more pieces than English, so the same sentence costs more and fills the window faster. And the model may have seen less writing in your language, so its habits there are thinner. Both gaps vary by language and tokeniser, so test the model you are using. You can still brief in your own language, and sometimes you should: a jagali has no clean English name.

Why it matters

Your own language carries design ideas that English flattens. Where a system supports English better, that is a limit of the system, not of your language. A good brief often uses both, on purpose.

See it in the studio

A student briefs a house "with a jagali" — the raised plinth-seat along the front of a Karnataka house. Translated as "a porch," it returns a Western porch with a swing. The word was right. The model simply knew very little about it. Name it, then describe it: "a raised masonry seat along the front wall, under the eave, where visitors sit before entering."

Watch for this

Assuming that a fluent answer in your language is a correct one. A model sounds fluent before it is right, in every language. Less training data makes that gap wider.

Try it

Write one short brief in English and the same brief in your own language. Run both. Compare how faithful each output is to what you meant, and where each one drifted. Then try a third version: your language for the idea, English for the constraints. If your tool shows a token count, write the same Kannada sentence once in Kannada script and once in Roman letters, and compare what each one costs.

Prove it

Name one design idea from your own region that a model handles poorly, and explain, using tokens and training data, why it handles it poorly — on the model you tested, not on models in general.

How it works

Researchers have taken the same text in different languages and counted the tokens. The counts differ several times over, and the languages the tokeniser saw least pay the most. The gap is tokeniser-specific: a newer tokeniser can cut the same Kannada sentence into far fewer pieces than an older one did. How much of any one model's training was in your language is usually not disclosed, so "trained mostly on English" is a reasonable guess about many systems, not a fact you can check for yours. This card is EVOLVING because both things are changing: newer models train on more multilingual data, and tokenisers are being rebuilt. The gap is what we measure today, not a fixed fact.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON