TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Ethics & Provenance · AGE 16—21 · POSITIONAL · HELD
Datasets reproduce social visibility
The people not in the data become the people the system is built against.
Our position
Bias in a model is not only "it looks Western". A dataset is a record of who got seen. Who is photographed and who is not. Which streets are mapped and which are called "encroachment". Whose home is a household on a form and whose is a structure. Whose language is on the internet. Who is in the archive and who was never written down. In India those lines run along caste, class, gender, disability, informality and language. A model trained on that record inherits its blind spots, and a designer who uses the model inherits them too. The Lab holds that a student must learn to ask who is missing from the data before asking what the data says.
Why we hold it
Architecture draws for whoever is on the survey. A site plan that shows the plot and not the tea stall at its gate, a housing brief whose "household" assumes a nuclear family, a render with nobody in a wheelchair and nobody selling anything — these are not the model's inventions. They are the data's habits, and the data's habits are society's. A designer who cannot see the gap will build it. Indian researchers who interviewed practitioners working with marginalised communities found whole communities missing or misrepresented in datasets, and safety apps populated by middle-class users marking Dalit, Muslim and slum areas as unsafe. That is the mechanism in one example: the people who are not in the data become the people the system is built against.
The strongest objection
A technical map cannot fix society and should not try. Talamana teaches how models work; caste, class and gender are the province of sociology, law and politics, and a literacy curriculum that takes a position on them is doing something other than teaching AI. Worse, it risks teaching a politics under the name of a mechanism, and a student who disagrees with the politics will distrust the mechanism. A card about datasets should say "datasets have gaps; check them" and stop there.
We answer it this way. The objection is right that the Lab cannot fix society, and right that a literacy map must not smuggle politics in as mechanism. Our claim is narrower than it sounds. Naming caste, class, gender, disability, informality and language is not a politics; it is the documented list of axes along which Indian data goes missing, and "check for gaps" without saying where to look is an instruction nobody can follow. We state the position as ours, with this objection beside it, so that a student who disagrees can see exactly what they disagree with and still use the method.
What would make us revise it
Evidence that the gaps are closing on their own: measured Indian datasets — built-environment, street, household — in which informal settlements, women, Dalit and Adivasi communities, disabled people and speakers of languages other than English appear in proportion to their presence, and models trained on them whose default street, house and household look like the ones most Indians live in. If the record starts to see everyone, a card telling students to look for who is missing becomes history rather than method, and it moves to LINEAGE. We check this against published dataset documentation every edition.
Try it
Take any site you have surveyed. List everyone who was on it that day, including the people working, selling, waiting or passing through. Then look at your site drawing and your generated context views. Count how many of those people, or their traces, appear. Write one line on who the drawing could not see, and one on why.
Take it to crit
Ask the student who is missing from their data, their drawing and their renders — and whether the absence was the model's, the survey's or theirs. Watch whether the answer names a group or a feeling.
How it works
A dataset is made from what exists in a collectable form: photographs on the open web, maps that someone drew, forms that someone filled, text in languages that are online. Each of those is already filtered by who had a camera, who was counted, who was mapped, and whose language was written. A model learns the filtered record and returns it with the gaps smoothed over, which makes the gaps harder to see than they were in the source. Sambasivan and colleagues, in a 2021 study of fairness in India built on thirty-six interviews, describe data points missing "because of social infrastructures and systemic disparities", entire communities "missing or misrepresented in datasets", and proxies — a surname, a pin code, an occupation — that carry caste, religion or class whether or not the dataset names them. TILTED MIRROR on this map measures the geographic lean; this card is the same mechanism read along social lines. What is not yet measured, to the Lab's knowledge, is the specifically architectural version — whose houses are in the building-image datasets — which is why this is held as a position, argued and dated, rather than asserted as a fact.
What this idea builds on
What this idea opens up
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.