TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Ethics & Provenance · AGE 16—18 · FACTUAL · DURABLE
The labour inside the dataset
A model's provenance includes the people who labelled, cleaned and filtered its data.
The idea
Two kinds of dataset. In supervised learning, someone says what is in the picture: draws the box around the hand, tags the road sign, marks the scan. The big image models were mostly pretrained the other way, on image–caption pairs inherited from the web, alt-text people had already typed for other reasons, filtered by a model called CLIP rather than labelled by hand. Human labour still enters: cleaning, filtering, evaluation, checking outputs. People do this work in Jharkhand and Nairobi, paid per task, invisible. "Where did the data come from?" honestly includes them. Provenance is whose hours too.
Why it matters
An architect already knows that a building carries the labour of the people who made it, whether or not the photograph shows them. The same is true of the model. Knowing this changes how you talk about "the AI" doing something — and knowing which labour went where keeps you from telling the story wrong.
See it in the studio
A tool tags your site photographs: vehicle, vendor, temporary structure. Each of those categories was learned from many thousands of images that someone labelled by hand, deciding what counted as a vendor and what counted as a structure. Their choices are now built into the tool, and into your site analysis. The image generator on the next tab learned differently — from captions nobody was paid to write — and its blind spots come from what the web bothered to caption.
Watch for this
The word "automatically". When a caption, a tag or a rating is described as automatic, ask what taught the automation. Sometimes a person, paid by the piece. Sometimes a caption someone typed years ago for a different reason. Both are provenance.
Try it
Watch Humans in the Loop, or read the reportage it grew from. Then list three things you used this week that needed labelled data, and one that was trained on inherited captions instead. For each, write one line: who most likely did the labelling — or wrote the caption — and what they would have been asked, or not asked, to decide.
Prove it
Explain the difference between a dataset labelled by people and one that inherited its captions from the web, say where human labour enters each, and why a model's provenance chain is incomplete without the people who did it.
How it works
Labelled data is how supervised learning gets its answer sheet — the boxes, tags and masks behind a site-photo classifier. The large image–text models were pretrained on web-scale pairs: LAION-style datasets were built from existing alt-text associations and filtered with CLIP, which is why web-scale learning became possible without anyone drawing boxes around billions of objects. Human labour did not leave; it moved — into cleaning and filtering, evaluation sets, benchmark-making, moderation, red-teaming and specialist review. All of it is piece-work, commissioned through platforms and subcontractors, which is why the workers rarely appear in a paper or on a product page. Karishma Mehrotra's 2022 reportage followed women doing labelling work at a data centre in Jharkhand; Aranya Sahay's 2024 film built its protagonist, Nehma — an Oraon woman who tags training images — from that reporting. The card beside this one, The people inside the model, covers the labour after training: preference rating, moderation and safety work, where the material itself can harm the worker.
What this idea builds on
What this idea opens up
- Nothing yet names this as a foundation.
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.