TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Ethics & Provenance · AGE 17—20 · FACTUAL · CONTESTED
Your work is in the scrape
Public image datasets were built from the web you use, and no one was asked.
The idea
The large public image datasets were built from the open web. For the best-documented, LAION-5B: a non-profit crawl, Common Crawl, had already copied billions of pages; LAION read that crawl's records, pulled out image links with their alt-text, downloaded the images only long enough to filter them with a model called CLIP, and published a list — URLs and captions, not the images. Model makers fetched the images from the list. Individual permission was generally not sought. That web may include your portfolio site, your Instagram, your college's exhibition page. Consent, where it exists, has mostly come afterwards.
Why it matters
"Scraped from the web" sounds like it happened to distant strangers. It happened on the web you use. Your own work may be in there.
See it in the studio
A final-year student posts a thesis portfolio online in May. By the following year a generated image carries the unmistakable colour grading of that portfolio. Nothing can prove the link, and nothing can undo it.
Watch for this
"I'll just opt out." Opt-out only reaches datasets that honour it, only from the date you ask, and only the copies that have not already been trained on. Anything already learned stays learned.
Try it
Open the LAION-5B documentation and read how the dataset was built: who crawled, what was filtered and with what, and what was actually published. Then search your own site's images by reverse image search and note where else they appear. Write three lines on what recourse you would actually have.
Prove it
Explain, step by step, how a public dataset ends up pointing at images no one was asked to give, and describe what it would mean for your own posted work to sit in a training set — and what recourse, if any, you would have.
How it works
The law is moving in three directions at once. In the EU, rights-holders can reserve their work against machine reading, and the 2024 AI Act asks general-purpose model makers to honour such reservations. In India, a DPIIT working paper in December 2025 proposed the opposite: a mandatory blanket licence for training with royalties paid into a collective, and no opt-out — a proposal, not law, as of this edition. And on 24 July 2026 the Delhi High Court's interim ruling in ANI v OpenAI treated training on news articles as prima facie fair dealing on the facts before it; the suit continues. Three places, three answers, none of them final. That is why the card is CONTESTED.
What this idea builds on
What this idea opens up
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.