TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Ethics & Provenance · AGE 17—19 · FACTUAL · EVOLVING

Old drawings as training data

A model trained on an archive learns only what the archive's makers chose to photograph.

The idea

Archives are becoming data. Plate books, institutional scans, a family's box of photographs: once digitised, each is a set of images a model can learn from, and documented web-scale datasets such as LAION-5B include scans from digitised collections. What a model learns is what the archive contains, and that is only what the archive's makers chose to photograph. Who may use a scan turns on three separate things: whether the photograph itself is still in copyright, the terms on which the collection was made available, and who owns the physical print — and none of them follows from the age of the subject. A studio's own archive — drawings, photographs, surveys — is a corpus it can search today and, with care, train on.

Why it matters

Whatever the archive left out, the model will never know — but it will still answer confidently. A student who knows what a corpus leaves out knows where the model is guessing.

See it in the studio

The Lab's own study corpus is built around the 1866 photographic survey of the Dharwar district and Mysore — the early plate survey of the Deccan by Taylor and Fergusson. Its subjects are Halebidu, Belur, Gadag and Vijayanagara. There is not one frame of Dharwad town, because the colonial camera went to temples, not to the administrative town. A model trained on images labelled "Dharwar" would learn a district of monuments and no city at all — and would produce one, confidently, if asked.

Watch for this

"It's from 1866, so it's free." Three questions hide in that sentence. Is the photograph out of copyright? That depends on the photographer's dates, not the building's — an old subject says nothing about a photograph taken last year. May you use this scan? That depends on the archive's access terms, which are a contract and can restrict a scan of a photograph that is itself free. Whose is the print? Owning the object is not owning the copyright, and a family or a photographer's estate can hold either without the other. An old building does not make the image free to use.

Try it

Pick a public-domain archive — the Internet Archive's texts hold many Indian survey volumes and district gazetteers. Choose one volume on a place you know. List what it photographs and what it does not: which buildings, which streets, which people. Write one sentence on what a model trained only on that volume would believe about the place.

Prove it

Explain, with one real archive as the example, how the choices of its makers would become the beliefs of a model trained on it — and name one thing about the place the model could never learn from it.

How it works

The large image models learned from web-scale datasets, and documented ones such as LAION-5B gathered scans from digitised collections among billions of other images; WHOSE WORK on this map carries the question of consent for that, and the courts' early answers. A studio archive is a smaller and different case. Searching it by meaning — THE ARCHIVE on Studio Practice — needs only good description. Training on it needs ownership of the drawings, permission from the people in the photographs, and an honest account of what the corpus is missing, because fine-tuning a model on a studio's own work teaches it the studio's habits and its blind spots alike. Rights are layered, and the layers are different kinds of thing: copyright in the original photograph or drawing; access terms for the scan and the catalogue record, which are contract, not copyright; and, where a database shows enough original selection or arrangement, a copyright in the database itself. "Public domain" is a claim about the first only. A faithful scan does not by itself earn a new copyright — whether it ever can is argued differently in different countries — so an archive that wants to control its scans usually does so by the terms of access, not by a fresh copyright. READ THE OLD DRAWING is what you do with one document; this card is what happens when thousands of them become training data. EVOLVING because archives are being digitised and licensed on changing terms, and because what training on them legally requires is still being decided.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON