TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 17—19 · FACTUAL · EVOLVING

Synthetic data

A model trained mostly on model output drifts towards the average of the average.

The idea

Training data used to mean things people made: photographs, drawings, sentences. Now a growing share of what models learn from is synthetic, meaning text and images produced by earlier models. Some is made on purpose, to fill gaps. Some arrives by accident, because the open web is filling up with generated material. Synthetic data is useful. It can be made to order and labelled exactly. It also has a known danger. A model trained mostly on model output drifts towards the average of the average. Rare things — the unusual house, the minority building culture — fade first.

Why it matters

The question "where does the model's world come from?" now has a second answer: partly from other models. A designer who works from generated references, using tools that were trained on generated references, is looking at copies of copies.

See it in the studio

A student's moodboard is twelve generated courtyards. Each looked plausible because it was built from the average of photographed courtyards. The next model, trained partly on images like those, learns an even narrower idea of a courtyard. The real one in the student's own town — with its odd proportions and its tulsi plant — is now two steps further from anything a tool will offer.

Watch for this

Assuming synthetic data is always a problem, or never one. Carefully made synthetic data is how some models get good at rare cases. Unfiltered model output is how they lose them. The question is who made it, and whether anyone was checking.

Try it

Generate "a traditional Indian house" ten times and save the set. Use three of those images as references and generate ten more. Compare the second set to the first, and both to a photograph of a real one. Describe what narrowed.

Prove it

Explain why a model trained on model output tends to lose its rare cases, and name one situation where synthetic data is the right choice anyway.

How it works

Researchers trained models on their own output again and again, generation after generation, and measured the tails of the data — the rare cases — disappearing. They named the effect model collapse. In practice, labs now mix synthetic data deliberately with curated human data and filter what they scrape. So the pure case is a warning, not a forecast. How much of the open web is generated, and how well filtering works, are both moving numbers. That is why this card is EVOLVING.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON