TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 17—19 · FACTUAL · EVOLVING
Synthetic data
A model trained mostly on model output drifts towards the average of the average.
The idea
Training data used to mean things people made: photographs, drawings, sentences. Now a growing share of what models learn from is synthetic, meaning text and images produced by earlier models. Some is made on purpose, to fill gaps. Some arrives by accident, because the open web is filling up with generated material. Synthetic data is useful. It can be made to order and labelled exactly. It also has a known danger. A model trained mostly on model output drifts towards the average of the average. Rare things — the unusual house, the minority building culture — fade first.
Why it matters
The question "where does the model's world come from?" now has a second answer: partly from other models. A designer who works from generated references, using tools that were trained on generated references, is looking at copies of copies.
See it in the studio
A student's moodboard is twelve generated courtyards. Each looked plausible because it was built from the average of photographed courtyards. The next model, trained partly on images like those, learns an even narrower idea of a courtyard. The real one in the student's own town — with its odd proportions and its tulsi plant — is now two steps further from anything a tool will offer.
Watch for this
Assuming synthetic data is always a problem, or never one. Carefully made synthetic data is how some models get good at rare cases. Unfiltered model output is how they lose them. The question is who made it, and whether anyone was checking.
Try it
Generate "a traditional Indian house" ten times and save the set. Use three of those images as references and generate ten more. Compare the second set to the first, and both to a photograph of a real one. Describe what narrowed.
Prove it
Explain why a model trained on model output tends to lose its rare cases, and name one situation where synthetic data is the right choice anyway.
How it works
Researchers trained models on their own output again and again, generation after generation, and measured the tails of the data — the rare cases — disappearing. They named the effect model collapse. In practice, labs now mix synthetic data deliberately with curated human data and filter what they scrape. So the pure case is a warning, not a forecast. How much of the open web is generated, and how well filtering works, are both moving numbers. That is why this card is EVOLVING.
What this idea builds on
- A starting idea.
What this idea opens up
- Nothing yet names this as a foundation.
Sources
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.