TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 19—21 · FACTUAL · CONTESTED

3D and world models

A world model holds an environment's state over time, and a splat is not one.

The idea

Beyond flat images sit two more kinds of system. 3D generators produce a representation of shape and appearance — a mesh, a point cloud, a radiance field, a splat — that you can orbit and, for a mesh, section and measure. A world model is something else: a learned model of an environment's state and how that state changes over time or under an action. It may never show a picture. Some recent systems render an interactive scene you can move through, but that is one use, not the definition. A splat is not a world model.

Why it matters

The day a model can hold a room is the day this whole strand changes. Knowing how far off that day is, and what people mean when they claim it has arrived, is part of being literate.

See it in the studio

A text-to-3D tool gives you a mesh of a pavilion you can open in your modelling software. It is watertight nowhere and its roof is a skin with no structure. It is still the first generated thing on this strand you can put a scale to. A "world" demo lets you walk down a generated street. In many demos, turn round and the street behind you has been invented again; the newest hold it for minutes. Ask how long, and what happens after.

Watch for this

The word "world" doing marketing work. A video model that produces believable motion is not, because of that, holding a world. Ask: can I go back to where I was and find it unchanged?

Try it

Generate one object three ways: as an image, as a 3D mesh, as a short video. For each, write what you can measure, what you can move through, and what stays put when you look away.

Prove it

Explain the difference between a generated image, a generated 3D model and a simulated world you can walk through, and name one thing a world model would offer a spatial designer that image generation never can.

How it works

3D generation usually lifts 2D generative priors into geometry, or trains directly on 3D data, which is scarce. A radiance field (NeRF) or a Gaussian splat stores appearance from many viewpoints; it is not a watertight solid you can section like a BIM model. "World model" has two meanings in use. The older research sense is a learned internal simulator that an agent plans with. The newer product sense is a generative system that renders consistent, explorable scenes — Google DeepMind's Genie 3 (August 2025) is described by its makers as a general-purpose world model whose environments stay largely consistent for several minutes as you move. CONTESTED: OpenAI framed its video model as a "world simulator". Many researchers dispute that predicting believable pixels amounts to modelling a world, and argue that a genuine world model must represent state, not appearance. This card names the disagreement rather than settling it.

What this idea builds on

What this idea opens up

Sources

Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.

Age grows from 11 at the centre to 22 at the edge, and six sectors show the learning strands. Tab into the map and the arrow keys step from idea to idea, following the links where there is one. Enter opens the idea under the cursor, and E reads out its links and the reason recorded on each. Press slash for Search, question mark for the full key list, and Escape to leave. Open Ideas for the complete readable list, including what each idea builds on and what it opens up.

LOGIKA · RBDS AI LAB INDIA
ON-RAMP · AGE 11 · FIRST ENCOUNTERS, NOT GATES — IDEAS · — DEPENDENCIES
DONE
OPENS NEXT
SOLID — STANDS ON