TALAMANA · THE AI LITERACY MAP FOR ARCHITECTURE AND DESIGN · Generative Mechanics · AGE 19—21 · FACTUAL · EVOLVING
How video generation holds time
A video model must keep the world the same from one frame to the next.
The idea
A video model must do what an image model never has to: keep the world the same from one frame to the next. Light must stay where the sun put it. Materials must not move. A column must not become two when the camera turns. That steadiness is called temporal coherence, and it is one of the places these systems break — alongside physics, cause and effect, camera geometry and object permanence: a brick wall that turns to render mid-pan, a window that slides along a façade. Generation runs across many frames at once, trying to hold one scene through time.
Why it matters
For architecture, a walkthrough is a claim about a space. If the building at second four is not the same building as at second one, the clip shows no space at all. It only shows the appearance of one.
See it in the studio
A generated approach to a school: gate, forecourt, verandah. Watch the columns. Frame by frame, the bay count changes, the verandah depth shifts, and the courtyard that should appear through the entrance never resolves. The clip looks good. The building has not been held.
Watch for this
Judging a clip at playback speed. Coherence failures hide in motion. Scrub it, or pull stills every half-second and lay them in a row.
Try it
Generate a ten-second walkthrough of a simple building. Extract a frame every second. Pin them up. Circle what changed that should not have.
Prove it
Explain why a generated clip can shift in material or lighting from one frame to the next, and name what temporal coherence asks a model to hold that a single still never has to.
How it works
Current video models extend diffusion across space and time. They denoise a block of frames together, so that attention runs between frames as well as within them, often in a compressed latent space. That is what buys coherence, and the coherence is still partial. EVOLVING: this is one of the fastest-moving corners of the field, and a year changes which failure is the visible one.
What this idea builds on
What this idea opens up
Sources
- practical-applications
- TEC.U
- OpenAI, Video generation models as world simulators (Sora), 2024
- Diffusion — How AI Paints from Noise
Open this idea on the map · The complete map · Logika · RBDS AI Lab, India · revised every edition.