Bridge 6 Before Session 3: Video

Time Makes Failure Visible

Video systems must maintain relationships across time. Architectures use different temporal conditioning strategies, and inconsistency can emerge gradually or abruptly. A face may shift, a detail may appear, or a physical relationship may break. Motion makes these changes easier to notice than a single frame does.

Authored teaching simulation. The five frames below were designed for this lesson; they are not output from a video model and do not represent one universal architecture. The sequence supports observation practice: name what changed before proposing why.
Step through 5 frames — watch identity drift over time
Frame 0 of 4
No drift yet
Cumulative drift — all 5 frames at once
Attribute
F0 (prompt)
F1 (+2s)
F2 (+4s)
F3 (+6s)
F4 (+8s)
Face
⚠ slight
✗ different
✗ different
Hair
⚠ longer
✗ wavy
✗ wavy
✗ wavy
Jacket
✓ red
⚠ brighter
⚠ red-orange
✗ orange
✗ orange
Earring
+ added
persists
persists
Setting
⚠ light
✗ path changed
Key line "Video systems must preserve relationships across time. When those relationships change, describe the output first; architecture or workflow evidence is needed to explain the cause."
Documented mechanism, with scope. Video architectures vary. Some diffusion systems model fixed blocks jointly across space and time; other systems and production workflows use different forms of temporal conditioning, reference media, tracking, masks, or compositing. The Video Diffusion Models paper, for example, describes a space-time diffusion architecture for fixed frame blocks. This authored sequence does not reveal which mechanism caused a failure in any real clip.

Now open the tool

Temporal Telephone is a human drawing relay, not a video-model implementation. Compare what participants preserve, change, or add under its reference conditions, then name what that human activity cannot establish about a real generator.