Interactive generative video produces frames while accepting inputs that influence what happens next. Unlike a fixed generated clip, the system responds to navigation, movement or other controls during the experience.
The capability overlaps with world models when the generated environment maintains spatial and temporal consistency. Low latency and coherent state are important for interactions to feel responsive and believable.
ELI5
Interactive generative video lets viewer input affect scenes that an AI system creates during the experience. Messages, votes, or game actions can guide what the system generates next.
For example, viewers might vote for a character to enter a forest, and a later generated scene follows that choice. The response is not instant because some video may already be queued and new scenes take time to render.



