Bilawal Sidhu tests Genie 3 as an interactive world model rather than a conventional video generator. Text and image prompts become short-lived environments that respond to keyboard input, letting him move through spaces inspired by spacecraft, city streets, the Moon, restaurants, cathedrals, scanned places, underwater scenes, and toy-scale driving tracks.
The strongest results come from visual priors the model has learned from video and imagery. Reflections, lens flare, changing exposure, water caustics, atmospheric effects, material detail, and camera motion often appear without being explicitly programmed. The system can also extend a scene beyond its initial frame, synthesize new viewpoints, and sometimes respond to nearby vehicles or surfaces.
The hands-on tests also expose important limits. Controls have noticeable streaming latency, characters and faces can drift, geometry changes while the camera moves, collisions are unreliable, and users can clip through objects or fall outside the generated space. Interactions remain shallow, and the one-minute session length prevents sustained exploration or persistent game state.
Sidhu sees Genie 3 as a new creative medium rather than an immediate replacement for Unreal Engine. It is already useful for concept exploration, virtual production, memory-like scene reconstruction, and rapid spatial prototyping, but a production game engine still offers deterministic physics, authored mechanics, durable worlds, and precise control.
Watch on YouTube


