Runway's Next World Model Can Be Directed Like a Film, or Played Like a Game
Runway's GWM Worlds 2 generates real-time interactive worlds that can be scripted like a film or played like a game, using the same underlying model.
Runway released GWM Worlds 2, a real-time interactive world model generating continuous 720p video at 24fps with 48kHz audio, extending the GWM Worlds research it first showed in December. The genuinely notable part for filmmakers specifically: Runway built three distinct ways to use it, and one of them is explicitly designed around directing, not playing.
Three Ways to Use the Same Model
Runway lays out the modes directly. Ahead of time, a user authors the full timestamped sequence of actions upfront, then the model generates the entire video and audio from that script, Runway's own stated use case here is "filmmaking, directing." Turn-based generation pauses at decision points for a user to choose what happens next, aimed at visual novels. Real-time generation runs continuously, responding to live input with no pauses at all, the mode built for games and interactive experiences, and the hardest of the three technically, since it requires reacting within tens of milliseconds.
That's a meaningfully different pitch than a typical AI video tool. Rather than one interface for one use case, Runway built a single underlying model flexible enough to be scripted like a film or improvised like a game, depending entirely on how a user chooses to input actions.
WorldPrompt: Separating What Persists From What Changes

The technical foundation is something Runway calls WorldPrompt, a formal structure splitting a generated world into two layers. Persistent context covers the genesis prompt, the environment, its subjects, and the physical rules governing it (gravity, collision, character abilities, camera perspective), plus a first frame to ground everything visually. On top of that sits a timestamped event stream: individual actions, movement, gestures, dialogue, sound, each addressed to a specific subject or to the scene itself, with start and end times that can overlap freely.
In Runway's own example, a director could script an entire street-corner conversation between two characters, specifying exactly when each line of dialogue happens, when a character sweeps leaves, when the camera follows a specific subject, entirely upfront, without touching a live keyboard binding during generation. That's a genuinely different authoring model than "type a prompt, get a clip," closer to blocking a scene than describing one.
Where This Sits Relative to Runway's Video Tools

GWM Worlds 2 is built on Runway's foundational audio-visual generation model, the same lineage as Gen-4.5, but it's a separate research preview, not a feature inside Runway's existing filmmaking product. Access is limited to a contact-form request, no public pricing or general availability has been announced.
Runway is candid about current limitations too: visual details and geometry can drift during quick camera movement, long-term consistency isn't fully solved, and the system doesn't support image references beyond the first frame.
Competitive Context

This continues a pattern BRC has tracked closely. GWM Worlds 2 follows GWM-1's December debut and Runway's Solaris interface-generation tool from earlier this week, both signs of a company explicitly betting on general world simulation as a bigger opportunity than video generation alone, backed by a $315 million Series E explicitly tied to that strategic pivot.
It also sits in similar territory to H3-World, the recent MiniMax H3 adaptation adding WASD keyboard control through a lightweight LoRA.
The approaches differ meaningfully: H3-World retrained a tiny fraction of an existing model's parameters to add fixed keyboard controls, while GWM Worlds 2 is a purpose-built model designed from the ground up to generalize to arbitrary, freely-described actions rather than a preset control scheme, a harder problem with a correspondingly more ambitious scope.
The Signal in the Noise

The ahead-of-time authoring mode is the part worth watching if you work in narrative filmmaking specifically. It reframes AI video generation less as "prompt, get a clip, prompt again" and more as scripting a scene's full timeline in advance, actions, dialogue, camera moves, all specified before generation runs.
Whether that actually becomes a practical filmmaking tool depends entirely on solving the consistency and drift issues Runway itself flags as unresolved, but the underlying idea, a single system flexible enough to serve both a director's script and a player's live input, is a genuinely different framing than most AI video tools currently offer.
Would scripting a full scene's actions and camera moves in advance, rather than generating and reviewing clip by clip, change how you'd approach previsualization or blocking?
The Details

- Model: GWM Worlds 2, released September 3, 2026
- Output: continuous 720p video, 24fps, 48kHz audio, no fixed session length
- Three use modes: ahead-of-time (filmmaking/directing), turn-based (visual novels), real-time (games/interactive experiences)
- Core framework: WorldPrompt, splitting persistent world context from a timestamped, freely-overlapping action event stream
- Extends: GWM Worlds (December 2025), adding generated audio and richer subject/scene control
- Multiplayer supported: multiple users can control separate subjects or the world itself simultaneously
- Access: research preview, contact-form request only, no public pricing
- Known limitations: visual/geometric drift during fast camera movement, imperfect long-term consistency, no image references beyond the first frame