Meet Atlas: World Labs' New Model Beats MiniMax H3 and Reframes Real Footage

World Labs' new Atlas model beats MiniMax H3 on camera control and can reframe real footage into camera angles that were never filmed.

Share
Meet Atlas: World Labs' New Model Beats MiniMax H3 and Reframes Real Footage

World Labs, founded by AI researcher Fei-Fei Li, released Atlas, a single model that generates camera-controlled video, reconstructs real scenes in 3D, and simulates space and time from ordinary footage. In head-to-head testing, human raters preferred Atlas's camera-controlled output over MiniMax H3 75% of the time, and over Gemini Omni Flash, Seedance 2.5, and FLUX 3 by even wider margins.

What Atlas Actually Does

Atlas takes one or more reference images and generates new views from any camera position or angle, up to a full minute of coherent video at 1440p. Unlike most video models, which interpret camera movement from a text description, Atlas accepts precise camera geometry as a native input, letting a user specify an exact camera path rather than describing one in words and hoping the model interprets it correctly. World Labs frames this as the difference between staging a scene as a director and "pulling the lever of a slot machine."

The model also reconstructs real-world spaces from as few as two or three photographs, generating both new 2D viewpoints and explicit 3D outputs, point clouds and 3D Gaussian splats, usable directly in VFX, robotics, and game production pipelines. World Labs claims Atlas outperforms specialized 3D reconstruction models built specifically for that narrower task, not just general-purpose video generators asked to do something outside their core design.

The Feature That Actually Matters for Filmmakers

The most concrete, immediately useful capability here is video reframing. Atlas can take footage from as few as three ordinary cameras, World Labs' own demo reel was shot on cell phones and action cameras mounted on tripods and clamps, and reconstruct the scene well enough to generate entirely new camera angles and freeze-frame "bullet time" moves that were never actually captured. That's a real, practical VFX capability normally requiring a dedicated multi-camera volumetric capture stage, made possible instead with a handful of consumer cameras and no specialized equipment.

The Robotics Angle, and Why It's Not a Detour

A meaningful chunk of Atlas's design targets robotics simulation, reconstructing physical spaces from casual phone video, then generating what a robot's onboard camera would see as it moves through that space, including how objects respond to physical interaction. That might read as unrelated to filmmaking at first glance, but it's built on the exact same underlying spatial reconstruction technology powering the reframing and camera-control features, evidence the core architecture generalizes well beyond a single narrow use case rather than being a video tool with robotics bolted on as an afterthought.

Competitive Context

World Labs' benchmark results are self-reported, evaluated using third-party human raters rather than an independent lab, worth keeping in mind before treating the MiniMax H3 comparison as settled fact.

But the underlying architectural choice, treating camera position as a native, precise input rather than a fuzzy text description, addresses a real, persistent limitation across nearly every other video generation model on the market right now, imprecise camera control has been one of AI video's most consistent complaints across tools BRC has covered this year.

The Signal in the Noise

Atlas represents a different bet than most AI video companies are making. Rather than chasing marginal quality gains on straightforward text-to-video generation, World Labs built toward spatial precision and 3D consistency as the core differentiator, a genuinely harder technical problem, but one with real payoff for anyone doing VFX, previs, or reconstruction work rather than just generating standalone clips. Access is limited to early partners for now, with no public pricing or general availability announced yet.

Would reliable, precise camera control be worth switching tools for, even if it meant working with a less mature or widely available platform than the AI video tools you already use?

Specs & Pricing

  • Model: Atlas, World Labs, released September 1, 2026
  • Architecture: multimodal autoregressive diffusion transformer, natively processes text, images, video, camera poses, and 3D depth maps
  • Camera-controlled generation: up to 1 minute of video at 1440p from 1-6 reference images
  • 3D reconstruction: functional from as few as 2-3 input images, scales to 100+ images for full environment recreation
  • Outputs: 2D video, point clouds, 3D Gaussian splats
  • Benchmark results (human preference over Atlas): MiniMax H3 75%, Gemini Omni Flash 81%, Happy Horse 1.1 86%, FLUX 3 93%, Seedance 2.5 94%
  • Access: early access only, request-based, no public pricing announced

Resources & Reads