Want Better AI Video Results? Learn to Prompt Like a Cinematographer, Not a Novelist
Static, unmotivated AI video shots usually come down to vague prompting. Creator Dan Kieft breaks down 30 camera movements and how to actually prompt them like a cinematographer.
Most people prompting AI video models describe what they want to see. Creator Dan Kieft's approach is different: describe how the camera should move, using the same specific, established vocabulary a real cinematographer would use on set. His recent breakdown covers 30 distinct camera movements and how to actually prompt each one, a genuinely useful reference for anyone whose AI-generated shots keep coming out static or motivated by nothing in particular.
Why Camera Language Matters More Than Descriptive Prompting
Kieft's core argument is direct: "The way you use the zoom in AI filmmaking really matters a lot when it comes to trying to portray the mood." That's the actual insight worth sitting with. A prompt like "camera slowly gets closer to the character" is vague enough for a model to interpret loosely. A prompt using an established term, dolly in, crash zoom, push past, carries specific, well-documented meaning the model has almost certainly seen labeled consistently in its training data, giving it a much narrower, more predictable target to hit.
This mirrors a broader pattern worth understanding: AI video models respond better to the specific technical vocabulary of filmmaking than to conversational description, because that vocabulary is exactly how real footage gets labeled and discussed across the training data these models learn from.
The Movements, Organized by Category

Kieft's 30 movements break into logical groups worth understanding by function rather than as a flat list.
Basics: Static, Pan, Whip Pan, Tilt Up, Tilt Down, the foundational vocabulary every other movement builds on.
Zooms: Slow Zoom, Fast Zoom, Crash Zoom, each carrying a distinct emotional register, a slow zoom builds tension gradually, a crash zoom delivers sudden shock or emphasis.
Dolly and physical moves: Dolly In, Dolly Out, Truck Shot, Pedestal, Slider, Push Past, movements that physically relocate the camera through space rather than just reframing from a fixed point, producing a different, more immersive sense of movement than a zoom alone.
Tracking and following: Arc, Follow Shot, Reverse Tracking, Side Tracking, Low Tracking, Vehicle Tracking, Chase Shots, camera movements built around keeping pace with a moving subject, useful for action or dynamic character movement.
Human and handheld: Handheld, Body-Mounted/Snorricam, movements that introduce deliberate instability or a subjective, first-person physical quality to a shot.
Aerial: Crane Up, Crane Down, Drone, Helicopter, FPV (first-person view), movements built around vertical scale and establishing shots or high-energy pursuit sequences.
Special effects shots: Tilt-Shift, Infinite Zoom, Earth Zoom Out, Time-Lapse, more stylized, less naturalistic movements suited to specific creative effects rather than standard coverage.
Why This Matters Specifically for AI Filmmaking

Traditional film education already teaches this vocabulary as a matter of course, it's how a director communicates shot intent to a cinematographer and camera operator on set. What's genuinely new here is applying that same established language directly as prompt input for a generative model, treating the AI as though it needs the same specific direction a human camera operator would.
That reframes what's actually happening when someone struggles to get good results from AI video prompting: it's frequently not a limitation of the model, it's a vocabulary gap between what the person describes and the specific cinematic terms the model was more consistently trained to recognize.
Worth Knowing

Kieft's tutorial is built around Higgsfield's platform specifically, testing prompts against Seedance and Kling as the underlying models, so exact prompt phrasing may need adjustment when applied to other platforms with different training data and interpretation tendencies. The underlying principle, using precise cinematic vocabulary over vague descriptive language, should transfer across platforms even if specific wording needs tuning.
The Signal in the Noise
The actual skill worth developing here isn't memorizing 30 specific terms, it's internalizing the habit of thinking like a cinematographer before writing a prompt: what is this camera movement supposed to communicate emotionally, and what's the specific, established term for that movement, rather than describing the desired visual outcome in plain language and hoping the model interprets the intent correctly. That's a genuinely transferable skill as new models and platforms continue to emerge.
Which camera movement have you found hardest to actually get right in AI video prompts? Curious what's tripped you up, drop it in the comments.