AI Video Camera Movement Prompts: a Free 28-Prompt Library for Music Videos
This post contains affiliate links — if you sign up through them we may earn a commission at no extra cost to you (disclosure). Researched and edited for accuracy with AI assistance.
Quick answer: The fastest way to make AI music-video footage look filmed instead of generated is to give every clip one clear camera instruction that matches the song section: static close-ups and slow dolly-ins for verses, push-ins for the pre-chorus, crane, orbit, and crowd shots for the chorus, slow motion for the bridge, and pull-backs for the outro. Below is a free 28-prompt copy-paste library organized exactly that way, with syntax notes for Seedance 2.0, Kling 3.0, and Veo 3.1. All three models respond to standard filmmaking vocabulary — dolly, crane, orbit, tracking — they just disagree on where the camera phrase goes in the prompt and how many moves you can stack.
Why do camera prompts matter more than style keywords?
Style words ("cinematic," "moody," "4K") are the least reliable part of any video prompt. Camera language is the most reliable part — all three models' official and community guides are built around standard cinematography terms, and the models respond to that vocabulary far more consistently than to style adjectives. Google's official Veo 3.1 guide calls the cinematography element "the most powerful tool for conveying tone and emotion" and recommends putting it first in the prompt (source). Seedance-focused guides say the same thing from the other direction: one clean movement with one modifier works, while stacked contradictory instructions ("smooth aggressive rapid dolly") turn into noise.
Camera prompting has become enough of a bottleneck that people now sell prompt handbooks for it on Substack (example). You don't need one. The library below covers the music-video use case for free, and if you're starting from zero, our full walkthrough on turning a finished song into a video covers everything around the prompts.
How does each model want its camera instructions written?
The 28 prompts below are written camera-first, which is the safest default across all three tools. Here's how each model differs, as of July 2026:
| Model | Where the camera phrase goes | Vocabulary it responds to | Stacking limit | Clip length (July 2026) |
|---|---|---|---|---|
| Seedance 2.0 | Its own block — guides recommend [Subject] + [Action] + [Camera] + [Setting/Lighting] + [Style] + [Audio] | Pan, tilt, zoom, dolly, truck, pedestal, crane, orbit, tracking, whip pan, dolly zoom, rack focus | Two ideal, three max — join with "+" or "while" | 4–15s per generation (source); Video Extension chains segments (source) |
| Kling 3.0 | Near the beginning of the prompt, before the scene description (source) | Slow pan, tracking shot, smooth zoom, push forward, handheld, drone descending (source) — best written as how the camera behaves over time, not as keywords (source) | One primary movement per shot; stacking produces inconsistent results (source) | Extensions in 5s/10s hops to roughly 3 minutes total (source) |
| Veo 3.1 | First — [Cinematography] + [Subject] + [Action] + [Context] + [Style] (source) | Dolly shot, tracking shot, crane shot, aerial view, slow pan, POV shot (source) | No explicit official limit; Google's own examples each use a single movement, so keep to one (source) | 4/6/8s generations; ~148s via 7s extensions (source) |
Two practical notes. First, every prompt below says "the singer" — replace that with your consistent performer description (same wording every clip; see our character consistency guide for why exact repetition matters). Second, if a prompt underperforms in Seedance specifically, reorder it subject-first per the six-block structure above and keep the camera phrase intact. For a deeper head-to-head on which model fits which kind of video, see Kling vs Seedance for music videos.
What camera prompts work for verses?
Verses are for restraint. Small moves, close framing, one light source. Save the big vocabulary for the chorus so it has somewhere to go.
- Slow dolly-in from medium shot to close-up on the singer's face, shallow depth of field, background held static — the workhorse verse shot; intimacy without motion sickness.
- Static tripod shot, the singer framed in a doorway, only their hands and jaw moving, moody practical lighting — a locked camera makes small performance details read loud.
- Handheld medium close-up with subtle sway, the singer sitting on a bed playing guitar, morning window light — "subtle" is the load-bearing word; drop it and the shake takes over.
- Slow lateral tracking shot past the singer leaning against a wall, foreground objects drifting through frame — foreground occlusion adds depth all three models render well.
- Rack focus from a spinning record in the foreground to the singer's face behind it, close-up — a Seedance-listed special move; in Kling, phrase it as "focus shifts from the record to the singer's face."
- Slow orbit around the singer at a kitchen table, about 90 degrees over the clip, warm tungsten light — specifying partial rotation stops the model from attempting a full dizzy loop.
What camera prompts build tension in a pre-chorus?
The pre-chorus has one job: compress. Push-ins are the whole playbook — the frame tightens, the energy has nowhere to go, then the chorus releases it.
- Slow push-in on the singer's eyes, ending in extreme close-up, background lights blooming out of focus
- Low-angle push-in as the singer stands up, camera rising with them, dust hanging in the air
- Tracking shot from behind as the singer walks toward a lit stage door, pace gradually quickening
- Dolly-in on the singer while the crowd behind them comes into focus — two instructions joined the way Seedance's "+ / while" syntax expects; simplify to the dolly alone for Kling.
- Handheld push-in with slight shake increasing, the singer inhaling before the drop, lights flickering
What camera prompts make a chorus feel big?
Chorus shots buy scale: height, rotation, and crowds. These are also the most expensive shots to iterate on, so lock your performer reference before burning credits here.
| Prompt (copy verbatim) | Why it works |
|---|---|
| Crane up from a close-up on the singer to a wide overhead of the full crowd, stage lights sweeping | The classic chorus reveal — small to huge in one move. |
| Fast orbit around the singer mid-jump on stage, confetti in the air, strobing lights | One of the few places "fast" earns its keep; the strobe hides orbit artifacts. |
| Wide aerial view pulling back over a festival crowd, the singer centered on stage, golden hour | Aerial vocabulary is on Veo's official keyword list and lands reliably in all three. |
| Whip pan from the drummer to the singer at the moment the chorus hits, stage lights flaring | Gives you a natural beat-synced cut point when you edit. |
| Tracking shot gliding through a dancing crowd toward the singer, faces passing close to the lens | Foreground faces sell crowd density better than a wide shot alone. |
| Low-angle crane shot rising past the singer to reveal a wall of speakers and lights | Low angle plus rise reads as power — good for a final chorus. |
| Dolly zoom on the singer at center stage, background stretching, lights streaking | A special-effect move Seedance guides list explicitly; use it once per video, maximum. |
What camera prompts suit a bridge or breakdown?
The bridge is subtraction. Slow the camera, empty the frame, and the return of the final chorus doubles in impact.
- Slow motion close-up, the singer's hair and jacket moving in wind, camera static
- Ultra-slow push-in on the singer standing alone in an empty venue, house lights half up
- Slow overhead top-down shot, the singer lying on the stage floor, spotlight tightening around them
- Static wide shot, the singer tiny in frame against a huge empty room, one light source
- Slow orbit in slow motion, rain hanging in the air around the singer — orbit + slow motion is a two-instruction stack; fine for Seedance and Veo, split it for Kling.
How should the outro camera move?
Outros pull away. Reverse the verse's dolly-in and the video closes the loop visually, even if the viewer never consciously notices.
- Slow pull-back from close-up to wide as the singer turns away, lights dimming
- Crane up and pull back over the empty stage, cables and road cases in view, night
- Tracking shot following the singer down a corridor and out of frame, camera slowing to a stop
- Slow zoom out through a doorway, the singer left small in the lit room, hallway dark
- Aerial pull-back rising above the venue roof, city lights spreading below, ending on a hold — the final hold gives you a clean frame to fade out on.
How do you chain these shots into a full video?
None of these models generates a full song in one pass, so a music video is really an edit of 15–40 short clips. As of July 2026, Seedance 2.0 generates 4–15 seconds per clip and chains segments with Video Extension (source), Kling 3.0 extends in 5–10 second hops toward roughly 3 minutes — with its extension tutorial warning that quality softens past the 60–80 second mark (source) — and Veo 3.1 reaches about 148 seconds through 7-second extensions (source). The mechanics of chaining without drift are covered in our guide to making AI videos longer than 10 seconds.
The music-video-specific advantage of Seedance 2.0 is that it accepts your actual track as a reference input, so motion and cuts can follow the song rather than a text description of it — our Seedance music video tutorial walks through that workflow section by section. Before you start generating, budget it: a 3-minute video at one camera move per clip is 15–25 generations plus retries, and our credit calculator will tell you what that costs on each plan. One honest warning: the Seedance free tier is one watermarked, non-commercial credit per 24 hours — enough to test one prompt from this library per day, not enough to build a video. Treat it as a preview, decide, then either subscribe for a single project month or walk away.
Estimate your render cost with our free credit calculator.
Frequently asked questions
Do these camera prompts work in Seedance, Kling, and Veo?+
Yes. All three models respond to standard filmmaking vocabulary — dolly, crane, orbit, tracking, pan, aerial. The differences are structural: Veo 3.1 wants cinematography first in the prompt, Kling 3.0 wants the camera phrase near the beginning and only one movement per shot, and Seedance 2.0 treats camera as its own block and tolerates two or three chained movements joined with a plus sign or the word 'while'.
Can I combine two camera movements in one prompt?+
Depends on the model. Seedance 2.0 handles two movements well (three is the ceiling) when you join them with '+' or 'while', as in 'crane up + slow pan right'. Kling 3.0 is most reliable with a single primary movement per shot — stacking produces inconsistent results, with the movement often coming out different from what the prompt describes. Veo 3.1 sits in between: one clear movement is safest. When in doubt, split the idea into two clips and cut between them.
How do I sync camera movement to the beat of my song?+
Two ways. Seedance 2.0 accepts your music track as a reference input, so generated motion can follow the actual audio. For any model, the editing-room method works: generate clips with one camera move each, then cut between clips on the beat. Whip pans and push-ins that land at the end of a clip give you natural beat-synced cut points.
Why does my camera prompt get ignored?+
Three usual causes: the camera phrase is buried at the end of a long prompt (move it to the front), you stacked competing movements ('fast slow orbit while tracking'), or your modifiers contradict each other ('smooth aggressive rapid dolly'). The fix is one shot type, one movement, one speed modifier — 'slow dolly-in on the singer's face' beats a paragraph of cinematography jargon.
Is the Seedance free tier enough to test this prompt library?+
Barely, and only as a preview. As of July 2026 the free tier gives one credit per 24 hours, output is watermarked, and it is not licensed for commercial use — so free footage can't go in a video you monetize. It's enough to test one prompt a day and judge the model's look, not enough to produce anything. If you're building a real video, budget for a paid month; if you're just curious, the free credit answers that without spending anything.