AI Music Video

The full music video workflow: a frozen character bible, shots mapped to song sections, and a style spine that makes clips cut together.

Build the video for a song in three locked blocks + a shot list. Video models have no memory between clips: consistency comes entirely from repeating these blocks verbatim.

CHARACTER BIBLE (write once, paste word-for-word into EVERY shot prompt: any paraphrase causes drift):
[40-60 WORDS, EXHAUSTIVE: "A woman in her late 20s with waist-length black braids, warm brown skin, gold hoop earrings, wearing a cropped white tank, oversized vintage denim jacket, and silver rings on both hands", include hair, skin, 2-3 clothing items, one distinctive accessory]

STYLE SPINE (same rule: identical in every prompt):
[FILM LOOK + GRADE + ERA: "shot on 35mm film, teal-and-orange grade, subtle grain and halation, anamorphic lens flares, music-video contrast"]

SHOT LIST: map one 8-second clip per song section:
| Section | Shot | Energy |
| Verse 1 | [INTIMATE: "medium close-up, she walks toward camera down an empty neon-lit street at night"] | restrained |
| Chorus 1 | [BIG: "wide shot, she dances on the rooftop against the skyline, camera slowly orbiting"] | full |
| Verse 2 | [NEW LOCATION, SAME INTIMACY] | restrained |
| Bridge | [THE TURN: "slow motion, rain starting, extreme close-up on her face looking up"] | stripped |
| Final chorus | [BIGGEST: return to the chorus location, wider, more motion] | peak |

EACH SHOT PROMPT = [SHOT TYPE + ACTION from the list] + [CHARACTER BIBLE verbatim] + [LOCATION] + [STYLE SPINE verbatim] + "no subtitles, no on-screen text, no captions."

PERFORMANCE SHOTS (for lip-sync sections): She sings with full emotional performance, directly to camera [or in profile], her mouth clearly visible. Generate muted: sync the actual track and lip-sync in post with a dedicated tool.

Rules: one location + one camera move + one action per clip. Change one variable between shots of the same section. If your tool accepts reference images, generate 1-3 stills of the character first and attach them to every clip: image anchors beat text anchors.

How to use

The full pipeline: lyrics → Suno for the track → this prompt structure for the clips (Veo 3.1's ingredients-to-video with character reference images is the strongest option) → lip-sync tool for performance shots → edit cuts on the beat. The verbatim rule is everything: creators who paraphrase the character block get a different person every clip. Budget roughly 2-3 generations per shot; verses forgive drift, choruses and close-ups don't.

Originated fromStan SedberryUpdated
Cinematicadvanced

More video prompts

[Attach your first frame image, and the last frame image if your tool supports it, then use this prompt:]

Animate from the first image to the last image.

What changes between them: [THE ONE TRANSFORMATION, e.g. "the closed box opens and the product rises out of it" / "the empty street fills with morning light" / "the character turns fro

First and Last Frame Control

Direct a video model between two fixed images so the motion, timing, and transition land exactly where you need them for an edit.

Videointermediate
Generate a sequence of shots featuring the same character.

**Step 1: write the character lock.** This exact text goes into every shot prompt, word for word, before anything else. Never paraphrase it.

CHARACTER LOCK:
"[NAME], a [AGE] year old [DESCRIPTION: build, height impression], with [HAIR: exact length, color, texture, and how it is

Multi-Shot Character Consistency

Keep one character looking like the same person across separately generated shots using a locked description block and fixed style spine.

Videoadvanced
Write and shot-list an explainer video.

What we are explaining: [PRODUCT, FEATURE, OR CONCEPT]
Who is watching and what they already know: [AUDIENCE + BASELINE]
The one thing they should understand afterward: [THE SINGLE TAKEAWAY]
What they should do next: [THE ACTION]
Length: [60 / 90 / 120 seconds]
Visual style: [SCREEN RECORDING / MOT

Explainer Video

Turn a product or concept into a 60 to 90 second explainer with a script timed to shots, so the visuals carry what the words cannot.

Videointermediate

Search prompts

Find a prompt by title, description, tag, or category.