How to Prompt Seedance 2.5 Video Model
Learn how to prompt Seedance 2.5 for text-to-video, references, camera motion, style, audio, keyframes, storyboards, video editing, and extensions.
Seedance 2.5 gives a lot of control, but only if asked for that control clearly.
You can tell it which image defines a character, which video supplies the motion, when a shot should change, how the camera should move, what the dialogue should sound like, what must stay consistent, and even which parts of an existing video you want changed.
The problem is that all of these jobs need slightly different instructions.
This guide presents a practical way to write Seedance 2.5 prompts without turning them into walls of text. It explains how to choose the right workflow, structure the prompt, map references, control timing and camera movement, and handle editing, keyframes, storyboards, audio, and extensions with real examples along the way.
Introduction
A prompt does not need to describe everything in the final video.
If the character is already defined by an image, that image can carry the appearance. If a reference video already contains the movement you want, use it for the movement. If neither exists, those details have to come from the text.
The same applies to camera movement, timing, audio, environments, keyframes, and existing footage. The prompt becomes useful when it fills in the information that the inputs do not already provide, and makes the relationships between those inputs clear.
So before writing the prompt itself, the first question is:
What are you giving Seedance, and what do you need Seedance to do with it?
That leads directly into the next section: choosing the right workflow.
1. Text-to-Video: Start From the Prompt Alone
Text-to-video is the simplest Seedance 2.5 workflow. There are no reference images, motion videos, keyframes, or audio files carrying information for you. The video is built from the prompt alone.
A useful starting point is:
Subject + location + event + genre/style + camera idea.
For example:
Realistic nature documentary style, natural lighting and shadows. On a warm afternoon, on a grassy slope in the forest, a chubby panda cub rolls down the hill.
That sentence already establishes:
- Subject: panda cub
- Location: grassy forest slope
- Event: rolling downhill
- Style: realistic nature documentary
- Lighting: natural afternoon light
For a very simple clip, a prompt can stay close to this level.
If the video needs a specific sequence of events, several shots, consistent characters, controlled effects, or synchronized audio, then the prompt needs more structure.
Expand the prompt only where you need control
A detailed text-to-video prompt can usually be organized into four layers:
1. Overall video idea
Establish the subject, setting, event, and visual direction.
2. Important details that apply throughout
Describe things that should remain stable across the video, such as:
- character appearance
- clothing
- environment
- lighting
- important objects
- overall visual treatment
Define these once instead of repeating them inside every shot.
3. Plot, shots, or timeline
For a simple video, describe the action normally.
For a longer sequence, break it into either:
Shot 1: ...
Shot 2: ...or:
0-4s: ...
4-8s: ...Both structures work. A shot can include its visuals, action, camera direction, dialogue, and sound.
We will go much deeper into shots and timestamps later.
4. Additional rules that apply across the video
This is where you can keep recurring instructions together, such as:
- consistent camera behaviouur
- atmosphere
- sound
- environmental conditions
- continuity
- behaviour of an important effect
These notes are useful when repeating the same instruction inside every shot would make the prompt harder to read.
There is no required heading for any of these blocks. The headings are only there to keep complicated prompts organized.
Example: A 20-second sequence made entirely from text
The idea is straightforward: one martial artist stands beside a waterfall and gradually takes control of the surrounding water.
The challenge is keeping that idea coherent across three different shots while the character, environment, water behaviour, camera, and sound all remain connected.
Establish the overall video first
The prompt starts with:
Create a 20-second photorealistic cinematic fantasy sequence in 16:9 widescreen. Text-to-video mode.
The video contains exactly three clearly separated cinematic shots connected by clean hard cuts. No dissolves, morphing transitions or continuous transformation between camera angles. Do not display the shot numbers or timecodes.This gives the sequence a clear shape before any detailed action is introduced:
- photorealistic cinematic fantasy
- three shots
- clean hard cuts
- no transition effects between the camera setups
Duration and aspect ratio can also be set through generation settings, so they do not need to be repeated in the text when those settings are already defined.
Define recurring information once
The same martial artist appears throughout the sequence, so his appearance is established once:
MASTER CHARACTER LOCK:
Show one original adult male martial artist with a shaved head, calm angular facial features and a lean athletic build. He wears a weathered charcoal-grey traditional training robe, loose grey trousers, dark red ankle wraps and bare feet.
His face, age, body proportions, shaved head and complete outfit remain identical in all three shots. He is the only person in the scene. He never duplicates or changes appearance.The location is handled in the same way:
MASTER ENVIRONMENT LOCK:
The entire sequence takes place beside one enormous natural waterfall deep inside a secluded jungle ravine. The martial artist stands on a broad, flat, wet stone platform surrounded by a shallow turquoise pool.
Ancient roots twist around dark moss-covered rocks. Fine waterfall mist hangs in the air. Soft daylight enters from above and creates subtle shafts of light through the mist. Keep the waterfall, stone platform, pool, roots and lighting direction consistent across all shots.MASTER CHARACTER LOCK and MASTER ENVIRONMENT LOCK are just organizational headings. They are not special Seedance commands.
The useful principle is:
If the same information applies throughout the video, define it once instead of repeating it inside every shot.
Break a longer sequence into clear shots
The first shot establishes the character and mood:
SHOT ONE — CALM BEFORE THE MOVEMENT:
Tight cinematic close-up from the martial artist’s chest to the top of his head.
His eyes are closed. He holds both hands vertically together in front of his chest in a simple meditation pose...
The camera slowly pushes closer. Several tiny water droplets begin floating upward around his hands...
Near the end of the shot, he slowly opens his eyes and looks forward with complete concentration.The second shot introduces the main interaction:
SHOT TWO — WATER AWAKENS:
Wide full-body shot showing the martial artist centred on the wet stone platform with the complete waterfall behind him.
He smoothly lowers into a deep, stable martial-arts stance...
The surrounding water responds directly to each movement.The final shot develops the same idea into the payoff:
SHOT THREE — THE GREAT WATER RING:
Medium-wide heroic angle with the complete martial artist visible from head to bare feet.
As he rotates his forearms, the separate water streams connect into one enormous circular ribbon rotating vertically around and behind him.The progression is easy to read:
setup → development → payoff
Each shot has one clear job instead of trying to make the entire sequence happen at once.
We will get into shot timing, cuts, and camera language in detail later.
Give unusual behaviour its own rules
The hardest part of this video is the water.
It has to react to the martial artist, remain physically believable, form one specific shape, and avoid turning into generic magical energy.
Those rules are grouped together:
WATER PHYSICS LOCK:
All water originates visibly from the natural pool and waterfall.
Every water movement must be caused by a corresponding hand, arm or body movement from the martial artist.
Use one coherent main water ring.
Water must remain transparent, physically heavy and naturally illuminated, with believable refraction, spray, mist, surface tension and gravity.WATER PHYSICS LOCK is also only a heading.
The technique is useful whenever one part of the video has its own logic.
That could be:
- a transformation
- liquid behavior
- a product changing state
- an object reacting to a person
- damage that must persist
- a recurring visual effect
Keep that logic together rather than scattering the same rule across several shots.
Also notice that the detailed body choreography here has a reason. The water is supposed to react directly to the martial artist's movements.
For ordinary actions, you usually do not need to describe every hand and foot movement. Save that level of precision for moments where the exact motion affects the result.
Keep recurring visual and audio direction outside the shots
The visual treatment applies to the entire sequence:
CAMERA AND VISUAL STYLE:
Extremely photorealistic live-action cinematic footage. Real human skin texture, individual eyelashes, natural hands, wet woven fabric, realistic stone, moss, roots and physically convincing water.
Cool grey and deep green colour palette with soft natural daylight, volumetric waterfall mist, deep shadows, subtle film grain and cinematic depth of field.The soundtrack is handled separately:
AUDIO:
Natural waterfall ambience, flowing water, robe movement, bare feet shifting across wet stone and deep realistic water whooshes synchronized with the martial-arts movements.
Add a restrained atmospheric cinematic drone that gradually builds with the water ring and ends with one deep impact as the final wave releases.
No dialogue, chanting or narration.This keeps instructions that apply everywhere out of the individual shot descriptions.
Use negative instructions selectively
The prompt ends with a long continuity and exclusion block:
STRICT CONTINUITY:
Exactly one martial artist, one consistent outfit, one waterfall environment and one coherent water ring. No opponents, fighting, character duplication, changing faces, changing clothing...A long negative list is not something every Seedance prompt needs.
It is usually better to clearly describe what you want first.
Add exclusions when they protect something important or prevent a specific unwanted result. No subtitles, No BGM, or preventing a second character from appearing are useful when those constraints actually matter.
Do not automatically paste a giant negative block into every prompt.
What to take from this example
The useful structure underneath the custom headings is:
CORE VIDEO IDEA
↓
RECURRING DETAILS
↓
SHOT PROGRESSION
↓
SPECIAL BEHAVIOR RULES
↓
GLOBAL VISUAL + AUDIO DIRECTION
↓
ONLY THE CONTINUITY CONSTRAINTS THAT MATTERFor a simple video, you might only need:
CORE IDEA
+
ACTION
+
CAMERAFor a controlled multi-shot sequence, the larger structure makes the instructions easier to follow.
2. Using Reference Assets
Once you attach images, videos, or audio, the prompt no longer has to describe everything from scratch.
The important part is telling Seedance what each asset is supposed to control.
References are identified by their upload order, such as Image 1, Video 1, or Audio 1. Use those names directly when assigning their role.
Give every reference a clear job
A simple setup might be:
Image 1 defines the character's appearance.
Image 2 defines the environment.If you only want part of an asset, narrow its role:
Image 1 defines the character's identity and hairstyle.
Image 2 defines the wardrobe only.The same applies when images and videos are mixed:
Image 1 controls the character's identity.
Image 2 controls the outfit.
Video 1 controls the body movement only.This becomes especially useful when the references disagree with each other.
If Video 1 contains the motion you want but has the wrong actor, clothes, and location, say so:
Video 1 controls body movement only.
Do not use the person, clothing, or environment from Video 1.A single asset can also control several properties when they already belong together:
Image 1 defines the character's face, hairstyle, skin tone, body proportions, and outfit.The goal is not to give every visible detail a separate reference.
It is to make the role of every reference clear.
Example: three images with three different jobs
This video uses three images:
| Asset | Role |
|---|---|
| Image 1 | Storyboard and sequence |
| Image 2 | Polyphemus |
| Image 3 | POV guard |
The prompt starts by connecting the storyboard:
follow the storyboard @Image1Polyphemus is then tied to another image:
POLYPHEMUS LOCK @Image2
Polyphemus remains visually identical throughout the entire film.
Exactly one large centered eye.
No changing proportions.And the first-person character is tied to the third:
POV GUARD @Image3
One ordinary young Greek night guard.
He is the only character whose POV we experience.The important part is the mapping:
Image 1 → sequence
Image 2 → giant
Image 3 → POV characterSeedance does not have to guess which image is responsible for which part of the video.
This matters more as the number of references increases.
Do not describe an accurate reference twice
If a reference already contains exactly what you want, let it carry that information.
Suppose Video 1 already has the exact action and camera movement you need. Instead of rewriting every movement:
First she raises her arm.
Then she turns around.
The camera begins moving around her...you can simply write:
Strictly follow the actions and camera movement in Video 1 and keep their sequence unchanged.Add more text only when you want to change, remove, or clarify something from the reference.
This keeps the prompt focused on information the model does not already have.
Multiple views and multiple subjects
Several images can represent the same subject.
For example:
Images 1-3 show the same character from the front, side, and three-quarter view.
Use all three as appearance references for the same character.This can help when the character needs to remain recognizable across changing angles.
When several different characters are involved, map them separately instead of relying on visual similarity:
Images 1-2 are Character A.
Images 3-4 are Character B.If audio is involved, the relationship can be mapped at the same time:
Images 1-2 are Character A and correspond to Audio 1.
Images 3-4 are Character B and correspond to Audio 2.For a small number of subjects, both single-view and multi-view references can work. With larger casts, simpler single-view references are generally easier to keep separated.
How many references can you use?
Seedance 2.5 supports a large number of assets in one request:
| Reference type | Limit |
|---|---|
| Images | 30 |
| Videos | 10 |
| Audio clips | 10 |
| Total assets | 50 |
Reference videos can have a combined duration of up to 30 seconds, and reference audio can also total up to 30 seconds.
The maximum is not necessarily the most stable setup.
For subject references, the more practical ranges are:
| Input | Generally more stable |
|---|---|
| Subjects referenced through images | 1-8 subjects |
| Subjects referenced through video or audio | 1-5 subjects |
| Subject video/audio reference length | 5-10 seconds |
Larger setups are possible, but the more subjects and relationships Seedance has to resolve, the more stability can drop.
So the rule is simple:
Use every reference that adds useful information. Do not add references just because you can.
3. Motion Reference: Use Another Video for Movement
A reference video can provide the movement instead of making you describe that movement from scratch.
The basic instruction is:
Refer to [the movement you need] from Video 1 to generate [the new scene], keeping the motion details consistent.The important part is specifying what movement to take from the video.
That might be:
- a character's body movement
- a running or walking pattern
- a dance
- a fight sequence
- a facial expression
- an object's movement
- the trajectory of an effect
- the timing and rhythm of an action
If the new video uses a different character, map the appearance and movement separately:
Image 1 defines the character.
Refer to the movement in Video 1 for the character in Image 1.The image answers who performs the action. The video answers how the action moves.
Reference the motion you actually need
You do not always need everything happening in the source video.
If only one action matters, name that action:
Refer to the running motion in Video 1...If the whole sequence matters:
Strictly follow the actions in Video 1 and keep their sequence consistent.This is especially important when the source contains other information you do not want to carry into the result.
For example, if Video 1 has the right movement but the wrong person and location:
Refer to the body movement in Video 1 only.The character and scene can then come from the prompt or other references.
Do not rewrite movement that is already clear in the video
If the reference already shows the exact motion you want, there is usually no reason to turn it back into a long written description.
Instead of manually describing:
The horse lifts its front legs, pushes forward, lands, extends its rear legs...the motion itself can remain in the reference video.
The prompt only needs to tell Seedance what part of that motion matters and what to do with it.
Example: transfer a running motion into a completely different result
Input
Output
The source video provides the horse's running form.
The prompt is:
Referencing the running shape of the horse in the video, generate a scene: a golden steed runs on the grassland, then freezes its magnificent running posture and turns into a horse-shaped gold pendant.The reference is not being used as the final scene.
It contributes the running motion, while the prompt changes the horse's appearance, environment, and final event.
This is what makes motion reference useful. The source video can provide a movement that would be tedious or difficult to describe precisely in text, while the rest of the video can be completely different.
Motion references do not have to look like the final video
A useful motion reference can be extremely rough.
For example, a tabletop setup made from toy cars and magnetic tiles can define:
- which objects move first
- their direction
- relative speed
- spacing
- collisions
- traffic flow
- the order of events
That movement can then be interpreted as a full-scale road scene with real vehicles and construction equipment.
The visual quality of the reference is not the point here. Its job is to communicate how things move through space and time.
This is useful when an action is difficult to describe but easy to demonstrate.
Motion and appearance can stay separate
A reference video may contain a person whose appearance you do not want.
That does not stop the video from being useful.
For example:
Image 1 defines the character's appearance.
Video 1 provides the character's movement only.The same idea works for objects.
A rough object, toy, or placeholder can provide motion while another image defines what the final object actually looks like.
This separation becomes even more useful when building scenes from several references.
When the reference contains more than motion
A video also carries camera movement, environment, lighting, style, and other information.
Do not assume Seedance should inherit all of it.
If you only want the action, say so.
If you want both the action and camera behavior, both can be referenced:
Refer to the character movements and shot language in Video 1...Those are two separate pieces of information coming from the same asset.
4. Camera-Motion Reference
Use a reference video when you want Seedance to follow an existing camera movement while generating a different scene.
For example, the scene can come from an image while the camera path comes from a video:
Image 1 → scene and visual center
Video 1 → camera movementA real prompt can be as direct as:
Referring to the camera movement in Video 1, create a concept video for a science and technology park, with the tall building in Image 1 as the visual center, also using a first-person diving perspective to reflect the sense of technology in the park.Inputs:
Camera Motion Reference Video
Reference Image
Output:
Here, Image 1 controls what the scene looks like, while Video 1 controls how the camera moves through it.
If only the camera movement is useful from the reference video, keep its role narrow:
Refer to the camera movement in Video 1 only.If both subject motion and camera motion should be followed, say that explicitly:
Refer to the subject movement and camera movement in Video 1.The important distinction is simple:
- Motion reference transfers how the subject moves.
- Camera-motion reference transfers how the camera moves.
5. Control Style of your video
Style can be controlled in two ways:
- Describe the look in the prompt
- Provide an image or video that already has the look you want
Both can work. The choice depends on whether the style is easy to describe or easier to show.
5.1 Style with text
If the visual treatment is clear enough to describe, no style reference is necessary.
A style can be named directly:
Photoreal live action
Stop-motion clay
2D cel animation
Low-poly 3D
Black-and-white mangaBut a style name alone leaves a lot open to interpretation. Add a few characteristics that define the version of that style you actually want.
Instead of:
Stop-motion clay.use something closer to:
Stop-motion clay with visible fingerprints, tactile miniature textures, and subtle frame-by-frame movement.The same applies to other looks:
Bright 2D cel animation with bold outlines, simple graphic shadows, and fluid expressive movement.Late-1990s low-poly 3D with chunky polygons, pixelated textures, simple vertex lighting, and slightly stiff retro-game movement.The pattern is simple:
STYLE
+
THE VISUAL TRAITS THAT DEFINE ITExample: changing style while the story continues
This sequence keeps one chef, one kitchen, and one cooking process while the rendering changes five times.
The important instruction appears near the beginning:
Keep the same chef, same kitchen layout, same cookware, same ingredients, and continuous cooking progress throughout the entire video.
The only major change is the visual rendering style.Then each section defines its own treatment:
0–3s — PHOTOREAL LIVE ACTION
3–6s — STOP-MOTION CLAY
6–9s — BRIGHT 2D CEL ANIMATION
9–12s — PS1 LOW-POLY 3D
12–15s — BLACK-AND-WHITE MANGAThe cooking does not restart when the style changes. The physical state continues from the previous moment:
Each new style must inherit the exact physical state of the previous moment.
The cooking must progress continuously and logically from raw pasta to finished dish.That distinction matters. Style is changing. The scene logic is not.
Full prompt
Create a 15-second cinematic mixed-media cooking sequence. One chef prepares one plate of spaghetti from beginning to end inside the same small home kitchen.
Keep the same chef, same kitchen layout, same cookware, same ingredients, and continuous cooking progress throughout the entire video. The only major change is the visual rendering style. The video moves sequentially through exactly five visual styles.
Never show multiple styles at the same time.
0–3s — PHOTOREAL LIVE ACTION
A realistic chef drops dry spaghetti into a pot of boiling water. Steam rises naturally. The chef stirs once with tongs. Warm practical kitchen lighting, realistic textures and motion.
3–6s — STOP-MOTION CLAY
Without changing camera position, composition, chef identity, or object placement, the entire scene smoothly becomes handmade clay stop-motion. The spaghetti is now partially cooked. The chef pours tomato sauce into a pan and stirs it. Visible clay fingerprints, tactile miniature textures, subtle stop-motion movement.
6–9s — BRIGHT 2D CEL ANIMATION
The same scene transforms into clean colorful hand-drawn cel animation. The chef lifts the cooked noodles from the pot and drops them into the tomato sauce. Sauce splashes slightly as the noodles land. Bold outlines, simple graphic shadows, fluid expressive animation.
9–12s — PS1 LOW-POLY 3D
The whole kitchen becomes late-1990s low-poly game graphics while preserving the exact scene layout. The chef tosses the spaghetti inside the pan. Chunky polygons, pixelated textures, simple vertex lighting, slightly stiff retro-game movement.
12–15s — BLACK-AND-WHITE MANGA
The scene transforms into a detailed black-and-white manga illustration in motion. The chef twists the finished spaghetti onto a plate, adds grated cheese and one basil leaf, then places the finished dish toward the camera.
Dynamic ink strokes, cross-hatching, subtle speed lines only during the final plating action. End clearly on the completed spaghetti dish.
Each new style must inherit the exact physical state of the previous moment. The cooking must progress continuously and logically from raw pasta to finished dish.
One chef only. One kitchen only. One spaghetti dish only. No duplicated ingredients, no reset of cooking progress, no changing kitchen layout, no character replacement, no mixed styles within the same shot, no text, no subtitles.The important lesson is not that style changes need timestamps. Timing is covered later.
What matters here is that every style is defined clearly enough to produce a visibly different rendering while continuity instructions protect everything that should remain unchanged.
5.2 Style with a reference image or video
Sometimes the target look is easier to show than describe.
In that case, give the style asset its own role:
Image 1 → subject
Image 2 → visual styleThe subject reference answers who or what should remain recognizable.
The style reference answers how that subject should be rendered.
Example: keep the subject, change the rendering style
Input:
Photoreal Reference Image
Animated Reference Image
Output:
The mapping starts simply:
Use Image 1 as the main subject reference and Image 2 as the visual style reference.Then separate what should remain from what should change.
For the subject:
Preserve her recognizable facial structure, hairstyle, bangs, outfit identity, proportions, and overall pose logic.For the style:
Use Image 2 specifically for the visual treatment: stylized rendering, illustration-like facial simplification, soft graphic shading, color treatment, texture treatment, and the overall aesthetic language.This is more useful than broadly saying:
Make Image 1 look like Image 2.The first version tells Seedance which properties belong to each reference.
A style reference does not have to control everything in the image
A style image also contains composition, pose, background, clothing, lighting, and other visual information.
If those should not transfer, narrow the instruction:
Use Image 2 for the visual style only. Keep the composition and subject identity independent from Image 2.Or narrow it further:
Use Image 2 for color treatment, shading, and texture only.The same principle works with a style video.
If its rendering treatment matters but its characters and actions do not, reference only the style.
Do not re-describe a style that the asset already shows clearly
Once the style reference is accurate, the prompt does not need to reproduce every visible detail in words.
Use text for the parts that still need clarification:
- which qualities should transfer
- which qualities should not transfer
- what must remain recognizable
- whether the style remains constant or changes during the video
Example prompt
Use Image 1 as the main subject reference and Image 2 as the visual style reference.
Create a short cinematic portrait video of the same young woman from Image 1. Preserve her recognizable facial structure, hairstyle, bangs, outfit identity, proportions, and overall pose logic. Keep the subject clearly identifiable throughout the video.
The video should begin in a photorealistic look close to Image 1, then smoothly transition into the stylized illustrated look of Image 2 while preserving the same subject and similar composition.
Use Image 2 specifically for the visual treatment: stylized rendering, illustration-like facial simplification, soft graphic shading, color treatment, texture treatment, and the overall aesthetic language. Do not replace the subject's identity with a different person.
Keep the framing close and intimate, with a portrait-focused composition and subtle motion only: gentle head movement, natural blinking, slight hair movement, and a soft camera drift or slow push-in. The transition from photoreal to stylized should feel visually smooth and intentional, not like a cut to a different person.
Maintain a clean background and consistent subject placement. Do not introduce extra characters, costume changes, text, subtitles, logos, or unrelated scene changes.
The main goal is to show that the same subject from Image 1 is being re-rendered in the style of Image 2.Choosing between text and a style reference
| If you have... | Use... |
|---|---|
| A style that is easy to describe | Text styling |
| A specific visual look you want to follow | Style reference |
| A reference look that needs adjustment | Style reference + text clarification |
The same rule from reference mapping still applies: give the style source a clear role, and let the other references control the things that should not change.
6 Audio: Dialogue, Voices, Music and Sound Design
Audio can be written directly into the same prompt as the visuals.
A video may need dialogue, music, environmental sound, sound effects, or some combination of them. Describe only the parts that matter for the scene.
If a sound belongs to one specific moment, place it near that moment:
0-4s:
The apartment door slams shut.
SFX: heavy wooden door impact.
4-7s:
Maya turns toward him.
Maya, quietly: "You came back."If a sound should continue across several shots, describe it once:
AUDIO:
Constant rain outside the apartment.
Low indoor room tone.
No BGM.This keeps moment-specific sounds attached to their actions without repeating recurring ambience throughout the prompt.
6.1 Dialogue and Voices
For dialogue, three things should be clear:
WHO SPEAKS
+
WHAT THEY SAY
+
HOW THEY SOUNDHere is a multi-character sequence where each character has a defined voice and a specific line.
[AUDIO]
English lip-sync.
Maya = young mezzo, alarm → wonder → joy
Riff = relaxed alto
Luma = airy synthetic soprano
Moss-7 = gentle radio baritone
Pip = bright fast high voice
Milo = dry sleepy low voice
Aurelia = elegant metallic alto
... uplifting score ducks under speech.The dialogue is then assigned to individual speakers:
[TEXT / DIALOGUE]
2.3s Maya “Whoa!”
4.2s Maya “Where am I?”
7.0s Riff “Welcome, rider.”
11.0s Luma “The gate heard you.”
14.7s Moss-7 “Let it grow.”
16.7s Pip “Mail for you!”
19.2s Milo “Wonder suits you.”
23.0s Aurelia “Choose your first thread.”
27.0s Maya “Then let's ride it.”
Only the speaker's mouth moves.Make the speaker clear
Instead of:
"Where am I?"write:
Maya: "Where am I?"The same dialogue can also be placed directly inside a shot:
6-9s:
Maya looks toward the open doorway.
Maya, nervous and speaking quietly:
"Did you hear that?"When several characters are present, explicitly attaching the line to one character removes ambiguity about who should speak.
Separate the words from the voice
These instructions do different jobs:
Maya: "Where am I?"defines the spoken line.
Maya = young mezzo, alarm → wonder → joydefines the intended voice and delivery.
Voice descriptions can specify only the characteristics that matter for the scene:
low calm male voicebright fast high voicesoft breathy voice, speaking slowlyolder male voice,
slightly rough timbre,
restrained emotion,
deliberate deliveryUseful properties include pitch, age impression, energy, emotion, pace, accent, tone, timbre, and delivery.
Voice delivery can change during the scene
The voice does not have to remain emotionally identical for the entire video.
For example:
Maya = young mezzo, alarm → wonder → joyOr the change can be attached directly to the performance:
Maya, nervous and speaking quickly:
"Where are we?"
She looks around and relaxes.
Maya, now speaking softly with wonder:
"This place is incredible."Keep speech attached to the correct mouth
For a shot containing several visible characters, the prompt can explicitly say:
Only the speaker's mouth moves.This makes the intended speaking behavior clear rather than leaving the mouth activity of the other characters unspecified.
Using a reference voice
If a reference audio clip already contains the voice you want, assign that audio to the character.
For example:
Image 1 depicts the protagonist John and uses the voice timbre from Audio 1.If you only want the timbre:
Use Audio 1 for John's voice timbre.For several characters:
Images 1-2 are Character 1 and correspond to Audio 1.
Images 3-4 are Character 2 and correspond to Audio 2.The reference mapping itself works the same way as the reference mapping covered earlier. Here, the important part is defining what the audio controls.
A reference audio asset can be used for audio information such as voice, dialogue, tone, timbre, music, or melody.
A real multi-character prompt uses this directly:
Inputs
Character:
Scene:
Output
The voice of the SUPERMODEL ... is @ Audio 1The useful distinction is:
visual reference → what the character looks like
audio reference → what the character sounds likeIf only the voice should transfer, say that instead of leaving the role of the audio reference broad.
6.2 Music, Ambience and Sound Effects
The non-dialogue soundtrack usually comes from three kinds of sound:
Music / BGM
Environment / Ambience
Action Sounds / Foley / SFXA scene may use all three or only one or two of them.
Music
Instead of giving the music only a broad label:
Cinematic music.Describe what it should do during the video:
A restrained electronic score begins quietly.
As the city is revealed, introduce a low rhythmic pulse.
The music gradually builds in energy.
During dialogue, the score ducks underneath the voices.
End on one sustained synth note.A useful structure is:
MUSIC STYLE
+
ENERGY
+
HOW IT CHANGES
+
HOW IT RELATES TO THE SCENEFor example:
uplifting score ducks under speechHere, the important instruction is not only the type of music. It also tells the soundtrack to reduce its prominence while dialogue is being spoken.
If an uploaded audio asset already contains the music or melody you want to reference, assign that role to the audio asset instead.
Ambience and synchronized sound effects
Environmental sound describes what the location naturally sounds like.
Sound effects can then be tied to particular visual events.
AUDIO:
low sustained wind across open ice,
deep groaning ice under pressure,
snow ticking on fabric,
one long sub-bass swell as the eyes open,
a rising low resonant hum as the gold spreads,
and the deep grinding of enormous mass
leaving the ground at the end.
No voice, no dialogue, no narration, no music.Some of these sounds belong continuously to the environment:
low sustained wind across open icesnow ticking on fabricOthers are triggered by specific visual events:
one long sub-bass swell as the eyes opena rising low resonant hum as the gold spreadsdeep grinding of enormous mass leaving the groundWhen an important action should produce a particular sound, connect the two directly:
As the steel door locks,
a heavy metallic clunk echoes through the corridor.When the spacecraft passes the camera,
a violent pressure burst and deep rushing whoosh hit at the pass-by moment.As the glass hits the table,
a short sharp glass impact is heard.The same idea can stay very simple for a natural scene:
Natural environmental audio only:
wind, rustling grass,
and the soft plop of the panda rolling.The amount of audio detail should match the scene. A quiet natural shot may need only a few environmental sounds, while an action-heavy shot may need several event-linked effects.
6.3 Controlling What You Do Not Want to Hear
Seedance can also be told what kind of audio should be absent.
For example:
No BGM; generate only environmental sounds and action sounds.For a silent result:
No audio.A useful version for a scene that should contain sound but no music is:
No BGM.
Generate only environmental sounds and action sounds:
wind, footsteps, clothing movement,
distant traffic and door impacts.This specifies both:
WHAT SHOULD NOT BE THERE
+
WHAT SHOULD BE THERE INSTEADThere is an important practical limitation.
More than 30 generations were tested with variations including:
No music
No melody
Musicless
Ambience only
SFX only
No music under any circumstancesand:
[SOUND]
Strictly naturally occurring sounds and Foley only.
No music allowed.Music still appeared in some results.
So:
No BGMshould be treated as a prompting instruction, not as a guaranteed mute switch.
Giving Seedance a positive replacement soundtrack such as ambience and action sounds makes the intended result clearer, but it does not turn the instruction into a hard technical constraint.
If a completely music-free result is required, the generated output should be checked rather than assuming the prompt was obeyed.
6.4 Changing Audio in an Existing Video
Audio can also be modified after the video already exists.
This can include changing dialogue, vocals, music, or sound effects while preserving the rest of the video.
A narrow dialogue edit can look like:
Only edit the man's dialogue in Video 1:
Change it to:
"Don't come over here."
Adjust the accent to an American English accent.The requested change is specific:
existing dialogue
→
new dialogue + new accentwithout asking for unrelated parts of the video to change.
Translating existing dialogue and matching the lips
A more advanced audio edit is translation.
Input - English Video
Output - Chinese Video
Translate the spoken dialogue in the video into Chinese,
with no subtitles.
Precisely adjust the lip movements to match the translated speech,
while keeping everything else unchanged.This instruction contains three separate requirements:
TRANSLATE THE SPEECH
+
MATCH THE LIPS TO THE NEW SPEECH
+
PRESERVE EVERYTHING ELSEThis is different from generating a new scene with translated dialogue.
The video already exists. The task is to change the spoken language and the corresponding lip movement without unnecessarily changing the rest of the shot.
The same structure can be used for another dialogue replacement:
Only change the woman's dialogue in Video 1.
Replace it with:
"I thought you'd never come back."
Keep her existing voice character.
Adjust her lip movements to match the new dialogue.
Leave the visuals, camera movement,
environment, sound effects and music unchanged.For this kind of edit, the prompt can be thought of as:
WHAT AUDIO CHANGES
+
WHAT IT BECOMES
+
WHAT MUST RESYNC
+
WHAT MUST REMAIN UNCHANGEDThat keeps the request focused on the audio change instead of reopening the whole video.
A simple way to think about audio
Before finishing the prompt, ask:
Who speaks?
How should they sound?
What should the music do?
What does the environment sound like?
Which important actions need matching sound effects?
Is there anything I specifically do not want to hear?Only answer the questions that matter for that particular video.
A conversation may need voices, quiet room ambience, and restrained music.
An action sequence may need almost no dialogue but several synchronized effects.
Tell Seedance what the scene should sound like with the same clarity used to describe what happens on screen.
7. Frame Control
Sometimes you do not need to describe the whole video from zero.
You may already know how it should begin, how it should end, or a few important visual states it should pass through.
That is where frame control helps.
Seedance gives you three useful ways to do this:
| If you already know... | Use... |
|---|---|
| exactly how the video should begin | First Frame |
| exactly how it should begin and end | First + Last Frame |
| several important visual states in between | Multiple Keyframes |
7.1 First Frame
Use First Frame when the opening image matters most.
The first frame gives Seedance a clear starting point. It already defines the subject, framing, outfit, lighting, environment, and mood. So the prompt does not need to waste words re-explaining all of that. It should mainly describe what happens next.
Input
Output
A simple way to write it is:
Image 1 is the first frame.
A young woman walks forward along the wet sidewalk at night.
Keep the same outfit, same street, same lighting mood, and same camera direction.
The camera slowly tracks backward as she walks toward it.
Her hair moves slightly in the night air.
Storefront reflections and streetlights shimmer on the pavement.The core idea is simple:
first frame
+
what happens after itIf the platform supports assigning the image as the actual first frame, that gives stronger control than just mentioning it in the text. It also means the output ratio follows the aspect ratio of that first-frame image.
Use this mode when the opening composition needs to be right.
7.2 First + Last Frame
Use First + Last Frame when you already know both endpoints and want Seedance to build the motion between them.
start frame
→
progression
→
end frameInput
First Frame:
Last Frame:
Output
This is useful when the beginning and ending states matter more than every intermediate detail.
A good prompt in this setup focuses less on re-describing the two images and more on describing:
- what changes
- how it changes
- how fast it changes
- what should stay consistent while the scene develops
For example:
[MODE]
First & Last Frame
[CAMERA]
Fixed tripod, zero drift.
The scene begins from the first frame and ends at the last frame.
Show the construction progressing naturally between the two states.
Workers build the pool area step by step.
Materials appear in believable stages.
The decking, stone edges, water, and furniture are added progressively.
Finish exactly on the completed final scene.The important habit here is this:
If the two endpoint images already show the beginning and the end, the prompt should spend most of its energy on the middle.
Also keep the two frame images compatible. They should use the same aspect ratio whenever possible, since the output follows the first frame.
7.3 Multiple Keyframes
Use Multiple Keyframes when two endpoints are not enough.
Sometimes you know more than the start and end. You may also know a few important visual states that should happen in order.
That is what keyframes are for.
A → B → C → DInstead of giving Seedance only the first and last destination, give it several checkpoints.
The basic structure is:
Use Images 1 to 3 in order as keyframes.Then describe the motion that connects them.
For example:
Use Images 1 to 3 in order as keyframes.
Begin in the calm first state.
Progress naturally into the second state.
Then continue into the final third state.
Keep the subject consistent.
Make the transitions smooth and visually coherent.
Let the motion between each state feel continuous rather than abrupt.The point of keyframes is not to describe every frame of the video.
The point is to define the important states the video should pass through, while letting Seedance invent the motion between them.
So the difference is:
First + Last Frame
A ---------------- BMultiple Keyframes
A -------- B -------- CUse First + Last Frame when the main goal is the journey between two fixed endpoints.
Use Multiple Keyframes when the video must pass through several defined stages in order.
8. Storyboard Control
A storyboard is useful when you already know the sequence of shots and events, but do not need every panel to become an exact frame.
It can guide things like:
shot structure
composition
actions
camera rhythm
story progressionThis is different from the keyframes covered above.
Storyboard → guides the sequence
Keyframes → define specific visual states more strictlyIf the video needs to pass through exact visual checkpoints, use keyframes. If the main goal is to show Seedance how the sequence should unfold, use a storyboard.
Give the storyboard a clear role
A storyboard can be combined with separate references for characters, environments, or other elements.
For example:
Image 1 → storyboard and overall sequence
Image 2 → Cyclops appearance
Image 3 → POV guardInputs
Storyboard ( Image 1 ) :
Cyclops ( Image 2 ):
Guard ( Image 3 ):
Output
The prompt begins by assigning the storyboard:
follow the storyboard @Image1Then other references handle details that should remain independent from the storyboard.
For the Cyclops:
NEVER show 2nd eye of POLYPHEMUS
(he only has one eye on forehead)
@Image2For the POV character:
Everything the audience sees is exactly what the young guard sees.
The audience is the guard.
@Image3This is the useful pattern:
STORYBOARD
→ sequence and shot progression
OTHER REFERENCES
→ character / environment / subject detailsThe storyboard does not need to control everything.
Fill in what the storyboard cannot show clearly
A storyboard already contains a lot of visual information, so there is no need to describe every panel again.
Instead, use the prompt to add details that are difficult to communicate through the storyboard alone.
In this example, the storyboard establishes the sequence, while the prompt adds a strict camera rule:
The entire film is ONE CONTINUOUS FIRST-PERSON POV shot
from the exact same young Greek guard.
NEVER switch to third person.
NEVER show an over-the-shoulder shot.
NEVER show the guard's face.It also defines the camera itself:
Use a clean 0.5× ultra-wide 13mm-equivalent
first-person camera at natural human eye level.Then the timeline fills in the actions, dialogue, reactions, sound, and continuity needed to turn the storyboard into a complete 30-second sequence.
That is usually the right division of work:
WHAT THE STORYBOARD ALREADY SHOWS
→ let the storyboard carry it
WHAT THE STORYBOARD CANNOT SHOW PRECISELY
→ add it in the promptA storyboard is not a strict frame sequence
Storyboard panels guide the generated sequence, but Seedance does not have to reproduce every panel exactly frame for frame.
It can interpret the movement and transitions between those panels.
A storyboard is not the right choice when exact intermediate compositions are critical.
For stricter visual checkpoints:
Use Images 1 to 6 in order as keyframes.For broader shot progression:
Use Image 1 as the storyboard reference
for the overall shot structure and sequence.Keep the storyboard easy to read
For multi-panel storyboards, keep the board relatively simple.
A good target is 15 panels or fewer.
Simple line-art or stick-figure storyboards can work well because the important information is easy to read:
subject position
+
composition
+
action
+
shot progressionExamples of unsuitable multi-panel storyboards:
Not Recommended
Not Recommended
Not Recommended
Avoid filling the storyboard itself with large amounts of text or unnecessary visual detail.
The prompt is a better place for information such as dialogue, exact timing, continuity requirements, sound, or special camera instructions.
Also make sure the prompt and storyboard agree with each other. If the storyboard shows one camera direction while the text demands something incompatible, Seedance has to resolve conflicting instructions.
For simple concept storyboards, the prompt can stay simple
Not every storyboard needs a long 30-second production prompt.
If the board already communicates the idea clearly and the main goal is for Seedance to develop it into a coherent sequence, the instruction can be much shorter:
Construct the complete story according to the storyboard sequence.
Use the shots in a coherent order and create natural movement
between the storyboard moments.The amount of text should depend on what is missing from the storyboard.
A detailed production sequence may need camera rules, dialogue, continuity, and timing.
A clear concept storyboard may need only a short instruction explaining how Seedance should use it.
9. 3D Clay-Model Reference
Some shots are difficult to control with text alone.
If you already know the camera path, subject movement, blocking, or shot rhythm you want, build that structure first as a simple 3D animation and use it as a reference.
The basic workflow is:
3D PREVIS
→
camera + motion + blocking + timing
→
FINAL RENDERThere are two useful ways to do this: coarse 3D blockouts and fine-grained 3D references.
9.1 Coarse 3D Blockout
A coarse blockout does not need detailed models or finished materials.
Simple geometric forms can represent the people, objects, and environment while the video carries the information that actually matters:
camera movement
shot rhythm
shot-size changes
subject motion
movement paths
camera blocking
lighting changesInputs
Input Video:
Scene Images:
Output
Side By Side View:
The important part of the prompt is defining exactly what the 3D video should control.
For example:
Use the 3D clay-model reference video <Video 1>
as the only reference for the video's camera movement,
shot rhythm, shot-size changes,
subject motion trajectory, and camera blocking.
Strictly preserve the shot order,
camera position changes,
movement patterns, and pacing of Video 1.The visual appearance can come from separate references.
So the division of work can be:
Video 1
→ camera + blocking + motion + rhythm
Images
→ character / environment / visual appearance
Text
→ final scene details and requirementsThis is useful for scenes where getting the movement structure right is more difficult than describing what the characters look like.
Choose what the 3D reference is allowed to control
A 3D reference contains several types of information at the same time.
You may want all of them, or only some of them.
If you only need movement:
Refer to the camera movement and motion in Video 1.If the lighting animation matters too:
Refer to the lighting changes,
camera movement, and motion in Video 1.If the 3D video is only a blockout, explicitly prevent its unfinished appearance from transferring:
Use Video 1 only for camera movement,
shot rhythm, camera blocking,
and character animation.
Do not reference its visual content.That separation is important.
A grey blockout may be useful because of how everything moves, not because you want the final video to look like grey 3D geometry.
9.2 Fine-Grained 3D Re-rendering
A fine-grained 3D reference serves a different purpose.
Instead of using simple geometry mainly to control structure, provide a more complete 3D video and ask Seedance to re-render it into a different final world or visual treatment.
DETAILED 3D VIDEO
→
RE-RENDER
→
FINAL VISUAL RESULTInputs
Output
The prompt can be much simpler:
Render Video 1.
No BGM; generate only environmental sounds and action sounds.
Rendering requirements:
The background is a nighttime cyberpunk city
in deep blue and purple tones,
filled with dense skyscrapers.
Huge holographic billboards and neon lights glow
between the buildings.
Several flying vehicles move through the sky.
The character is a small raccoon
dressed in a black stealth suit,
moving cautiously across the rooftop.Here, the 3D video already provides much more of the underlying scene and motion.
The text mainly tells Seedance how that existing structure should be rendered.
Keep fine-grained 3D references clean
When using a more complete 3D reference, make sure Seedance can see the actual scene clearly.
Avoid unnecessary overlays such as:
trajectory lines
coordinate lines
camera cones
other viewport guidesThose are useful while working inside 3D software, but they do not need to become part of the reference video.
Coarse vs Fine 3D
| If you need... | Use... |
|---|---|
| Camera paths, blocking and difficult motion | Coarse 3D blockout |
| Shot rhythm and movement structure | Coarse 3D blockout |
| A more complete scene re-rendered into another look | Fine-grained 3D reference |
The main difference is what the 3D input is expected to contribute.
COARSE 3D
→ primarily controls how the video moves
FINE 3D
→ provides a more complete base for re-renderingIn both cases, the prompt should make clear which information from the 3D video is important and which parts should be replaced by the final visual treatment.
10. Video Editing
Editing is useful when the video already works and only part of it needs to change.
Instead of describing a new video, describe the difference between the current version and the version you want:
WHAT EXISTS
→
WHAT SHOULD CHANGE
+
WHAT SHOULD STAY THE SAMEThe more specific the edit scope is, the less of the existing video Seedance has to reinterpret. Seedance 2.5 supports adding, removing, or modifying visual elements, and partial edits can be limited to a particular time range when needed.
10.1 Editing with Text
If the change can be described clearly in words, edit the video without adding another visual reference.
A simple clothing edit can be as direct as:
Change the orange shirt worn by the standing woman in Video 1
to a dark navy work jacket.
Leave everything else in the shot completely unchanged:
same camera move, same lighting, same second character, same room.The structure is:
orange shirt
→
dark navy work jacket
everything else
→
preserveThe same approach works for a particular part of the environment:
Replace the view through all the windows in Video 1
with heavy grey storm clouds and rain over the same ridgeline.
Keep the interior of the cabin,
both characters,
the interior lighting,
and the camera movement unchanged.Here the edit is limited to the view outside the windows.
Several edits can be grouped into one controlled request
If several changes are needed, list exactly what should change instead of broadly asking Seedance to redesign the scene.
For example:
Make three changes to Video 1 and nothing else.
1. Remove the green two-way radio from the wall shelf.
Leave the shelf and wall unchanged.
2. Add a white enamel mug of coffee
with steam rising from it
onto the map table beside the brass instrument.
3. Change the seated man's navy fleece
to a dark red checked flannel shirt.
Keep his face, hair, and position the same.
Everything else in the shot stays unchanged.The pattern is:
CHANGE 1
CHANGE 2
CHANGE 3
+
PRESERVE EVERYTHING ELSEThis is more precise than saying:
Change a few things in the room.Limit an edit to part of the video when necessary
If something should change only during a specific section, attach the edit to that time range:
From 4-6 seconds,
change the man's action from drinking coffee
to mopping the floor.
Leave the rest of the video unchanged.That is enough here. The timing rules themselves do not change just because the task is editing.
10.2 Editing with Reference Images
Sometimes text can describe what should change, but not precisely what the replacement should look like.
That is when reference images help.
Inputs
Input Video
Dark-clothed
Environment
Brown Clothed
Output
In this example, the original video already contains the fight choreography and rhythm.
The edit prompt replaces the visual content around that motion:
Replace the two-person fight in Video 1
with an empty-handed probing exchange
before a cold-weapon duel.
Replace the scene with a medieval stone castle platform,
ancient courtyard, mountain-fortress platform,
or simple stone-brick duel arena.
Refer to Image 1 for the environment.Then the two fighters are replaced:
Replace the man in dark clothing in the video with Image 2.
Replace the man in light-colored clothing with Image 3.
Keep the original actions and rhythm unchanged.The result separates two jobs:
ORIGINAL VIDEO
→ actions + performance rhythm
REFERENCE IMAGES
→ new environment + new character appearancesThe existing movement does not need to be described again when it is already present in the video. The prompt can focus on the parts that are being replaced.
Preserve the parts that already work
For editing, phrases such as these are useful:
Keep the original actions and rhythm unchanged.Leave everything else in the shot unchanged.Only change the clothing.Preserve the camera movement, lighting, and character position.They define the boundary of the edit.
Before writing an editing prompt, ask two questions:
What exactly am I changing?
What exactly do I want Seedance to preserve?If those two answers are clear, the edit prompt usually does not need to be long.
11. Video Extension
Extension is useful when the current video already works and the story or action should continue beyond its existing ending.
The original clip becomes the starting point. The prompt mainly needs to describe what happens next.
EXISTING VIDEO
+
EXTENSION LENGTH
+
NEXT ACTIONScene 1 ( 10 sec )
Scene 2 ( 8 sec )
Scene 3 ( 10 sec )
Output
11.1 Continue From the Existing Ending
A basic extension prompt can be very direct:
Extend Video 1 by 5 seconds.
A bee flies in and lands on the flower.
The camera moves into a macro close-up as its legs
and abdomen collect golden pollen.
The bee takes off and the camera follows it
toward another flower.The important part is that the prompt describes the new continuation.
There is no need to rewrite everything that already exists in Video 1.
If the current clip already establishes the character, environment, camera style, and movement, those can continue from the source video.
For a continuation where the join itself matters, make that intent explicit:
Extend Video 1 forward by 10 seconds.
Continue directly from the final moment of Video 1.
Maintain the current movement and camera direction
through the extension point.
Continue the action naturally into the next event
without restarting the scene or repeating the previous action.Then describe the actual new event.
11.2 Keep the Join Seamless
A good extension should feel like the original video simply kept running.
The things that matter most around the join are:
subject continuity
movement continuity
camera continuity
environment and lighting
audio continuityOnly call out the ones that might otherwise change.
For example:
Continue from the final frame without resetting the camera.
The character keeps moving in the same direction
at the same pace.
Preserve the existing lighting and environment.
Continue the ambient sound naturally across the join.This is different from asking Seedance to create another independent shot. The new action should grow out of the state already present at the end of the input video.
Forward and Backward Extension
Extension can continue the video in either direction.
FORWARD
existing video → new endingor:
BACKWARD
new beginning → existing videoFor forward extension, describe what happens after the current ending.
For backward extension, describe what should happen before the existing beginning while making sure it can naturally lead into the original video.
A Few Extension-Specific Settings
Extension keeps the aspect ratio of the video being extended, so do not choose a new ratio for the continuation.
The extension duration can be chosen separately.
For better visual and audio continuity, use MOV for the input and output when that option is available.
There is also one audio detail worth checking after generation: the volume of the extension can differ slightly from the original clip. This difference is generally smaller when the source video was itself generated by Seedance 2.5.
After extending, check two places in particular:
the visual join
the audio joinIf both pass cleanly, the extension should feel like part of the original video rather than a second clip attached afterward.
Final Takeaway
Seedance 2.5 gets much easier when you stop trying to control everything with text.
If an image already defines the character, let the image do that job. If a video already has the movement you want, reference the movement. If you already know the important frames, use frames or keyframes. If the video already exists, edit or extend it instead of generating the whole thing again.
Then the prompt only has to explain what is still missing.
If you want to put these workflows into practice, you can access Seedance video model at photogpt and also get access to Minimax H3, Google Omni Flash, Grok Imagine and more.
Last updated on
How to Prompt the MiniMax H3 Video Model
Learn how to structure MiniMax H3 prompts with reference images, keyframes, dialogue, camera movement, environmental sound, and background music.
PhotoGPT Advanced Settings for Image Generation
In this guide, we'll go through the advanced options available in your dashboard for higher control over your generated images.