How to Prompt the MiniMax H3 Video Model
Learn how to structure MiniMax H3 prompts with reference images, keyframes, dialogue, camera movement, environmental sound, and background music.
MiniMax Hailuo H3 is a multimodal video generation model available in PhotoGPT.
It supports reference images, keyframes, video input, and audio input in one context window for continuous videos up to 15 seconds.
To use MiniMax H3 effectively, you need to understand the prompt structure the model follows.
In this guide, we will learn how to prompt MiniMax H3 using:
- A prompt and reference images
- A prompt and keyframes
How to structure the prompt
The prompt is divided into 4 parts.
- Reference Images / Keyframes Images
- Integrated Multimodal Description
- Overall Soundscape or Environmental and physical sounds
- Non Diegetic Music or Background sound
The complete structure of the H3 prompt is:
REFERENCE-IMAGE ALIGNMENT INSTRUCTION
integrated_multimodal_description:
The visible and synchronized audiovisual timeline.
overall_soundscape:
The full environmental and physical sound environment.
non_diegetic_music:
The audience-only background score, or N/A.To keep the explanation simple and easy to follow, we will start with Integrated Multimodal Description.
A) Integrated Multimodal Description
MiniMax H3 uses integrated_multimodal_description for the main audiovisual timeline. It describes visuals, actions, shots, speakers, dialogue, singing, and diegetic audio.
Reusable single-field formula
integrated_multimodal_description:
[Shot 1]
STYLE + INITIAL FRAMING + SUBJECT + ENVIRONMENT.
REFERENCE OR OPENING-STATE CONTINUITY.
SMALL NATURAL MOVEMENT.
ACTION ONSET.
SPEAKER DESCRIPTION + SPEAKER ID + DELIVERY:
<d>[Language] Exact dialogue.</d>
VISIBLE PHYSICAL RESPONSE TO THE ACTION OR SPEECH.
CAMERA MOTION TYPE + AMPLITUDE + SPEED.
INTERMEDIATE COMPOSITIONAL CHANGES.
SYNCHRONIZED DIEGETIC SOUND OR VISIBLE ENVIRONMENTAL RESPONSE.
FINAL CHARACTER STATE + FINAL COMPOSITION.For multiple shots:
integrated_multimodal_description: [Shot 1] ...
[Shot 2] At 00:03.500, the camera cuts to...
[Shot 3] At 00:06.000, the shot switches to...Let’s understand it properly with a really good and detailed prompt:
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, cinematic and photorealistic. The young female hiker shown in <Picture 1> remains on the narrow mountain ridge, preserving her exact facial identity, natural skin texture, brown hair, red insulated hiking jacket, backpack straps, body proportions, front-facing orientation, lighting, background mountains and initial composition.
The camera begins in the close head-and-shoulders framing established by <Picture 1>. She remains unaware of the camera and keeps her eyes fixed on one distant point in the valley ahead. A light mountain wind moves only a few loose strands of her hair and slightly shifts the fabric around her collar.
She takes one natural breath, gently raises her chin and loudly calls into the valley. The young woman with a clear, strong, natural American voice (S1) shouts with joyful energy: <d>[English] Helloooooooo!</d>
Her lips, jaw, cheeks, throat and chest move naturally with the sustained call. Her arms remain relaxed beside her body. Her feet, hips, shoulders and head remain oriented in the same direction; she does not rotate, step, turn, or follow the camera with her eyes.
As she begins shouting, the camera pulls out with large amplitude at slow speed while pedestaling upward with small amplitude at slow speed. It travels directly backward along the same optical axis, progressing continuously from a close portrait to a chest-up frame, then to a full-body view, and finally to a wide view of the ridge and surrounding valley.
The camera does not orbit, truck sideways, pan around her, or use a digital zoom. Realistic parallax reveals the steep grassy slopes, distant layered mountains and deep valley while the horizon remains stable.
As the camera moves farther away, her direct voice becomes quieter and more distant. After a short natural delay, the final syllable reflects first from the left mountain wall and then more softly from the right side of the valley, creating separated stereo echoes that gradually become quieter and more diffuse.
She completes the call and remains standing in exactly the same place and orientation. By the end of the shot, she appears as a small red figure on the vast green ridge while the final distant echo fades naturally across the valley.This is a 9-second, 16:9 cinematic shot of a hiker calling “Helloooooooo” from a mountain ridge.
Breaking down the prompt
In the breakdown below, important fields are marked as Structural field. These fields should not be overlooked while writing the prompt.
1. Field label (Structural field)
integrated_multimodal_description:This tells H3 that everything following it belongs to the main audiovisual timeline. This field can contain visuals, actions, speakers, dialogue, camera movement, shot changes and synchronized diegetic audio.
2. Shot declaration (Structural field)
[Shot 1]Every prompt should begin with a shot identifier. The first shot does not need a timestamp. Only later shots need precise cut times.
Correct:
[Shot 1] Live-action, cinematic...Incorrect:
[Shot 1] At 00:00.000...3. Visual style
Live-action, cinematic and photorealistic.This establishes the rendering language immediately. Use a direct visual category rather than vague praise like “million-dollar quality.”
Useful style labels include:
Live-action
Cinematic
2D-animated
3D CG
Claymation
Watercolor
Vintage filmThe guide recommends establishing the overall style and initial composition at the beginning of Shot 1.
4. Reference-image anchor (Structural field)
The young female hiker shown in <Picture 1> remains on the narrow mountain ridge...This grounds the video in the supplied image. Immediately identify what must remain consistent:
- character identity
- appearance
- clothing
- body proportions
- position
- lighting
- background
- composition
For image-to-video, the recommended progression is:
first-frame anchor
→ action begins
→ continuous development
→ result5. Continuity locks
preserving her exact facial identity, natural skin texture, brown hair...This prevents the model from interpreting the reference as loose inspiration.
Only lock details that matter. Do not write dozens of generic negative instructions before describing the actual video.
A useful order is:
identity
→ wardrobe
→ position
→ orientation
→ environment
→ composition6. Initial physical state
She remains unaware of the camera and keeps her eyes fixed on one distant point...This defines the character’s relationship with the camera before movement begins.
It solves behaviors like turning toward the moving camera, tracking the lens with her eyes, rotating her body during a pullback, performing directly for the viewer.
Observable physical instructions work better than saying:
Do not look AI-generated.7. Small environmental motion
A light mountain wind moves only a few loose strands of her hair...Subtle secondary movement makes the shot feel alive without destabilizing the character.
Good secondary motion can be like loose hair moving, fabric shifting, breathing, smoke drifting, rain falling, leaves reacting to wind.
Avoid giving every object strong motion simultaneously.
8. Action onset
She takes one natural breath, gently raises her chin...This creates a believable physical lead-in to the main action.
Instead of making the character suddenly shout, the model receives a short causal sequence:
breath
→ chin rises
→ mouth opens
→ voice beginsThat usually improves human motion and lip synchronization.
9. Speaker identity (Structural field)
The young woman with a clear, strong, natural American voice (S1)...(S1) creates a stable speaker identity. On the speaker’s first vocal appearance, specify useful voice characteristics outside the dialogue tag like age or character type, gender, pitch, timbre, accent, speaking rate, delivery.
A character keeps the same speaker ID throughout later shots.
10. Dialogue syntax (Structural field)
<d>[English] Helloooooooo!</d>Inside <d>, include only:
language tag + exact spoken wordsDo not place acting instructions inside the dialogue tag.
Correct:
She shouts with joyful energy: <d>[English] Helloooooooo!</d>Incorrect:
<d>[English, shouting loudly and happily] Helloooooooo!</d>11. Physical vocal performance
Her lips, jaw, cheeks, throat and chest move naturally...This translates “realistic shouting” into observable anatomy.
It is more useful than:
Perfect lip sync and realistic acting.You are telling the model what realism looks like jaw opens naturally, cheeks respond to pressure, throat engages, chest contracts during breath, lips shape the sustained vowel
12. Body-orientation lock
Her feet, hips, shoulders and head remain oriented in the same direction...This is stronger than simply saying:
She does not turn.It defines the physical parts that must remain fixed and prevents the model from satisfying one instruction while rotating another part of the body.
13. Camera motion type
The camera pulls out...Use pull out when the camera physically travels backward.
Use zoom out only when the camera remains stationary and the lens changes focal length.
motion type + amplitude + speed| Dimension | Available expression | Description |
|---|---|---|
| Motion type | Zoom In / Zoom Out | The focal length changes while the camera body remains stationary |
| Motion type | Push In / Pull Out | The camera moves forward or backward |
| Motion type | Pan Left / Pan Right | The camera remains in place while the lens pivots horizontally |
| Motion type | Truck Left / Truck Right | The camera translates horizontally |
| Motion type | Tilt Up / Tilt Down | The camera remains in place while the lens pivots vertically |
| Motion type | Pedestal Up / Pedestal Down | The entire camera moves upward or downward |
| Motion type | Arc Shot | The camera moves in an arc around the subject |
| Motion type | Tracking Shot | The camera follows a moving subject |
| Motion type | Static Shot | The camera position and lens remain still |
| Motion type | Shake Slightly / Shake Strongly | Slight or strong camera shake |
| Motion type | POV | The subject's point of view |
| Motion type | Roll Clockwise / Roll Counterclockwise | The camera rolls around the lens axis |
| Amplitude | with small amplitude | Small-range change |
| Amplitude | with large amplitude | Large-range change |
| Speed | at slow speed | Slow movement |
| Speed | at fast speed | Fast movement |
If you want to learn about camera prompts in more detail, you can learn from here.
14. Camera amplitude
Camera amplitude means how far the camera moves, or how much the framing changes during that movement.
small amplitudeUse for subtle framing changes.
large amplitudeUse when moving from a close portrait to a wide landscape.
15. Camera speed
Speed describes the pacing of the camera movement.
at slow speed
at fast speedHere, slow speed supports a smooth cinematic reveal rather than a sudden artificial transformation.
16. Combined camera movements
pulls out with large amplitude at slow speed while pedestaling upward with small amplitude at slow speedThis describes two simultaneous physical movements:
- Pull out: the camera moves backward.
- Pedestal up: the entire camera rises vertically.
The movements are written as natural actions inside the shot rather than as disconnected technical labels.
17. Framing progression
close portrait
→ chest-up frame
→ full-body view
→ wide ridge viewThis gives the model clear intermediate stages.
Without intermediate framing, the model can abruptly shrink the subject instead of moving the camera physically.
18. Camera-path restrictions
The camera does not orbit, truck sideways, pan around her, or use a digital zoom.These restrictions directly prevent unwanted camera behavior.
Each word has a distinct meaning:
- Orbit / arc: camera moves around the subject.
- Truck: camera moves sideways.
- Pan: stationary camera pivots horizontally.
- Digital zoom: image enlarges or shrinks without physical parallax.
19. Environmental reveal
Realistic parallax reveals the steep grassy slopes...This tells the model what becomes visible because of the camera movement.
A camera instruction is stronger when paired with its visual result:
camera movement
→ newly revealed informationFor example:
The camera trucks right, revealing the open doorway.I recommend using camera movement when the viewpoint changes continuously, rather than creating an unnecessary cut.
20. Synchronized diegetic audio
As the camera moves farther away, her direct voice becomes quieter...This sound is synchronized with a visible event, so it can remain inside integrated_multimodal_description.
The key relationship is:
camera distance increases
→ direct voice becomes quieter
→ reflections arrive after a delay
→ echoes decayThis is more effective than simply writing:
Add stereo echo.21. Spatial audio direction
first from the left mountain wall and then more softly from the right...This tells the model:
- which side produces the first reflection
- which side produces the second
- their relative timing
- their relative volume
- how the echoes decay
It prevents the model from merely duplicating the voice equally in both channels.
22. Final state
By the end of the shot, she appears as a small red figure...Always define the final visual result.
A complete single-shot description needs:
opening state
→ action
→ camera development
→ final stateWithout a final state, H3 has less guidance about where the shot should land.
So far, we have covered the structural fields, the prompt order, and the wording variations that help provide clearer control.
Another segment of this guide is to understand how to write prompt when keyframes are included.
B) How to Incorporate Keyframes into the Prompt
There are three keyframe modes available with the MiniMax H3 video model.
- First Frame
- First Frame, Last Frame
- Last Frame
Each keyframe has its own fixed guidance structure and the guidance appears before integrated_multimodal_description. It tells H3 exactly where each uploaded image belongs on the video timeline.
A) First-frame mode
For the target video, at 0.00 seconds into the target video,
<Picture 1> from [Shot 1] is fully referenced.What each part means
0.00 seconds:The reference image is the exact opening moment.<Picture 1>:The first uploaded image.[Shot 1]:The image belongs to the first shot.fully referenced:The model should preserve the image’s subject, composition, wardrobe, environment and visual arrangement rather than treating it as loose inspiration.
Example Prompt:
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] An 8-second, 16:9 live-action cinematic shot begins exactly from the composition established by <Picture 1>. Preserve the same real young woman, facial identity, natural skin texture, brown hair, headband, red insulated hiking jacket, backpack straps, mountain ridge, distant valley, daylight, and front-facing camera position.
The camera begins with the same close head-and-shoulders framing. The woman is standing firmly in one position on the mountain ridge, looking past the camera toward one fixed distant point in the valley. Her feet, hips, shoulders, neck, head, and gaze remain oriented in the same direction throughout the video. She does not rotate, pivot, step, or follow the moving camera with her eyes.
A light mountain wind moves a few loose strands of her hair and subtly shifts the fabric around her collar. She takes one natural breath, slightly raises her chin, and joyfully calls into the open valley. The young woman with a clear, powerful, natural American voice (S1) shouts: <d>[English] Helloooooooo!</d>
Her lips, jaw, cheeks, throat, and chest move naturally with the sustained vowel. Her expression remains believable and restrained, like a real hiker calling across a large valley. Her arms stay relaxed beside her body, and she does not raise her hands toward her mouth.
At the exact moment her voice begins, the camera pulls out with large amplitude at slow speed while pedestaling upward with small amplitude at slow speed. The camera travels directly backward along the same optical axis in front of her, progressing continuously from the close portrait to a chest-up frame, then a full-body view, and finally a wide aerial view of the mountain ridge.
The camera does not move sideways or arc around her. Because it remains directly in front of her, she stays naturally front-facing without turning her body. Realistic environmental parallax gradually reveals the steep grassy ridge, deep valley, and distant layered mountain peaks.
As the camera moves farther away, her direct voice becomes progressively quieter, more distant, and slightly softer in its high frequencies. After a believable delay, the sustained final vowel reflects from the left mountain wall, followed by a softer and slightly later reflection from the right side of the valley. The reflections spread naturally across the stereo field, becoming quieter, darker, and more diffuse with each return.
She finishes the call and remains standing in exactly the same position and orientation. At the end of the shot, she is a small red figure on the vast green ridge while the last distant echo fades into the valley air. The image retains natural skin texture, realistic fabric movement, stable anatomy, authentic 24 fps motion blur, restrained color grading, and physically believable camera movement.B) First-frame and last-frame mode
Picture 1 from Shot 1 aligns with 0.00 seconds;
Picture 2 from Shot 1 aligns with 8.00 seconds.What each part means
Picture 1:Your opening close-up image.0.00 seconds:The video must begin from Picture 1.Picture 2:Your ending wide aerial image.8.00 seconds:The video must arrive at Picture 2 at the exact end.Shot 1for both images
Both images belong to one continuous camera shot.
This alignment instruction establishes the two endpoints.
Example Prompt:
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] An 8-second, 16:9 live-action cinematic single shot begins exactly from the close front-facing composition, character appearance, clothing, lighting, and mountain environment established by Picture 1, then develops continuously toward the wide aerial composition established by Picture 2.
Preserve the same real young woman, facial identity, natural skin texture, brown hair, headband, red insulated hiking jacket, backpack straps, body proportions, and mountain location throughout the transformation in camera distance.
She stands firmly in one fixed location on the ridge and looks toward one fixed distant point directly beyond the camera. Her feet, hips, shoulders, neck, head, and gaze remain oriented in the same direction for the complete shot. She never turns, pivots, rotates, steps, or tracks the retreating camera with her eyes.
A gentle mountain wind moves several loose strands of hair and lightly shifts the fabric around her jacket collar. She takes a natural breath, raises her chin slightly, and loudly calls into the valley. The young woman with a clear, strong, natural American voice (S1) joyfully shouts: <d>[English] Helloooooooo!</d>
Her mouth opens naturally, and her lips, jaw, cheeks, throat, breathing, and chest respond realistically to the sustained call. The performance remains restrained and physically believable. Her arms stay relaxed beside her body.
As her voice begins, the camera pulls out with large amplitude at slow speed while pedestaling upward with small amplitude at slow speed. It travels directly backward along the optical axis without moving sideways, orbiting, or changing direction.
The composition progresses smoothly through clear intermediate stages: close facial portrait, chest-up framing, waist-up framing, full-body view, medium-wide ridge view, and extreme-wide aerial landscape view. Each stage results from physical camera travel and realistic parallax, not artificial subject shrinking.
As the distance increases, the steep grassy ridge, surrounding slopes, deep valley, and distant mountain layers become progressively more visible. The camera movement gradually narrows every visual difference between the evolving shot and Picture 2.
The direct voice begins loud, close, dry, and centered. As the camera retreats, it becomes naturally quieter and more distant. After a short physical delay, the final sustained vowel reflects from the left mountain wall, then arrives more softly from the right side of the valley. Later reflections become progressively quieter, darker, wider, and more diffuse.
The woman completes the call without changing her position or orientation. During the final seconds, the camera continues approaching the exact altitude, distance, horizon level, ridge geometry, subject scale, subject placement, lighting, and landscape composition established by Picture 2.
At exactly 8.00 seconds, the shot settles precisely into Picture 2. The woman appears as the same small red figure in the exact final position on the vast mountain ridge while the last faint stereo reflection disappears into the open valley.C) Last-frame mode
<Picture 1> from [Shot 1] aligns with the 8.00-second mark.What each part means
<Picture 1>:The uploaded image is the ending image, not the opening image.[Shot 1]:The reference belongs to the final shot. In a one-shot video, that remains Shot 1.8.00-second mark:At exactly eight seconds, the generated video must match the uploaded image.
The integrated_multimodal_description explains what happens before the supplied image and how all movement settles into it.
Example Prompt:
How the reference pictures align with the target video — <Picture 1> (from [Shot 1]) aligns with the 8.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] An 8-second, 16:9 live-action cinematic single shot begins with a close front-facing portrait of the same real young woman in the red insulated hiking jacket shown as the small figure in <Picture 1>. She stands on the same narrow grassy mountain ridge under the same natural daylight, with the same distant valley and layered mountain environment visible behind her.
The opening camera is positioned directly in front of her at eye level. Her face has natural human proportions, realistic skin texture, slight facial asymmetry, wind-reddened cheeks, brown hair held back by a headband, and several loose strands moving gently in the mountain wind. Her backpack straps and red jacket remain physically consistent with the final figure visible in <Picture 1>.
She is firmly planted in one location and looks past the camera toward one fixed distant point in the valley. Her feet, hips, shoulders, neck, head, and gaze remain oriented in the same direction throughout the video. She never turns, pivots, steps, or follows the camera.
She takes a natural breath, slightly raises her chin, and joyfully calls into the valley. The young woman with a clear, powerful, natural American voice (S1) shouts: <d>[English] Helloooooooo!</d>
Her mouth, jaw, cheeks, throat, chest, and breathing move naturally during the sustained vowel. Her arms remain relaxed beside her body. Her expression is energetic but believable, without theatrical screaming or exaggerated facial strain.
As the voice begins, the camera pulls out with large amplitude at slow speed while pedestaling upward with small amplitude at slow speed. It moves directly backward along the same optical axis, maintaining the woman near the center of the frame without requiring her to rotate.
The framing changes continuously from a facial close-up to a chest-up view, a full-body view, a wide ridge composition, and finally an extreme-wide aerial landscape view. Physical parallax reveals increasing areas of the grassy ridge, steep slopes, deep valley, and distant mountain range.
Her direct voice gradually becomes quieter and more distant as the camera retreats. The final sustained vowel reflects first from the left side of the valley and then more softly from the right, producing separated stereo echoes with believable delays. Each later reflection becomes quieter, darker, and more diffuse.
During the final seconds, the woman remains completely still in her original orientation as the camera progressively converges on the precise final distance, height, viewing angle, horizon position, terrain arrangement, lighting, atmosphere, subject scale, and subject placement established by <Picture 1>.
At exactly 8.00 seconds, all movement settles into the exact final composition of <Picture 1>. The woman is the same small red figure standing on the vast ridge, and the final distant echo fades naturally as the image lands precisely on the supplied last frame.So far, we have covered how to structure the prompt and how to add keyframe guidance.
C) Environmental and physical sound
overall_soundscape: Steady high-altitude wind moves across the exposed ridge, accompanied by subtle jacket-fabric movement and one audible preparatory breath. The woman’s sustained call produces delayed natural reflections from the left and right mountain walls, with each echo becoming quieter, darker and more diffuse as it travels across the valley.Why this is separate
Inside integrated_multimodal_description, you connect sound to precise visible events:
She takes a breath and shouts...Inside overall_soundscape, you describe the complete acoustic environment:
wind + clothing + breath + valley reflectionsUse N/A only when the entire video is explicitly intended to be silent.
D) The background sound
Use 1–3 English sentences to describe background music that the characters cannot hear and only the audience can hear.
Focus on instrumentation, speed, rhythm, and dynamic changes; do not use abstract mood words or explain the emotional function of the score. Singing, instruments, radio, television, or phone music audible to the characters are diegetic events and should appear in the multimodal description. Use N/A when there is no non-diegetic music.
non_diegetic_music: Sparse piano notes at a slow tempo, joined by sustained low strings that gradually increase in volume before fading out.E) On Screen Text
Place any banner, sign, label, subtitle, or neon text that is actually visible on screen in English double quotation marks. Preserve the original text and punctuation verbatim, without translation.
A red neon sign reading "营业中" glows above the doorway.You now have the complete structure needed to write clearer prompts for the MiniMax H3 model.
References
Last updated on
Complete Lighting Prompt Guide for Portrait Photography
Master every major portrait lighting style and learn exactly how to describe it in an AI prompt. Golden hour, Rembrandt, chiaroscuro, softbox, high key, headshot lighting and more — with full prompt examples for each.
PhotoGPT Advanced Settings for Image Generation
In this guide, we'll go through the advanced options available in your dashboard for higher control over your generated images.