How to Prompt Grok Imagine Video 1.5
Learn how to prompt Grok Imagine Video 1.5 for image-to-video, text-to-video, reference images, complex scenes, camera control, audio, and video extension.
Grok Imagine can build a video from nothing, bring a still image to life, carry a reference into a new scene, or continue an existing video. The mistake is prompting all four the same way.
A prompt does not need to describe everything in the final video.
If an image already defines the characters, style, lighting, or composition, that information does not need to be rebuilt from text.
If there is no visual input, those details have to come from the prompt.
If several reference images are attached, their roles need to be clear.
And if a video already exists, the prompt can concentrate on what should happen after its current ending.
So before writing the prompt itself, the useful question is:
What information does Grok already have, and what still needs to be explained?
That leads directly into the first workflow.
1. How to Prompt Image-to-Video
Prompt
Visual aesthetic is wet on wet watercolor style. Maintain the watercolor style in every shot with the characters looking like moving watercolor paintings and adhering to the style perfectly.
A man and a woman jokingly discuss how they think they were about to fall in love before they were separated. They are joking to hide that they actually mean it. Insert some pauses to make it seem more natural.
Cinematic, prestige level quality directing and cinematography. Establishing shot, Tight close-ups and handheld camera work for emotional resonance. No wide shots. No music. Micro-expressions in their faces sell the theme more than the volume of their voices.Reference Image
Image-to-video already begins with a visual state.
The source image may carry the characters, clothes, environment, composition, lighting, color, and visual style.
The prompt becomes useful when it fills in what the image cannot show yet:
- action
- performance
- movement
- camera behavior
- sound
- anything that must remain consistent
In this example, the image already establishes the two characters and the watercolor treatment. The missing information is what happens once the frame begins moving.
The prompt does not re-describe both characters. It lets the image carry that information.
The watercolor style is repeated for a different reason: it is being treated as something that should remain stable once motion begins.
Everything else describes information that the still image cannot provide.
The two characters joke to hide what they actually feel. Pauses control the rhythm of the conversation. Close-ups and handheld movement make the performance more intimate. Micro-expressions keep the acting restrained, while No music asks Grok to keep music out of the soundtrack.
| Input | Job |
|---|---|
| Source image | Characters and visual appearance |
| Style instruction | Keeps the watercolor treatment consistent |
| Conversation | Defines the event |
| Emotional subtext | Defines the performance |
| Pauses | Controls conversational rhythm |
| Camera direction | Controls how the performance is filmed |
| Audio instruction | Defines what should or should not be heard |
If the source image is already accurate, additional text is needed only when something should change, move, or be specifically protected.
For a simple image-to-video clip, the prompt can stay close to:
Action + camera + sound + important preservation
A more involved scene can expand from there.
2. How to Prompt Text-to-Video
Prompt
Tracking shot alongside a futuristic matte-black flying car, a low wedge-shaped hovercraft with no wheels, seamless glass canopy and thin white light strips along its edges, four glowing thrust pods beneath it, lit the whole time. It rises fast out of a sea of night fog and banks toward an enormous black equilateral pyramid hovering above the bay. Fog tears off its hull, the two towers of a suspension bridge pass far below, the pyramid fills the sky ahead. A hangar mouth opens low on the pyramid's face, rimmed with thin white light, and the car glides inside and settles onto the floor of the small dark cave.
Monochrome black and white, moonlight through haze, wet metal reflections, sharp focus.
Sound: deep electric thrum from the pods rising with speed, wind buffeting the hull, one distant foghorn, low sub-bass hum from the pyramid, a soft settling hum as the car touches down inside.Text-to-video starts without a visual input carrying information for the prompt.
The scene has to be built from text.
A useful starting point is:
Subject + environment + event + camera + visual direction + sound
For a simple clip, that can still fit naturally into a few sentences.
In this example, the prompt first establishes the subject and immediately places the camera beside it.
Then the event develops in order:
rises from the fog → banks toward the pyramid → passes over the bridge → approaches the hangar → enters → settles
Several actions are happening, but they all belong to one continuous event.
That is different from stacking unrelated actions into the same clip.
The visual treatment is defined once:
Monochrome black and white, moonlight through haze, wet metal reflections, sharp focus.The sound is handled separately:
Sound: deep electric thrum from the pods rising with speed, wind buffeting the hull, one distant foghorn, low sub-bass hum from the pyramid, a soft settling hum as the car touches down inside.This keeps the prompt readable while still giving the soundtrack specific sources.
Instead of writing:
cinematic sci-fi audiothe prompt names what should actually be heard.
The same applies to motion.
Moves fast leaves the scale of movement open.
Rises fast, banks toward, and settles describe different stages of motion more clearly.
Intensity words such as slowly, rapidly, gently, fully, or with tremendous force are useful when the amount of motion matters.
For a simple text-to-video clip, this level may be enough.
If the video needs several shots, strict continuity, physical interactions, or synchronized audio across a longer sequence, the prompt needs more structure.
3. Structuring and Controlling the Prompt
A complex prompt becomes easier to manage when related instructions are kept together.
There is no required heading system.
The headings only help organize information that would otherwise be scattered across the prompt.
Example: A Four-Cut Bridge Sequence
Prompt
Make a project in Grok and name it
Grok Agent Prompt (Imagine 1.5):
SCENE CONTEXT A masked young man my subject from second reference image sits on the upper steel girder of a truss bridge over New York, facing forward along the roadway. Below him a dark horse gallops up the traffic lanes toward him from the far end of the span. He drops from the girder onto the horse's back as it passes beneath him and rides on. Camera is positioned on the same span, first in front of him, then behind him, then overhead.
ACTIVE REFERENCES : young man, slim medium build, thick messy black curly hair under a washed blue denim cap, black square-frame sunglasses with opaque orange tinted lenses, black paisley bandana tied over the lower face, worn black leather jacket, straight-leg blue jeans, brown lace-up leather boots. 100% matches the reference: blue-gray steel truss bridge over New York, three lanes of stopped cars and yellow taxis, raised pedestrian walkway on the left, Manhattan skyline in the far haze. 100% matches the reference.
LOCATION MAP Foreground: the horizontal steel girder sits on, rivets and peeling blue-gray paint, spanning the frame. Midground: three lanes of stopped traffic on the roadway 12 meters below, yellow taxis and dark sedans nose to tail, brake lights burning red. Background: the truss framework repeating in diminishing arches down the span, Manhattan towers flattened in haze at 400 meters depth, haze density 45%. The horse runs from the far end of the span toward the camera along the center lane gap. Light comes from a high overcast sky, camera stays on the shadow side of the girder, operator axis along the bridge's centerline.
FIRST FRAME / BLOCKING is centered in frame, seated on the girder, boots hanging over the edge, forearms resting on his knees, torso squared to camera, head level, chin forward. Wind moves his hair and jacket hem from frame one. Traffic crawls below him at 8 km/h from frame one. Composition rule: subject on the center vertical, steel beam cutting horizontally across the lower third.
FORMAT MODE Sequence of cuts, no timecodes — CUT 1, CUT 2, CUT 3, CUT 4, described in order. OPTICS CUT 1: MS to MCU, 47° at start pushing to 29°, neutral human perspective, rectilinear, no drift mid-segment. CUT 2: MS from behind, 47°, rectilinear, no drift mid-segment. CUT 3: WS on the horse, 29° compressing to 18°, then widening back to 47°, no drift mid-segment. CUT 4: WS overhead, 63° observational, rectilinear, motion-blur on the ground plane, no drift mid-segment.
CAMERA CUT 1: eye-level with the girder, 4 meters in front of, slow dolly-in at 2 km/h from waist framing to head-and-shoulders, focus locked on his face. CUT 2: camera repositions behind his right shoulder, 2 meters back, eye-level, static, looking down the span past his head, focus racks from his shoulder to the roadway ahead. CUT 3: camera holds behind him, pushes toward the approaching horse at 4 km/h, then pulls back at 6 km/h to reframe both the horse and silhouette in the same frame. CUT 4: camera rises to 20 meters directly above the roadway, looking straight down, tracking forward with the horse at 35 km/h. Tonal character: wide highlight roll-off, muted color science, shadows holding detail.
ACTION CUT 1 — Subject motion: Character sits still on the girder, then his eyes narrow behind the orange lenses, the skin at his temples tightening and his cheeks lifting under the bandana as he squints at something far down the span.
Camera motion: continuous dolly-in toward his face. CUT 2 — Subject motion: holds his squint, head level, shoulders still. Camera motion: static behind him. In the roadway below and ahead, a dark horse enters the far end of the span and gallops in a straight line up the gap between the center and left lanes at 35 km/h, hooves striking asphalt, mane and tail streaming back. CUT 3 — Subject motion: shifts his weight forward on the girder, hands releasing his knees and gripping the beam edge. Camera motion: push toward the horse as it closes distance, then pull back so the horse and the seated silhouette share the frame. The horse covers the last 60 meters and passes directly beneath the girder. CUT 4 — Subject motion: pushes off the beam and drops 12 meters straight down, jacket flaring, legs opening, and lands astride the horse's back in one motion. His knees clamp the horse's flanks, his left hand catches the mane, his torso absorbs the impact and settles into the horse's rhythm within two strides. The horse continues galloping forward at 35 km/h between the stopped cars. Camera motion: overhead tracking forward, holding horse and rider centered.
PERFORMANCE The squint is muscle-level: orbicularis tightening, crow's-feet creasing at the outer corners, cheeks rising, brow lowering a few millimeters. Skin shows pore-level detail and a wind-raw flush across the visible band of face between the bandana and the lenses. Catch-lights from the overcast sky sit on the curved surface of the orange lenses. Visible breath fabric movement: the bandana pulses faintly with each exhale. On landing, his jaw sets and his shoulders drop as the impact travels through him.
PHYSICS He falls under full gravity, accelerating through the 12 meters, jacket and hair pushed upward by the airflow. The landing transfers real mass: the horse's stride compresses for one beat, its head dips, then it recovers and drives forward. Hoof contact throws small grit off the asphalt. Contact shadows sit hard under the horse and under each stopped car. Wind moves his jacket hem and hair continuously at 30 km/h.
LIGHTING High overcast sky as a single soft overhead key at 5600K, no hard sun. Fill bounces up off the pale roadway and the car roofs. Brake lights add small red pools of practical spill onto the asphalt around each vehicle. The steel girders read a cool blue-gray under the flat key, their undersides falling into open shadow.
COLOR GRADE Desaturated, cool-leaning palette: blue-gray oxidized steel paint holding the overcast key, brake-light crimson as the only saturated accent burning through the haze, the orange lens tint as a second warm point against the cold field. Blacks lifted slightly, highlights rolled off in the sky.
WARDROBE Black leather jacket, broken-in with visible creasing at the elbows and a soft collar, worn open over a plain black shirt. Washed blue denim cap, brim curved. Black paisley bandana, cotton, knotted at the nape and moving with the wind. Straight-leg medium-wash denim jeans. Brown leather work boots, scuffed at the toes.
AUDIO Wind rushing across the steel girder. Layered traffic hum and idling engines. Hooves striking asphalt in a fast four-beat rhythm, growing louder as the horse approaches. A car horn blaring once. A heavy impact thud as he lands. The horse's sharp exhale.
STYLE Photoreal, cinematic, anamorphic framing, heavy fine film grain across the full frame, hazy atmospheric depth, muted desaturated grade.
OUTPUT SETTINGS 4K, anamorphic, real-time speed throughout all four cuts.
POSITIVE LOCKS keeps the denim cap, orange-tinted opaque sunglasses and black paisley bandana on through all four cuts, face covering intact from frame one to last frame. The horse runs in a straight line up the roadway from the far end of the span toward the camera, staying in the gap between lanes. In CUT 2 the camera sits behind and the horse approaches from ahead of him, entering frame at the far end of the span. He remains seated on the girder through CUTS 1, 2 and 3, and leaves the girder only in CUT 4. He lands on the horse's back and stays mounted, upright, riding forward. Cars stay stopped in their lanes throughout. One overcast time of day and one weather state across all cuts.Inputs
Character Reference:
Environment Reference:
Additional Reference:
The idea itself is straightforward:
A masked man sits on the upper girder of a bridge. A horse approaches through the traffic below. The man waits, drops onto the horse as it passes underneath him, then rides forward.
The challenge is keeping the character, bridge, horse movement, camera, physics, sound, and shot order connected across the full sequence.
Establish the overall video first
The opening SCENE CONTEXT explains the entire event before the detailed instructions begin.
That gives the prompt a clear shape:
man on girder → horse approaches → jump → landing → ride
Everything else supports that sequence.
Define recurring information once
The same character and environment remain throughout the video.
Instead of rebuilding them inside every cut, the prompt defines them in ACTIVE REFERENCES, LOCATION MAP, WARDROBE, and POSITIVE LOCKS.
That keeps recurring information out of the individual action blocks.
Break the sequence into clear cuts
The action is divided into four jobs:
Cut 1: establish the character and his reaction.
Cut 2: introduce the horse.
Cut 3: bring the horse directly beneath him.
Cut 4: complete the jump, landing, and ride.
The progression stays easy to follow because each cut has one clear purpose.
Keep camera instructions with the shot they belong to
The camera changes along with the action:
- Cut 1 uses a slow dolly toward the face
- Cut 2 holds from behind
- Cut 3 pushes toward the horse and then pulls back
- Cut 4 rises overhead and tracks the rider
A camera does not need to move in every shot.
A locked or static camera is useful when subject motion is already carrying the scene. When a specific camera move matters, it should be named clearly.
Give difficult behavior its own rules.
The hardest moment is the landing.
The prompt gives that interaction a separate PHYSICS block.
The character falls under gravity. Clothing and hair respond to the airflow. The horse absorbs the impact for one stride, then recovers.
That creates a physical chain:
fall → impact → reaction → recovery
The same technique is useful whenever one part of the scene has its own logic:
- liquid behavior
- collisions
- transformations
- fabric movement
- breaking objects
- product interaction
- damage that should persist
The rule belongs together instead of being repeated across several cuts.
Keep global audio outside the cuts
The soundtrack applies to the sequence as a whole:
- wind across the bridge
- traffic
- idling engines
- approaching hooves
- a horn
- landing impact
- the horse's exhale
Keeping this in one AUDIO block prevents the same ambience from being repeated inside every shot.
Specific sound sources are more useful than broad wording such as cinematic sound.
Use continuity instructions for things that actually matter
The POSITIVE LOCKS block protects:
- the cap
- sunglasses
- bandana
- horse path
- character position before the jump
- stopped traffic
- time of day
- weather
A large lock block is not necessary for every Grok prompt.
It is useful here because several shots need the same state to survive across the full sequence.
What to Take From This Example
The useful structure underneath the custom headings is:
Overall event → recurring details → shot progression → special behavior rules → global audio/style → continuity that matters
For a simple video, only a few of those parts may be necessary.
For a controlled multi-shot sequence, the larger structure makes the instructions easier to follow.
4. Using Reference Images
Once reference images are attached, the prompt no longer has to describe everything from scratch.
The important part is telling Grok what each image is supposed to control.
A reference may contain a person, clothing, pose, lighting, background, and composition at the same time.
If only one of those properties matters, the role should stay narrow.
Reference-to-video is different from image-to-video. A reference image guides the character, object, style, or other visual details in the generated scene without becoming the starting frame itself.
Example: Three References With Three Different Jobs
Prompt
CN Medium-close-up. Single take, no cuts. For positioning and lighting, refer to Image1. The camera is positioned to the side, right next to Image2 and Image3—the entire venue is not visible; only a warm, blurred background is shown. They place the green beer bottles on the table—quickly, casually, barely glancing at them. Then Image2 immediately comes into her own—breaking into a huge smile, shouting as she moves, bright and completely unfiltered: “Yay—yay—yay! We finally won this war!” She’s already running and jumping forward—a playful, bouncy jog, light yet unstoppable, pure energy. Halfway through, she glances back at Image3. Image3 meets her gaze—a mature, slightly restrained smile, warm yet measured—and then they embrace each other warmly; they look incredibly happy. The camera follows them from the side throughout the sequence, with warm light softly falling on both their faces. Slight handheld drift. ARRI 35, anamorphic lens. Realistic style. No 3D animation. Real textures, real weight, real motion blur. SFX only. No music.Reference Images
Positioning and lighting:
Main active character:
Second character:
The prompt does not describe the references twice.
Image 1 already carries the positioning and lighting.
Images 2 and 3 already carry the two characters.
The text is mainly used for the parts those images cannot show:
- performance
- action
- relationship between the characters
- camera behavior
- sound
Image 2 is given an energetic, unfiltered performance.
Image 3 is given a warmer and more restrained response.
The camera follows both characters from the side.
The reference mapping removes the need to guess which image should control which part of the final video.
A simple setup can be written as:
Image 1 defines the character.
Image 2 defines the outfit.
Image 3 defines the environment.Or:
Image 1 controls the product.
Image 2 controls the character.The goal is not to give every visible detail a separate reference.
The goal is to make the role of every supplied reference clear.
5. Video Extension
Video Extension starts from a video that already exists.
The character, environment, lighting, camera position, movement, and visual style are already established. The extension prompt mainly needs to explain what should continue or happen next.
Example: Extending a Continuous Violin Performance
The final video continues the same violin performance beyond the original ending while keeping the character, room, lighting, and camera style consistent.
How to Prompt This Continuation
A prompt for this kind of continuation can be written as:
Continue seamlessly from the final frame with no cut, restart, or change of scene.
The same woman continues playing the violin naturally, maintaining the same bowing rhythm, hand placement, posture, and subtle body movement.
Keep her long blonde hair, white lace dress, violin, facial appearance, and proportions consistent.
Preserve the same elegant sunlit room, piano in the background, warm afternoon light, window highlights, floor reflections, and soft cinematic depth of field.
Continue the existing camera movement smoothly from the same position. Maintain the intimate portrait framing with a gentle natural drift, then gradually ease back toward a slightly wider view as the performance continues.
The violin performance should continue naturally across the extension point with no visible interruption or reset.
No cuts, no new characters, no wardrobe changes, no change of location, and no sudden lighting shift.
End with her still playing the violin naturally in the same room.Original Video
The original portion already establishes:
- the woman and her appearance
- the white lace dress
- the violin performance
- the room and piano
- warm natural lighting
- camera position and movement
- the overall visual style
The extension does not need to describe all of those details again.
It mainly needs to continue the action from the point where the original video ends.
The progression is:
original performance → extension begins → violin playing continues → camera continues → longer final performance
The important part is that the scene does not reset at the extension point.
The woman continues playing instead of starting the performance again. The room and lighting remain consistent, and the camera continues naturally with the action.
For a simple continuation, the prompt can stay short:
She stands up slowly, grabs her coat, and walks toward the door. The camera follows.This works because the existing video already contains the scene.
When more control is needed, the extension prompt can also describe:
- what action should continue
- how the camera should continue
- which character details should stay the same
- which parts of the environment should remain unchanged
- how movement should carry through the join
A useful structure is:
next action + camera continuation + important continuity
The goal is for the extension to feel like the original video simply kept running, rather than a new clip starting from scratch.
6. How to Fix Common Prompt Problems
A bad result does not always mean the whole prompt is bad.
Sometimes the action is right but the camera is wrong. Sometimes the character looks good but the sound does not. The easiest way to improve the next generation is to keep what already works and change only the part that does not.
| Problem | First thing to change |
|---|---|
| Motion is too weak | Action verb, speed, or intensity |
| Too many things happen | Remove unrelated actions |
| Camera feels random | Specify the camera behavior |
| Camera should stay still | Use locked or static |
| Face or product changes | Add targeted preservation |
| Audio feels generic | Name real sound sources |
| References get mixed up | Give each reference a clearer role |
| Shot order breaks down | Separate the sequence into clearer beats |
| Physical interaction looks wrong | Describe both action and reaction |
| Extension feels disconnected | Continue from the existing end state |
Keep simple prompts simple
A simple shot usually does not need a long prompt.
For example, a portrait may only need:
She slowly turns toward the camera and smiles.
Locked medium close-up.
A light breeze moves a few loose strands of hair.
Sound: quiet room tone and soft fabric movement.
Keep her face and outfit unchanged.There is no reason to add a location map, several camera instructions, detailed physics, or a long continuity block when the scene does not need them.
Give complicated scenes enough structure
The opposite problem is trying to control a complicated video with only a few vague sentences.
If a scene depends on several cuts, exact positions, camera changes, physical interactions, synchronized sound, or continuity, the prompt needs enough structure to keep those instructions clear.
The right prompt length depends on how much control the video needs.
Change one thing at a time
When a generation is already close, changing several parts of the prompt at once makes it harder to know what actually improved the result.
If the motion works but the camera does not, change the camera instruction.
If the camera works but the product or character changes, change the preservation instruction.
If everything looks right except the sound, change the audio.
Keeping the parts that already work makes each new generation easier to compare and improve.
Final Takeaway
Grok Imagine gets easier to prompt when each input is allowed to carry the information it already contains.
If an image already defines the visual state, let it do that job.
If there is no image, establish the scene through text.
If several references are supplied, give each one a clear role.
If a sequence becomes complicated, organize the prompt so recurring rules and shot-specific instructions stay separate.
If the video already exists, continue from its current ending instead of rebuilding the scene.
Then the prompt only has to explain what is still missing.
If you want to put these workflows into practice, you can try Grok Imagine Video 1.5 directly on PhotoGPT, or open the PhotoGPT AI Video Generator to explore Grok alongside Seedance, MiniMax H3, Veo, Kling, and other video models.
Last updated on
How to Prompt Seedance 2.5 Video Model
Learn how to prompt Seedance 2.5 for text-to-video, references, camera motion, style, audio, keyframes, storyboards, video editing, and extensions.
PhotoGPT Advanced Settings for Image Generation
In this guide, we'll go through the advanced options available in your dashboard for higher control over your generated images.