PhotoGPTGuides

How to Prompt the Kling 4.0 Video Model

Learn how to prompt Kling 4.0 for text-to-video, image references, Omni Reference, keyframes, storyboards, camera movement, dialogue, sound, and video editing.

You can picture the scene. But what do you actually write to make it happen?

The character needs to look the same throughout. The camera needs to follow the action. A spoken line needs to come after a pause. How do you explain all of that to Kling 4.0 without writing a page of instructions?

This guide walks you through it with complete prompt examples. You’ll learn what to describe, when to use reference images and videos, and how to connect movement, camera direction, and sound. Each example explains the wording so you can write a prompt for your own scene.

Build your first Kling 4.0 prompt

Let's begin with text-to-video: creating a scene from words, without an uploaded image. We will use the street reunion below to build the description one part at a time.

Describe who is there and what happens

“Two friends meet again” gives us the idea, but not the action. They could already be hugging when the clip begins. They could simply wave to each other from a distance.

For this reunion, the important sequence is recognise each other, approach, then hug. Write those actions in order:

Two friends who have not seen each other for a long time meet at the corner of a quiet street at dusk. They smile, walk quickly toward each other, and share a warm embrace.

The street and time of day give the encounter a setting. The verbs explain what should happen there. The approach matters because we want to see the reunion begin, not just its final pose.

Decide how the camera should show it

A continuous shot keeps filming without switching to another view through a cut. A handheld look includes the small movements of a camera held by a person.

For this scene, the direction is:

A handheld camera follows them with subtle camera shake.

“Follows them” makes the viewpoint move with the encounter. “Subtle” keeps the movement suited to a warm reunion rather than a frantic chase. The camera is the viewpoint we watch from; this instruction does not ask for filming equipment to appear in the picture.

Add the appearance and sound

Soft, warm light gives this scene its evening appearance. Footsteps and street ambience place the action in an inhabited street. Here, ambience means the ongoing background sounds of the location.

We can ask for English dialogue without choosing every word. When a particular line matters, write the line itself. The dialogue section later in this guide explains how to assign words to the right person.

Put the instructions together

A continuous 7-second shot set in the UK. Two friends who have not seen each other for a long time meet at the corner of a quiet street at dusk, surrounded by a diverse crowd going about their daily lives. They smile, walk quickly toward each other, and share a warm embrace. Their facial expressions and body movements are natural. A handheld camera follows them with subtle camera shake. Soft, warm lighting with a nostalgic cinematic film look. Include natural footsteps, street ambience, background music, and English dialogue.

Read the prompt in groups: setting, action, camera, appearance, then sound. Each group makes a different decision. You do not need special headings inside Kling to keep those decisions understandable.

A reunion that develops from approach to embrace

The scene develops from the approach into the embrace; pedestrians remain part of the street setting.

Try this prompt: Open Kling 4.0 in PhotoGPT, copy the reunion prompt above, and change the setting to match your idea.

When can the prompt stay short?

A short prompt suits a scene whose details you are happy to leave open:

An old man ranting about AI.

This gives the character and subject of the performance. It does not specify his room, clothes, gestures, camera view, or exact words. The accompanying clip includes an armchair and a speaker on a side table, but neither object is requested in that sentence.

That distinction matters when you reuse the prompt. A detail visible in one generated video is not automatically a requirement in the words that describe it. Write the detail when your own scene depends on it.

A short request leaves the setting open

The room, armchair, and side-table speaker are visible in the clip but are not specified by the one-sentence prompt.

Choose the right way to start your video

Choose the input method around what you already have. This changes what the prompt needs to explain.

What you haveStarting pointWhat your words need to add
A written ideaText-to-videoThe people, place, action, and way the scene is shown.
An image that should open the clipImage-to-video / first frameWhat happens after that starting view.
An opening image and an ending imageFirst & Last FramesThe movement or change that connects them.
Images of important momentsMulti-keyframesThe order of those moments and the action between them.
Separate appearance and movement referencesOmni ReferenceWhat to take from each upload.
A video with something to changeVideo editingThe target of the edit and the parts to retain.

A first frame and a character reference have different jobs. The first frame establishes the opening picture. A character reference establishes who appears, without requiring that exact picture to open the video. The studio portrait below illustrates an opening view; the basketball character sheet later supplies a player for a different scene.

Animate a photo without changing the whole scene

Look at the studio portrait below. It already establishes the performer, headphones, white top, microphone, and recording booth. There is no need to invent a second description of those objects.

The moving part is the performance: changes in expression, head position, and gestures. For the visual direction, write:

Start from the studio portrait. Keep the performer beside the microphone, wearing the same headphones and white top. Let his facial expression and head position change during the performance, with occasional small hand gestures entering the lower part of the picture.

Keep the microphone, the round screen in front of it, and the booth in their existing positions. Hold the camera at the opening position rather than moving around the performer.

The first paragraph allows the person to move. The second identifies the objects and viewpoint you want to retain. “Keep everything unchanged” would mix those two requests together.

This describes the visible performance. It does not specify a song or spoken script. Add your actual words when the audio needs to follow a script, rather than expecting a portrait to determine them.

An opening portrait establishes the studio setup

Opening image

A performer in headphones and a white top beside a microphone and pop filter.

Studio-performance video

Compare the fixed studio setup with the changing expression and gestures. The image establishes appearance, not the words or a reference voice.

Give the performance a clear purpose

For a talking introduction, name the format and attitude directly:

A vlogger-style video in which the woman cheerfully introduces herself to the audience. Maintain a natural, authentic look and realistic image quality.

The living-room clip illustrates this kind of performance. The woman addresses the camera and changes expression as the introduction develops.

“Cheerfully introduces herself” explains what she is doing and how she should come across. It still leaves the words open. That is useful for a general introduction, but a brand message needs its actual script.

A living-room vlog introduction

Look at the changing expression and direct-to-camera presentation. This is a performance example, not a product demonstration.

Need to remove a distraction from the opening photo? Our AI Photo Editor guide shows how to edit the still image before you animate it.

Describe the camera movement in plain language

Think about where the viewer is watching from and whether that viewpoint moves. You can explain this without memorising a list of film terms.

The examples in this guide give you several distinct choices:

ExampleWhat the viewpoint doesHow to describe that choice
Studio performanceRemains in front of the performer while he moves.Keep the camera in its opening position.
Street reunionFollows the friends as they come together.Follow their approach and embrace in one continuous shot.
Outdoor vlogTravels with the woman while keeping her close to the camera.Use a walking selfie view, with the street visible behind her.
Drifting carMoves from above and behind the car towards its side and front.Follow the reference video's changing viewpoint around the car.

These are alternatives, not a checklist of movements to put in the same prompt.

Separate the subject's movement from the camera's movement

In the drifting example, the car turns and the viewpoint also travels around it. “The car turns” only describes the first movement. It does not tell us where the camera goes.

Write each job separately when you need both. When using a video reference, explicitly say that it supplies the driving movement and the camera movement, rather than leaving the meaning of “follow this video” open.

Keeping a subject the same size in the picture also does not mean the camera stays still. In a walking selfie, the viewpoint travels with the person even though the face remains close. Ask for a stationary camera when its position must remain fixed.

Decide whether the view changes through movement or a cut

A continuous shot shows the journey from one view to the next. A cut switches directly to a different view.

The container-yard example later uses several views: the approach with the briefcase, the handover, the struggle, and the departure. When writing a sequence like that, describe the action in each view and state where the cuts should happen. Do not also request an uninterrupted take for the same sequence.

For pictures showing the difference between wide shots, waist-up views, and close-ups, see our camera shot and framing guide.

Use Omni Reference to give each upload a clear job

A reference video shows more than an action. It also contains a person, clothing, a location, lighting, and a camera view. Tell Kling which parts you need.

In the basketball example, the reference video shows a dunk on an outdoor court. The new version uses a different player inside an arena. Three inputs separate those decisions:

UploadWhat to use it for
Image 1: player sheetThe player in black sportswear. The different views show the same person.
Image 2: arenaThe indoor court, hoop, spectators, and lighting.
Video 1: outdoor dunkThe approach, jump, ball movement, and pace of the dunk.

Image 1 and Video 1 are names for the uploads in this explanation. Use the matching file names or reference labels in your interface.

Here is how to give these inputs separate jobs:

Use the player from Image 1, keeping his face, hairstyle, black sportswear, and body proportions. Place him on the indoor basketball court in Image 2.

Have him perform the dunk shown in Video 1. Follow the approach, take-off, movement of the ball during the jump, and finish at the hoop, using the action's pace from the video.

Take the court and surroundings from Image 2, not the outdoor location in Video 1. Keep the player, ball, and hoop visible during the jump and dunk.

The last paragraph resolves the conflict between the two locations. Without it, you have supplied an outdoor court and an indoor court without stating which belongs in the finished scene.

The framing request serves the action: readers need to see the ball reach the hoop. A face-only view would hide the part of the movement that makes this a dunk.

Separate the player, location, and movement

Image 1: player reference

Front, side, and face views of the same player in black sportswear.

Image 2: arena reference

An indoor basketball court with a hoop and spectators.

Video 1: outdoor dunk reference

Indoor version

Compare the approach, jump, and ball movement across the clips. The appearance and indoor court are defined by the two images.

Working with several uploads? Our PhotoGPT Flow guide shows how to connect reference images and generation steps on a canvas.

How to prompt a multi-character scene

When a scene has several people, explain who each person is before describing what they do. Then keep using the same name or role for that person. Repeatedly writing “the man” becomes confusing when three different men appear.

The reference sheet below is one image containing three portraits, two vehicles, and filming equipment. Its top row provides the three character appearances. Assign them to the roles shown in the accompanying sequence:

Portrait on the sheetRole to use in the promptWhere that person appears
Upper left: man in a dark suitDriverBehind the wheel of the red car, then in the final crouching pose.
Upper middle: man with headphones and a red lanyardHeadphone crew memberBeside the monitors during the crew cutaway.
Upper right: grey-haired man in a brown jacketOlder crew memberGesturing beside the monitors and the headphone-wearing crew member.

These are labels for the people, not extra uploads. Along the bottom, the blue motorcycle and red car are props; the three car views show the same vehicle. The equipment provides the appearance of the filming setup.

Give each character a role in the sequence

The portraits establish appearance, but they do not explain who drives, who stays beside the monitors, or when the camera returns to the driver. Put those decisions into the action description:

Use the attached reference sheet for the cast and props. The upper-left portrait is the Driver. The man with headphones and a red lanyard in the upper-middle portrait is the Headphone crew member. The grey-haired man in the upper-right portrait is the Older crew member. Use the blue motorcycle, red sports car, and filming equipment shown along the bottom row. The three views of the red car show one vehicle.

Create a fast-paced action sequence on a city square at dusk. Begin with close views of the red car's turning wheel, the pedals, and the Driver at the wheel. A helmeted camera operator films from the blue motorcycle as it moves alongside the car. Cut to the Older crew member gesturing beside the monitors, with the Headphone crew member next to him. Return to the moving red car as an explosion erupts behind it. Finish with the Driver landing in a low crouch on the paving, with the burning car behind him.

Keep each person's face matched to their own portrait when they reappear. Dress the Driver in a dark leather jacket for the action. Keep the Headphone crew member's headphones and red lanyard, and the Older crew member's brown jacket. Present each view as a full-screen shot, without the reference-sheet grid.

The role names connect the opening portraits to the later actions. “Return to the Driver” refers to the same person after the crew cutaway. It does not ask for all three people to remain visible in every shot.

Notice the clothing instruction, too. The driver's reference portrait shows a suit, while the action video uses a leather jacket. The prompt asks to retain his face while separately requesting the jacket. Asking for every detail of the portrait to remain unchanged would conflict with that choice.

The helmeted camera operator is a separate visible role in the clip. The sheet does not identify which face belongs to that operator, so do not assign the headphones portrait to the rider simply because both are involved in filming.

Give each reference character a distinct role

Reference sheet: three characters, vehicles, and equipment

One reference sheet with three male portraits along the top and a blue motorcycle, views of one red sports car, and filming equipment below.

Multi-character action video

The sheet shows three distinct character appearances. Follow their separate roles as the sequence moves from the driver to the crew and back to the driver.

Match each replacement character to a specific performer

The fight example contains two character sheets and a video with two performers. Each sheet shows one person from several angles. It does not introduce several different people.

You also need to say who takes which role. Use a visible feature and an opening position from the reference clip so the assignment remains understandable once the action speeds up.

Image 1 defines the silver-haired fighter in black shorts and white gloves. Image 2 defines the short-haired fighter in green sportswear and pink hand wraps.

In Video 1, replace the performer in red shorts who starts closest to the camera with the fighter from Image 2. Replace the opponent in dark trousers farther from the camera with the fighter from Image 1.

Follow the exchange of movements and its timing from Video 1. Set the fight on a wooden platform surrounded by forest. Keep each replacement fighter's appearance tied to her own reference sheet.

The clip supplies the roles and movements. The sheets supply the replacement appearances. The forest setting is a written instruction here; this set does not include a separate forest-reference image.

Assign each fighter to the correct performer

Image 1: silver-haired fighter

Multiple views of a silver-haired fighter in black shorts and white gloves.

Image 2: short-haired fighter

Multiple views of a short-haired fighter in green sportswear and pink hand wraps.

Video 1: ring-fight reference

Forest version

The performer nearer the camera at the start becomes the short-haired fighter. The other role becomes the silver-haired fighter.

Use a video reference for a complex camera move

The drift example adds another job for the reference video: it guides the viewpoint as well as the vehicle.

The daytime clip begins above and behind a light-coloured car. The view follows the drift, moves closer to its side, and reaches the front. The two images show a replacement racing car and a neon-lit arena.

Use Video 1 for both the car's drifting movement and the camera's route. Use Image 1 for the dark racing car with red and blue graphics. Use Image 2 for the neon-lit arena and its lighting.

Recreate the drift in that arena. Follow the camera from the higher rear view towards the side of the turning car, then around towards its front. Keep the replacement car's shape and graphics consistent.

Take the environment from Image 2 rather than the daytime track in Video 1.

“Both the car's drifting movement and the camera's route” is the important instruction. You are asking the video to guide two related movements, while the images determine the appearance and location.

The route is communicated by the moving viewpoint in the reference clip. Naming its role is more useful here than adding a long list of camera terms.

Let the video guide the drift and the viewpoint

Image 1: racing-car reference

Four views of a dark racing car with red and blue graphics.

Image 2: arena reference

A circular neon-lit driving area surrounded by containers and lights.

Video 1: daytime drift reference

Neon-arena version

The video provides the moving reference. The still images show the replacement car and the environment; neither is an arrow diagram.

Turn a rough 3D preview into a finished scene

A white-model video is a simplified 3D preview. Its figures and surfaces are unfinished, but it can already show the action, camera views, and timing.

The container-yard preview has two figures and a briefcase. One figure carries the case into the scene. A handover becomes a struggle; the case falls; the carrier retrieves it and leaves. The character sheets provide the two finished appearances.

Use Video 1 for the container-yard layout, briefcase action, camera views, and timing.

Replace the figure who arrives carrying the briefcase with the woman in Image 1. Replace the other figure with the man in Image 2. Keep their appearances tied to their respective sheets.

Follow the approach, attempted handover, struggle, dropped case, and departure shown in the preview. Render the scene with detailed people, clothing, and metal containers instead of the simplified white surfaces.

The opening role identifies which figure becomes the woman. That is clearer than “replace one figure,” because both people later interact with the case.

The preview also shows changes of view. Following that sequence does not mean asking the camera to travel continuously between every angle.

Use the preview for action and the sheets for appearance

Image 1: briefcase carrier

Front, side, and face views of a woman in light athletic clothing.

Image 2: other character

Front, side, and face views of a muscular man in dark shorts.

Video 1: rough 3D preview

Finished-scene version

Compare the case carrier, handover, struggle, and departure. The preview supplies the planned views; the sheets supply the two appearances.

Depth-map videos are another supported structural input. They represent scene depth rather than finished colours. When you use one, distinguish that structural guidance from the image supplying the final appearance. The container-yard demonstration above uses a white-model preview, not a depth map.

Connect first and last frames by describing the change

First & Last Frames lets you choose the opening and ending pictures. Your prompt still needs to explain what happens between them.

Before writing, compare the two images. Check the subject's position, pose, surroundings, camera view, and lighting. Decide which differences are intentional. Then describe what causes those differences during the clip.

Use this as a planning check:

What to decideWhat belongs in the prompt
What movesName the subject, camera, or object responsible for the change.
What happens between the imagesDescribe the connecting action in order.
How the view changesState whether the camera travels continuously or the sequence uses a cut.
What stays consistentName the identity, objects, or scene details that should carry through.

Check for contradictions between the endpoint images and your written instructions. When a difference is intentional, describe how it happens rather than asking every visible detail to remain unchanged.

Use these decisions to write the transition for your own two images. Keep the endpoint requirements separate from any intermediate moments you intend to guide with additional keyframes.

Use keyframes for the moments that need a specific image

A keyframe shows how an important moment should look. Kling 4.0 supports up to ten keyframe images, including moments between the opening and ending.

The glowing-room overview shows five visual checkpoints: an illuminated palm, symbols on a monitor, an octopus beside a bed, a burger in a refrigerator, and an eye on a wall. The overview is one image containing five views. It should not be mistaken for five separate uploaded files.

For your own keyframe workflow, prepare the required moment images and put them in the intended order. Then describe the connections between them.

Follow the visual moments in this order: the glowing drawings on the open palm, the monitor with a red heart and blue spiral, the pink-purple octopus beside the bed, the cartoon burger in the open refrigerator, and the golden eye on the wall.

Let the viewpoint travel through the room as the glowing shapes lead it from one moment to the next. Begin close to the palm and finish facing the eye on the wall. Keep the warm room lighting and the luminous illustrated appearance throughout.

The images specify the main appearances. The video also contains connecting animation: a star travels away from the palm, and glowing forms carry the viewer towards later parts of the room. These connections are not all pictured in the overview.

Decide how much of that connecting action you need to describe. You can require important transitions while leaving minor movement open. Avoid assuming that a sheet of key moments already specifies every action between them.

Five key moments, shown together in one overview

Overview image containing five views

One composite showing a glowing palm, monitor symbols, octopus, fridge burger, and eye on a wall.

Glowing-room video

The overview shows the main appearances. The video also contains animated connections between those moments.

Use a storyboard to plan the sequence

A storyboard arranges planned moments and camera views into panels. A panel can represent a new shot, but it can also show the next moment within a continuous shot.

The cheese-moon example has two distinct inputs. One is Biscuit's character sheet: a clay dog in a white spacesuit with a giant yellow fork. The other is an eight-panel storyboard.

The storyboard follows Biscuit from the moon's surface into a cheese tunnel, past a mouse, and back to the surface. Its written plan calls for continuous camera coverage. Eight panels therefore do not mean eight cuts.

Use Image 1 for Biscuit's appearance: the clay dog, white spacesuit, and giant yellow fork. Use Image 2 for the story and camera plan, following panels P01 through P08.

Show Biscuit investigating the cheese moon, leaping with his fork, opening a hole, and tumbling into the tunnel. Follow the encounter with the mouse and the return to the surface. Finish with Biscuit and the mouse sitting together eating cheese, then pull back to reveal more of the moon.

Connect the storyboard moments in one continuous camera sequence. Keep the clay-animation appearance. Present the story as a full-screen scene, leaving out panel borders, panel numbers, and production notes.

The first paragraph separates character design from story order. The last explains how the panel sheet should become video rather than remain a visible grid.

When you review a result, compare it with the plan. Matching the main characters and events does not by itself show that every planned camera connection was followed. A storyboard expresses the intended sequence; it is still something to check against the video.

A character design and an eight-panel camera plan

Image 1: Biscuit character sheet

Biscuit, a clay dog astronaut in a white suit, carrying a giant yellow fork.

Image 2: storyboard

Panels P01 to P08 plan Biscuit’s cheese-moon journey and continuous camera coverage.

Cheese-moon video

The sheet defines Biscuit. The panels plan the story and request continuous coverage; compare that intention with the video rather than assuming every transition matches.

Plan cuts and timing before adding more shots

For a cut-based sequence, a written shot list can describe the same information without drawings: the view, the action in that view, and the point where the next view begins.

The container-yard preview makes the distinction useful. Its close view of the hands serves the briefcase exchange, while the wider view reveals the struggle. Each view needs a reason to be there. Adding another angle simply because there is room for it does not add another story event.

Kling 4.0 supports 3–30-second generations. Choose the duration around the action and any script. Read the words aloud and allow time for movement and reactions. A number written beside a shot communicates your requested timing; inspect the finished video rather than assuming a frame-accurate cut.

Give speech its own instructions

The reunion prompt requests English dialogue but leaves the actual words open. To use a fixed script, identify who speaks and put their words beside that description.

The reunion gives us two visually distinguishable people: one in a blue jacket facing the camera during the approach, and one in a black jacket initially facing away. A script can use those descriptions without inventing character names.

Replace the brackets with your own lines. This is a writing template, not a transcript of the clip.

Friend in the blue jacket, speaking in English: "[First line]"
Friend in the black jacket, replying after the first line: "[Reply]"

Put a pause, interruption, or change of delivery beside the line it affects. Language chooses what is spoken; an accent and delivery direction describe how it should sound. Neither substitutes for the words themselves.

Keep appearance references separate from voice references

A portrait can guide a person's appearance. It does not supply a recording of their voice.

When you upload a voice reference, state which character should use it. Write the required words separately. The image guides appearance, the audio reference guides the voice, and the script supplies the words.

For speech heard over a scene rather than spoken by a visible character, label it voice-over. Then state whether the on-screen characters should remain silent.

Connect sound to the scene

For the reunion, footsteps belong to the approach, while street ambience continues around the encounter. The prompt's background music is a separate layer from those location sounds.

When adding sound instructions, specify which sound accompanies which action. State when the dialogue needs to remain clearer than the music. Then listen to the result: check the complete lines, speaking order, and whether the sounds occur with the intended actions. A still frame cannot answer those questions.

Describe the text the viewer should read

Write the exact words and decide where they belong. Text can fill the screen, sit over a moving picture, or appear on a physical object within the scene.

The typography video cycles through several designs. Look specifically at the “DREAM BIGGER” card around 5.5–6 seconds, shown below. It uses blue and yellow lettering on cream, a blue oval, and a small black star.

For that card's appearance, write:

Set "DREAM" in large blue capital letters above and slightly to the left of "BIGGER" in yellow capitals. Use heavy black outlines and offset black shadows on a cream background.

Add a blue oval around the words and a small black four-point star near its upper-right edge. Put "SAME PEOPLE BRIGHTER DAYS" in small stacked black lettering at the lower left. Keep the main phrase fully visible.

This describes one card within the clip, not the whole sequence of titles. For a longer sequence, provide the wording and order of each title you need rather than expecting a single phrase to define the others.

Inspect one title within the larger typography sequence

The prompt discussion concerns the DREAM BIGGER card. The full clip contains additional titles that would need their own instructions.

Put product lettering on the product

In the orange-drink example later in the guide, “KLING POP” belongs to the bottle. Its position changes with that object. In a full-screen title, the letters belong to the graphic composition instead.

State that distinction when writing. Use the supplied product image for its existing label design, and identify the intended surface when adding separate artwork. Check the lettering during movement, not just in the opening frame.

For an opening image with a headline already in place, use our guide to adding text to images to set its wording and position before animation.

Describe lighting and materials through visible details

The outdoor vlog uses warm sunlight at the start, then softer light as the woman continues along the street. Its close selfie view moves with her. This is a different treatment from the fixed studio viewpoint.

To describe that look:

Create a walking selfie on a city street. Keep the woman close to the camera, with the street visible behind her. Begin with warm sunlight on her face, then let the light soften as she walks into shade. Keep natural skin texture and gentle movement with her steps.

The changing light belongs to the walk. Asking for identical brightness on her face throughout would fight against the transition into shade.

A walking selfie moves through changing light

The viewpoint travels with the woman while the street remains visible. The light on her face becomes softer during the walk.

For material direction, return to Biscuit's character sheet and cheese-moon video. The rounded clay forms, sculpted fur, and thick spacesuit are visible design choices. Describe those properties when you need that handmade appearance instead of adding unrelated style words.

A material instruction for that character can stay focused:

Keep Biscuit's hand-shaped clay appearance, including the sculpted fur, rounded paws, thick white spacesuit, and yellow fork.

Set resolution through the output controls. Describe light, texture, and materials in the prompt. “4K” does not explain where the light comes from or what an object is made of.

For more portrait examples, our lighting prompt guide shows how to describe soft window light, warm evening light, and stronger shadows.

Recreate a video around a different product

The reference advertisement shows a woman presenting a coffee can towards the camera, with a brown liquid splash. The replacement image shows an orange-drink bottle photographed in a coastal setting.

The new version keeps the café scene and changes the featured drink. That requires separating the product from the background in its photograph.

Use Video 1 for the woman, café setting, product presentation, and camera movement. Replace the coffee can she holds towards the camera with the orange-drink bottle from Image 1.

Keep the bottle's shape, orange appearance, and "KLING POP" label design. Show the bottle open during the liquid splash. Change the brown splash around the featured drink into an orange-coloured splash.

Follow the reference video's camera movement around the woman. Keep its café setting and surrounding props. Use Image 1 for the bottle, not for its coastal background.

Notice the cap. The product photo shows a capped bottle; the video presents it open during the splash. Preserving a product's design does not mean freezing it in the exact state shown in a photograph.

The target is also specific: the can held towards the camera. The instruction does not ask to replace every cup or drink elsewhere in the café.

Replace the featured drink and the splash

Image 1: replacement-product reference

A capped orange KLING POP bottle photographed with oranges and a coastal background.

Video 1: coffee-can reference

Orange-bottle version

The product photograph is capped, but the video shows the bottle open during the splash. The café comes from the reference video, not the product photograph.

Edit the target without redefining the whole video

In the street-band example, several pedestrians cross between the camera and the musicians. The edit clears those foreground people while keeping the band and the people farther back in the square.

Identify the target by its relationship to the scene:

Edit Video 1. Remove the foreground pedestrians crossing between the camera and the three street musicians. Fill the areas they cover with the appropriate continuation of the square and the band.

Keep the three musicians, their instruments, the microphone, the fountain, and the people behind the band. Preserve the camera's approach and the sequence of the performance.

“Foreground pedestrians crossing between the camera and the musicians” is more precise than “remove the people.” The broader version could include the performers or the crowd you want to retain.

The second paragraph states what should stay. It is a review checklist, not a guarantee that those details will be untouched. Compare the same moment in both videos, especially while the passers-by cover the musicians.

Clear the foreground and retain the band

Before: pedestrians obscure the musicians

After: foreground cleared

The intended target is the people between the camera and the band. Compare matching elapsed times to inspect what remains.

Kling 3.0 vs. Kling 4.0: what changes in your prompts?

Keep describing subjects, action, camera movement, appearance, and sound in ordinary language. The useful differences concern the length of a sequence and the inputs you can use to explain it.

What you are planningKling 3.0 starting pointHow to use the capability in Kling 4.0
A longer performanceA 3–15-second generation range.Plan within 3–30 seconds, allowing the action and dialogue to finish.
An important intermediate momentStart-and-end-frame guidance.Use up to ten keyframe images when moments between the endpoints need a specified appearance.
Several camera viewsMulti-Shot and Custom Multi-Shot were already available.Continue deciding which views need cuts and which belong in one continuous shot.
Separate appearance and movement3.0 Omni already combined image, video, and subject references.Assign explicit roles to the inputs, including structural video references for action and camera planning.

Native audio and visible text were already part of the earlier model's capabilities. They still need creative direction: the right speaker, actual script, intended text placement, and enough time to understand the result.

Make the next revision about the part that missed

Describe the difference between your intention and the video before changing the prompt. That gives you a specific instruction to revise.

What needs correctingWhat to clarify
The dunk uses the outdoor location instead of the arena.Image 2 supplies the court; Video 1 supplies the movement.
The replacement fighters take the wrong roles.Assign each one to an identifiable performer at the start of the reference clip.
The storyboard remains visible as a grid.Its panels guide a full-screen scene; borders and production notes are not part of the result.
The replacement bottle stays capped during the splash.Preserve the product design but show it in the open state needed for that action.
The edit removes people behind the band as well.Limit removal to the pedestrians between the camera and the musicians.

These are focused instructions to try, not diagnoses of a guaranteed cause. Keep the other inputs and settings the same for the next attempt so you can compare the change.

Check sound by listening, check words while they move, and compare edited footage at matching times. Save the prompt, inputs, settings, and result together. That gives your next revision an actual record to work from.

Have a scene in mind? Create your video with Kling 4.0. Start with the example closest to your idea and replace the scene details with your own.

Sources

Kling AI: Meet All-New Kling 4.0

Last updated on

On this page