PhotoGPTGuides

Using Audio Input for Video Generation Models

Use audio references in PhotoGPT to control voice, music, rhythm, ambience, and sound design in H3, Seedance 2.0, and Seedance 2.5 videos.

PhotoGPT can take an audio file as a reference for video generation. The uploaded audio is not the final soundtrack. It is a control signal that guides what the generated audio should sound like: a speaker's voice, the rhythm of a music track, the mood of an ambience bed, or the texture of a sound effect.

The workflow is the same for every model, but the prompt syntax and limits differ.

  1. Select a video model that supports audio references (H3, Seedance 2.0, or Seedance 2.5).
  2. Click the Reference Audio icon in the prompt panel.
  3. Upload one or more audio files. The dashboard checks duration and file count automatically.
  4. Write the prompt and tell the model what the audio is for. Use <Audio 1> for H3 and @Audio1 for Seedance.
  5. Turn on Generate Audio if you want sound (Seedance only). It is on by default. Switch it off for a silent video.
  6. Generate the video and review the result.

When first or last frame mode is active, reference audio is disabled. Switch to reference mode if you want to use audio.

Which models support audio input

ModelAudio refsPer-clip durationTotal durationAccepted formatsAudio-onlyOutput durationOutput audio
MiniMax H3Up to 32–15 sUp to 15 sWAV, MP3No, needs image or video4–15 sNative stereo 32 kHz
Seedance 2.0Up to 32–15 sUp to 15 sMP3, WAVNo, needs image or video4–15 sNative
Seedance 2.5Up to 102–30 sUp to 30 sMP3, WAV, M4AYes4–30 sNative

What audio input does and does not do

In the PhotoGPT video generator, Reference Audio is an input file you upload to shape the output audio. It does the same kind of job as a reference image or reference video: it carries information the prompt would otherwise have to describe in words.

What the audio can guide:

  • Voice and timbre: a character sounds like the clip.
  • Dialogue line and delivery: the model can lip-sync to a line you write in the prompt, while the audio clip steers the pace and tone.
  • Music style and rhythm: the generated soundtrack can follow the tempo, instrumentation, and energy of the reference.
  • Ambience and texture: wind, room tone, crowd noise, or abstract sound design can set the mood.
  • Sound effects: the texture and timing of impacts, animal sounds, or mechanical noises.

What the audio is not:

  • It is not a track that gets pasted onto the final video. The model reads the audio and regenerates a new audio output.
  • It is not a guarantee of exact lyrics, melody, or speech. The prompt still needs to describe the desired dialogue, music, or SFX.
  • It is not a substitute for image or video references when a model needs them.

Reference audio is separate from the Generate Audio switch that appears for some models. The switch controls whether the model produces a soundtrack at all. The reference audio controls what that soundtrack should sound like.

Picking the right model for the job

Use caseBest modelWhy
Mobile app UGC adH3 or Seedance 2.5H3 gives strong lip-sync and 2K output; Seedance 2.5 supports longer takes and audio-only references.
Product ad with voiceoverH3Locks product detail, typography, and camera motion at up to 2K.
Music videoH3 or Seedance 2.5H3 can follow a vocal performance; Seedance 2.5 handles longer music-driven shots.
Sound-driven dance and motionSeedance 2.530-second duration and strong beat-to-motion control.
Acoustic music performanceH3 or Seedance 2.5H3 keeps vocal and instrument detail; Seedance 2.5 supports longer takes.
Multi-character dialogueSeedance 2.5Up to 10 audio references let you assign a voice to each character.
Dubbing and translationH3 or Seedance 2.5Both can replace the audio of an existing video; Seedance 2.5 gives more headroom for longer clips.
Meme, reaction, and UGCSeedance 2.0 or 2.5Fast, lo-fi, platform-native output.

Use cases

Each example below shows how to give the audio reference a clear job. The prompt snippets are simplified; the H3 and Seedance 2.5 guides contain the full detailed syntax.

Mobile app UGC ad

A talking-head UGC ad for a mobile app uses one image of the app, one image of the reviewer, and one audio clip of the reviewer’s voice.

@Image1 is the mobile app screenshot: a clean finance-app dashboard with a blue card, transaction list, and account balance.
@Image2 is a medium close-up of a friendly young man in a black hoodie and beanie, sitting in a car, speaking straight to camera.
@Audio1 is his casual, fast-talking UGC voice.

Generate a 10-second vertical UGC ad. The man in @Image2 holds a phone showing the app in @Image1 and talks to camera with the same delivery as @Audio1. He says: {This app literally sorted my entire budget before I finished my coffee.} Keep it handheld, selfie-style, lo-fi. No music, no subtitles.

Model tip: This example comes from a Seedance 2.0 workflow. H3 and Seedance 2.5 can also produce talking-head UGC with a voice reference.

Product ad with voiceover

The full sleep-gummy ad was built from several shorter clips and joined in an editor. The prompt below is for one 10-second segment.

@Image1 is the sleep gummy bottle on a warm bedside table: clear amber glass, purple label, soft morning light.
@Image2 is the woman in bed, dark hair, olive-green top, relaxed expression, warm bedroom.
@Audio1 is the calm, reassuring male voiceover.

Generate a 10-second product ad. Shot 1: tight product shot of @Image1 on the bedside table. Shot 2: the woman from @Image2 wakes slowly, smiles, and reaches for the bottle. The voiceover from @Audio1 says: {Finally, a gummy that helps you fall asleep and stay asleep.} Soft bedroom ambience. No music.

Model tip: H3 is a strong choice for product detail and clean voiceover. Seedance 2.5 is better for a longer multi-shot ad.

Music video

The full music video was made from several clips. The prompt below is for one 10-second segment.

@Image1 is the main performer in a dark gym, wearing a black robe and gold boxing gloves, industrial lights behind him.
@Image2 is a close-up of the same performer, muscular build, gold boxing shorts, intense expression.
@Audio1 is the vocal performance of the song.

Generate a 10-second music video clip. The performer in @Image1 shadowboxes to the beat of @Audio1. At 00:04 the camera cuts to the close-up in @Image2. He sings with precise lip sync: {I come in first, that's the way it goes.} Gold lighting, dramatic, energetic. No subtitles.

Model tip: For music videos, write the lyrics in the prompt and use the audio as the vocal reference. Seedance 2.5 handles longer takes; H3 gives strong lip-sync for short performances.

Sound-driven dance and motion

A dance track can drive the camera, the steps, and the lighting.

@Image1 is the dancer: a young woman with red-tinted hair, silver puffer jacket, black crop top, black cargo pants, white sneakers.
@Audio1 is an upbeat electronic dance track.

Generate a 15-second music-driven dance clip. The dancer from @Image1 performs in an empty warehouse with blue and pink neon tubes. Her steps, arm movements, and hair flips land on the beat of @Audio1. The camera stays wide and follows her motion. Match the energy and rhythm of @Audio1.

Model tip: This example was made with H3's audio-to-motion feature. Use H3 for beat-synced motion, or Seedance 2.5 for longer references and audio-only input.

Acoustic music performance

A vocal and guitar performance can be synced to an original song. The clip below is 17 seconds, so it needs Seedance 2.5 or several shorter clips on H3 and Seedance 2.0.

@Image1 is an old photo of the musician: a man in a black T-shirt with a white maple-leaf emblem, playing a weathered acoustic guitar.
@Audio1 is the original acoustic song: finger-picked guitar and soft male vocals.

Generate a 15-second acoustic performance clip. The musician in @Image1 plays the guitar and sings along with the melody and timbre from @Audio1. Close-up on his hands and the fretboard. Keep the white background, soft natural light, and an unplugged feel. No extra music.

Model tip: Seedance 2.5 supports up to 30 seconds, so a 17-second performance fits in one clip. H3 and Seedance 2.0 work for shorter takes up to 15 seconds.

Multi-character dialogue

You can assign one audio reference per character so each speaker keeps a distinct voice.

@Image1 is a naturalistic photo of two siblings on a concrete rooftop at dusk, moving boxes nearby, city skyline behind them.
@Audio1 is the brother's warm, calm voice.
@Audio2 is the sister's gentle, clear voice.

Generate a 15-second two-character scene. The brother and sister from @Image1 sit on the rooftop ledge at golden hour. The brother says: {How many sunsets you think we watched from here?} with the voice from @Audio1. The sister replies: {Not enough.} with the voice from @Audio2. Subtle wind and distant city traffic. No music.

Model tip: Seedance 2.5 supports up to 10 audio references, which is useful when many characters speak.

Dubbing and translation

Give the model the original video and a new audio reference in the target language.

@Video1 is the source English video of a 3D animated man in a modern office. Do not use the original English audio.
@Audio1 is the warm, clear Chinese voiceover.

Generate a 15-second Chinese dub of the video. The man in @Video1 stands in the office and speaks with precise lip sync. Dialogue language: Chinese. He says {欢迎回来。今天我想和你谈谈这个季度的计划。} Preserve the camera and lighting from @Video1. No background music.

Input - English

Output - Chinese

Model tip: The key is to tell the model to translate the speech, resync the lips, and keep everything else unchanged.

Meme, reaction, and UGC

A short vocal clip can drive the energy and timing of a caricature or reaction clip.

@Image1 is the character: a stylized man with an exaggerated face, tall blonde bouffant hair, wearing a striped shirt and blazer.
@Audio1 is the funny, exaggerated voice clip.

Generate a 7-second meme clip. The character in @Image1 talks directly to camera with the wild delivery from @Audio1. He says: {You ever wake up and your hair already made plans without you?} Keep the background a flat pastel blue, simple studio lighting. No music, no subtitles.

Model tip: Seedance 2.0 and 2.5 are both good for lo-fi UGC and meme clips. Use Seedance 2.5 if you want a longer take or more control.

Audio file prep checklist

  • Format: WAV or MP3 works for all three models. Seedance 2.5 also accepts M4A. The PhotoGPT upload dialog accepts any audio/* file.
  • Duration: each clip must be inside the model's per-clip window (2–15 s for H3 and Seedance 2.0, 2–30 s for Seedance 2.5).
  • Combined duration: the total of all uploaded clips cannot exceed 15 s for H3 and Seedance 2.0, or 30 s for Seedance 2.5.
  • Number of clips: up to 3 for H3 and Seedance 2.0, up to 10 for Seedance 2.5.
  • File size: keep individual clips under 15 MB and the whole request under the dashboard limit.
  • Quality: clean, uncompressed, no clipping. Trim leading and trailing silence.
  • Rights: use only audio you have the right to reference.
  • Role clarity: decide whether the clip is for voice, music, rhythm, or ambience, and say so in the prompt.

Common mistakes

  • Expecting the audio to be copied exactly. The model uses the clip as a reference. Always write the desired dialogue, music, or SFX in the prompt.
  • Not naming the audio in the prompt. Uploading the file is not enough. You must reference it as <Audio 1> (H3) or @Audio1 (Seedance) and give it a job.
  • Uploading a clip that is too long or too large. The dashboard will grey the chip out and show a warning.
  • Mixing first/last frames with reference audio. Reference audio is part of reference mode. Switch away from first/last frame mode if you want to use it.
  • Using audio-only on H3 or Seedance 2.0. These models require an image or video reference alongside the audio.
  • Forgetting the Generate Audio switch on Seedance. If Generate Audio is off, the output will be silent even when you upload a reference audio.
  • Heavy compression or clipping. The model reads timbre and rhythm more easily from a clean clip.

Prompt syntax cheat sheet

ModelReference labelDialogueMusicSFXSubtitles
H3<Audio 1>, <Picture 1>, <Video 1>, <Subject 1><d>[English] ...</d>non_diegetic_music fieldoverall_soundscape fieldNot supported
Seedance 2.5@Audio1, @Image1, @Video1{ ... }( ... )< ... >【 ... 】
Seedance 2.0@Audio1, @Image1, @Video1{ ... }( ... )< ... >【 ... 】

For H3, the audio reference belongs inside the full reference-mode prompt. For Seedance, the bracket syntax is optional but helps the model separate music, SFX, dialogue, and on-screen text.

Next steps

Last updated on

On this page