Use HeyGen Video to create a short scene, then decide separately how its generated sound belongs with your finished song. If the song is already mixed, treat that recording as the soundtrack reference. A clip arriving with dialogue, ambience, or effects is not automatically ready to drop beneath a vocal.
HeyGen describes heygen-video-1 as a general-purpose model built on MiniMax H3 and post-trained by HeyGen. It supports text-to-video, image-to-video, and reference-to-video, with generated sound. This is distinct from HeyGen's avatar-rendering engines. Official model documentation.
This guide develops an original railway-platform scene and a practical review process. It is based on documentation, not a paid generation test, and does not claim a built-in HeyGen integration in LyricVideoMaker.
Choose what you want to supply, not just a model name
Decide which part of the shot is already settled. A written idea, an approved opening image, and a reference for a recurring object give a generator different information.
HeyGen's documentation describes the modes this way:
| Starting material | Mode | Main purpose |
|---|---|---|
| A written scene description | Text-to-video | Generate a new scene |
| One image plus a prompt | Image-to-video | Begin the clip on the supplied image |
| Reference images, videos, or audio plus a prompt | Reference-to-video | Guide the new scene with supplied material |
The documentation describes 5–15-second clips. Check your selected interface for the supported fields and output settings before preparing the request. Mode and duration details.
For a lyric-led video, you might already have an approved image with a quiet area for words. Starting from it lets you judge whether the motion preserves that useful composition. If all you have is a written concept, do not pretend that a reference image exists: start with the text route or prepare the image first.
A reference is also not proof of continuity. Compare the returned subject, clothing, and important objects with the approved material before accepting a shot.
Decide whether the scene should be heard
For our fictional scene, a traveler waits on an empty railway platform after rain. A train's lights become visible in the distance. The song already contains a delicate vocal and a quiet instrumental arrangement.
There are several possible sound decisions:
| Editing intention | What you want from the scene | What to check afterward |
|---|---|---|
| Let the song carry everything | Visual atmosphere only in the final edit | Mute clip audio and verify the song is unchanged |
| Add a little space around the song | Restrained environmental sound | Listen for masking of soft words or sustained notes |
| Make an instrumental break feel physical | A deliberate sound event | Check its level, timing, and relationship to the following vocal |
| Include a spoken line | Dialogue with a clear purpose | Confirm it does not interrupt lyrics or imply an unintended speaker |
These are editorial choices, not guaranteed generation modes. For a first test, choose one. Asking for dramatic rain, station announcements, a train horn, and orchestral music can leave you with a soundscape that has no room for the original track.
You can request a sound treatment in the prompt, but listen to the result. When the final video should use only the song, muting the clip's audio in the editor is the direct way to enforce that decision.
Write the visual action and the sound brief together
Use one principal event so you can tell whether the clip works. In this example, the distant lights grow brighter while the traveler stays still. There is no close-up singing or complex hand action to judge at the same time.
An original draft prompt:
A wide view of a nearly empty railway platform after rain, before sunrise.
One adult traveler in a dark blue coat stands beside a small gray bag.
The camera remains stationary beneath the platform canopy. The traveler
looks toward the far end of the track as two distant train lights become
brighter. Keep the traveler, bag, and platform layout consistent.
Leave a quiet dark area of the canopy in the upper part of the composition.
No readable signs or added lettering.
Sound: soft water dripping from the canopy and a distant low train rumble.
No announcements, spoken dialogue, singing, or background music.
The sound should remain restrained rather than building into a dramatic score.
Set the available duration and frame shape through the actual controls. This is a proposed creative brief, not a tested preset or special API syntax.
HeyGen's prompting guidance specifically recommends describing room tone and effects and requesting no music when a generated bed is unwanted. It also identifies close hand work and long on-screen text as less reliable cases. Prompting guidance.
The example avoids those extra tasks deliberately. Add the real lyrics during editing instead of asking the scene generator to draw several changing lines into the image.
Review the clip twice before placing it in the song
First review the picture with its generated sound turned down. The question is whether the visible action works: do the lights move plausibly, does the bag remain the same object, and is the traveler recognizable throughout?
Then listen to the clip's sound on its own. Note any extra music, speech, sudden loud effect, or abrupt audio ending. Do not assume that an appealing image makes the audio useful.
Finally audition the candidate against the intended section of the song. For the platform scene, check the quietest vocal phrase, not only the loud chorus. A low rumble that feels subtle alone may interfere with the song's bass, while sharp drips may distract from consonants.
| Problem | Smallest useful response |
|---|---|
| Picture works; sound is unnecessary | Mute the clip audio in the edit |
| Ambience helps but overwhelms the vocal | Reduce or remove that layer and compare again |
| The visual event arrives too late | Test trimming or placement before regenerating |
| A new music layer appears | Remove it from the edit; clarify the sound request if you need another generation |
| Important text space becomes busy | Adjust the composition or lyric placement |
| A recurring object changes | Review reference clarity and simplify the action |
Keep an accepted candidate before trying a variation. Change the problem you can name rather than rewriting the entire scene every time.
Keep song timing separate from clip duration
Suppose you choose a ten-second passage beginning at song time 0:40. A useful visual event four seconds into the clip belongs at 0:44 when the clip is placed without trimming or speed changes. That is a placement calculation, not a promise that the model will produce the event at exactly four seconds.
Review the actual returned file. If you trim its opening, record the new position of the visual event before aligning it with the song. Our timestamp-planning guide explains the difference between song time and clip time; its Seedance-specific generation controls do not apply to HeyGen.
If you need a visible singer matched to an existing vocal, evaluate that as a separate synchronization task. Do not assume a general scene-generation model behaves like a dedicated lip-sync service. See the music-video lip-sync guide for preparation and review checks.
Read the quote for the service you actually use
An official launch promotion, a provider's list price, and a third-party application's credit charge can be different numbers. Before a paid generation, inspect the selected mode, resolution, duration, and displayed charge in that service.
Record the amount charged and whether you kept any of the output. That gives you a useful basis for judging your own workflow. A temporary promotional rate or one creator's generation time is not a reliable estimate for every future shot.
This article does not reproduce a fixed price table because the creative decision does not require one. You need to know whether a particular clip solves your scene problem and what your current account will charge for the attempt.
Assemble the scene around the approved recording
After accepting the clip, review the transition into the next shot and check the displayed lyrics at the intended viewing size. Preserve the original song file so a new generated sound layer does not quietly become the release soundtrack.
LyricVideoMaker provides an audio-upload and editable lyric-video workflow. The HeyGen process here is external; use its output only through the tools and import options your editing workflow actually supports. For the full production sequence, see turning a song into a music video.
Should every music video scene have audible effects?
No. A silent visual layer over the song can be the right choice. Add effects when they contribute something specific, and compare the result with the song alone.
Does reference-to-video guarantee an identical character?
Treat it as guidance and review the result. Check identity, clothing, and the details that connect adjacent shots rather than accepting the mode name as a quality guarantee.
What makes a good first test?
Choose one bounded scene with one visible event, then review picture, generated sound, and the final mix separately. For the railway example, get the distant lights and the quiet vocal passage working together before building the rest of the sequence.
