LyricVideoMaker
← All articles
Tutorials

MiniMax H3 Audio References for Music Videos: What to Prepare

Prepare a song excerpt and visual references for MiniMax H3, distinguish audio guidance from soundtrack reuse, and check timing before building a full music video.

LyricVideoMaker Team

For a MiniMax H3 music video, prepare a short excerpt of the final recording, pair it with the visual input your chosen workflow requires, and explain the audio's role. A reference can communicate sound or performance information; it is not proof that the output preserves your master recording or places every action on the intended beat.

MiniMax describes H3 as a multimodal system working with text, images, video, and audio, with native stereo sound generation. That makes the relationship between the supplied audio and the desired output an important part of the brief. Official H3 introduction.

This guide is based on official documentation and editorial planning. The examples are original exercises, not paid-generation results or a measured synchronization test.

Check the input mode before preparing the song

The official open-source specification distinguishes first/last-frame generation from Ref2VA, the reference-to-audio-video mode. For Ref2VA, it lists up to three audio clips, each 2–15 seconds long and totaling no more than 15 seconds. Audio must accompany an image or video; it cannot be the sole input. Mixed reference inputs are limited to 12 files. These are the documented checkpoint conditions, not a guarantee that every hosted interface offers the same controls. Official H3 input specifications.

Before exporting assets, inspect the interface you will actually use:

  • Does it accept reference audio for the selected H3 workflow?
  • What visual material must accompany it?
  • What duration and file constraints does it display?
  • Does it describe the audio as a reference, a source to reuse, or something else?

A prompt that mentions an audio filename does not attach the file. Likewise, a video tool with an H3 label may expose only a subset of the underlying model's workflows. Resolve that before spending time preparing a large reference pack.

Decide what the reference audio should contribute

Write the audio's job in one sentence. This makes it easier to spot a result that is attractive but wrong for your song.

Intended job What to explain What still needs review
Guide the energy of a scene Which musical change should influence the visual action Whether the action develops at a useful pace
Guide a visible performance Which excerpt contains the performance you want represented Mouth movement, phrasing, and expression
Reuse source sound where the workflow supports it Which supplied track or passage should remain in the output Whether the returned sound actually preserves it
Supply a sound characteristic Which aspect matters, such as a voice quality or ambience Whether the output introduces unwanted new material

These are ways to describe intent, not four promised product modes. MiniMax's introduction describes reference and editing relationships through natural language. Your provider's actual input and output controls determine what you can request. H3's reference design.

For a finished music release, keep the approved song file available independently. Do not replace it with the audio from a generated clip without listening and deciding that the change is wanted.

Export an excerpt with a clear beginning and ending

Use the selected final recording, not an earlier version with similar lyrics. Choose a short passage containing the event you want the visuals to respond to: a vocal entrance, a pause, a rise in intensity, or an instrumental change.

For an imaginary track, suppose the useful passage runs from 0:48 to 0:58. Save that ten-second excerpt and note where it came from. A musical change at song time 0:52 occurs four seconds into this file.

Keep a simple asset note:

Source: approved final song mix
Excerpt: song 0:48–0:58
Length: 10 seconds
Important event: the arrangement opens up at excerpt second 4
Visual goal: a restrained scene becomes more spacious
Final soundtrack: retain the approved song in the editing timeline

The times are illustrative. Use your own recording's actual boundaries. If you later add silence, trim the start, or change playback speed, update the note before comparing synchronization.

Play the exported excerpt locally. Listen for a clipped first consonant, an abrupt ending, or an unexpected silent lead-in. Those are concrete problems to fix before interpreting a generated result as a model failure.

Pair the audio with a visual brief that has one main event

Our original scene is a small observatory at dawn. A seated figure looks toward a circular shutter. As the music expands, the shutter opens and reveals the sky. The first test does not require singing, dancing, several locations, or complicated object handling.

Prepare an approved image of the observatory or another compatible visual input, then attach it and the excerpt using your interface's controls. The following prompt uses descriptive asset names rather than assuming a provider-specific reference-tag syntax:

Use the supplied observatory image for the room layout and the seated figure.
The supplied ten-second song excerpt is the musical reference for this scene.
Keep the beginning restrained: the figure remains seated and looks toward
the closed circular shutter. As the arrangement becomes fuller, the shutter
opens to reveal the dawn sky. Finish with a clear view through the opening.
Use one steady wide view. Keep the figure's appearance and the room consistent.
The important visual event is the opening shutter; do not add a singing performance.

This is a proposed brief, not a tested preset. It tells the model what relationship you want, while leaving the final timing judgment to the edit. If the interface requires asset tags, map the actual uploaded files according to its documentation.

If the first result is hard to evaluate, simplify the action. A shutter already partly open may be a better test than a complex mechanical opening. Label that as a changed brief so you do not mistake it for an identical comparison.

Check three things separately in the returned clip

First watch the visuals without judging the soundtrack. Is the shutter still the same object? Does the seated figure change? Is the opening understandable?

Next listen to the returned audio. Does it contain your expected material, a changed performance, extra ambience, speech, or a new music layer? Native audio generation means there is another part of the output to inspect.

Finally place the clip against the original song passage. For the example, examine the shutter around song time 0:52 and review the transition into the next shot.

Observation Next step
The visual event works but happens late Try placement or trimming before another generation
The opening is early but the ending fits Review both boundaries before shifting the entire clip
The scene is coherent but generated music conflicts Keep the approved song and remove unwanted clip audio in the editor
The character or room changes Review conflicting visual references and simplify the request
A singing performance does not match the words Check the exact excerpt and use a dedicated synchronization review

Our timestamp-planning guide explains how to distinguish song time from clip time. The timing arithmetic is useful here, but Seedance-specific controls do not transfer to H3.

For singing faces, see the music-video lip-sync guide. A convincing voice or character reference does not, by itself, verify the mouth against every phrase.

Confirm the resolution workflow before requesting a final version

The official open-source documentation describes H3-Base output at 768p and a separate H3-Regenerate-2K stage for 2K output. A hosted service may package these stages differently. Check what the selected option actually produces rather than assuming every H3 request is a single native-2K pass. H3 system workflow.

Review the final-resolution result again. Keep the accepted draft until you have checked its replacement's motion, framing, duration, and audio. A higher-resolution version still needs to fit the same musical passage.

Build the full video around the approved recording

Once a representative clip works, continue with the surrounding sequence. Keep each excerpt's source position beside its visuals so that later edits do not quietly drift away from the selected song version.

LyricVideoMaker provides an audio-upload and editable lyric-video workflow. This external H3 reference method does not imply a built-in H3 integration. For the full production sequence, use the song-to-music-video guide.

Can I attach only the song and expect a complete video?

Check the input mode. The documented Ref2VA checkpoint requires visual input alongside reference audio. Even when a hosted tool offers a simpler workflow, verify its requirements rather than assuming a song upload is sufficient.

Does audio reference guarantee exact beat synchronization?

No guarantee is established by accepting the reference. Inspect the action at the musical event you care about and make the final placement in the timeline.

Should I use several audio references at once?

Start with the fewest files that communicate the task. Multiple excerpts need distinct roles and clear boundaries. For the observatory example, one short passage gives you enough information to judge whether the central visual event works.

Give your song a visual story

Bring your audio, work on the lyrics, and build a video around your song.

Create a lyric video

Keep reading