LyricVideoMaker
← All articles
Tutorials

H3 Max Turbo Music Video Prompts: Three Practical Examples

Write H3 Max Turbo prompts for music video shots, use first and last frames, and keep your song separate from generated sound with the right controls.

LyricVideoMaker Team

For an H3 Max Turbo music video shot, describe what changes on screen, choose the input mode that supports it, and set the soundtrack separately. A prompt can direct a curtain opening or a camera approaching a performer. It cannot substitute for attaching a first frame, setting the output size, or placing the right excerpt of your song in the edit.

This guide uses three original examples built around a miniature theater. They are starting prompts, not results from a generation test. The endpoint details below refer to fal's H3 Max Turbo documentation checked on October 2, 2026; another website's model picker may expose different controls.

Decide what the shot needs before choosing its input

Imagine a song whose chorus is about finally stepping forward. Your visual idea is a small paper theater with a blue curtain and a single wooden figure. You could show the curtain opening, animate a still you already like, or end on a specific composition for the next cut.

Those are different requests:

What you already have Useful starting point What the prompt should explain
Only the scene idea Text to video The set, subject, action, and camera
A finished opening image Image to video What moves after the opening frame
Opening and closing images First and last frames A plausible action connecting the two

fal describes H3 Max as a version of MiniMax H3 post-trained by its own team. Treat Max Turbo as a specific model and endpoint, rather than assuming every feature of the broader H3 family transfers to it. fal's model announcement.

If your actual requirement is multimodal reference guidance, our H3 audio-reference guide covers that separate workflow. The examples here concern text and keyframe inputs for Turbo.

Example 1: Reveal a scene from text

Use this when the miniature theater does not yet have a fixed appearance. Choose text to video, then paste:

A handmade miniature theater sits on a dark wooden table. Its blue paper
curtain is closed. The curtain separates slowly from the center, revealing
one small wooden figure standing beneath a warm overhead lamp. The figure
raises its head once and then remains still. A stationary front-facing
camera shows the entire stage throughout. Visible paper fibers and painted
wood give the scene a tactile stop-motion appearance. The table and theater
frame remain still. No writing appears on the set.

The action has a readable beginning and end: closed curtain, reveal, settled figure. That gives you somewhere to enter and leave the clip. The prompt also assigns movement to the curtain and figure while keeping the camera steady.

For a shorter musical passage, remove the head movement before adding more speed instructions. Two events compressed into a small gap can become harder to read than one clear reveal. For a longer passage, allow the revealed stage to hold instead of adding an unrelated second location.

If lyrics will overlay this scene, leave them for the text layer in your editor. Asking the scenery itself to carry the chorus makes the shot harder to reuse when you correct a lyric or change its timing.

Set technical controls outside the scene description

On fal's Turbo text-to-video endpoint, duration and aspect_ratio are separate inputs. Resolution choices are 480P, 768P, and 1080P; the documentation describes 1080P as refinement from a native 768P source. Writing “4K” into the scene does not add an output option.

For a straightforward prompt comparison, keep prompt_expansion_mode set to disabled. The endpoint also provides balanced and quality modes that rewrite prompts before generation. If you enable expansion, inspect expanded_prompt when it is returned so you can see what was sent onward. These are fal controls, not required words inside the creative prompt.

Record the selected settings with the saved clip. When a result is wrong, this lets you distinguish an unsuitable scene description from a different input size, duration, or prompt-rewriting setting.

Example 2: Animate an opening image without redesigning it

Suppose you have already made a still of the open theater. Use it as the starting frame rather than asking a new text-only generation to recreate the set.

Continue from the supplied opening image. The wooden figure slowly turns
its head toward the warm lamp above the stage. Its feet remain planted.
The blue curtain hangs still on both sides. The camera moves gently closer
while keeping the figure centered and its full body visible. Preserve the
paper theater's proportions, the figure's painted surface, and the existing
lighting. End with the figure holding the upward gaze.

Notice that this version does not ask the curtain to open again. That event has already happened in the supplied image. Describe the next movement from the visible starting state.

fal's Turbo image-to-video schema names the opening input image_url and the ending input end_image_url. It also supports an ending image alone. The output canvas follows the opening image when one is supplied, or the ending image for an end-only request.

Prepare the reference in the composition you actually need. If the wooden figure fills a landscape image from head to foot, a later vertical crop may cut it off. Resolve that framing problem before evaluating whether the head turn worked.

Example 3: Connect two frames with a believable action

For a first-and-last-frame attempt, use two images of the same theater from the same viewpoint. In the first, the figure stands at the rear of the stage. In the second, it stands near the front edge. Keep the curtain, light, and figure design consistent between the stills.

The wooden figure takes several small, deliberate steps from the rear of
the miniature stage toward the front edge, arriving at the position shown
in the ending image. Its body remains upright and its painted appearance
stays consistent. The camera is stationary. The blue curtain and theater
walls do not move. The warm overhead light remains constant. Finish with
the figure standing still at the front of the stage.

The two images provide endpoints; the prompt supplies a proposed route between them. They do not guarantee that every intermediate frame will remain coherent.

Watch the middle of the result carefully. If the figure stretches instead of stepping, reduce the distance between the two reference positions. If the set rotates, check whether the images accidentally imply different camera angles. Changing the references may be more useful than adding another paragraph of instructions.

Keep your song excerpt under your control

The current fal Turbo schema includes target_audio_url: it replaces the output soundtrack with supplied audio, trimming a longer input from its beginning or padding a shorter one with silence. That is a soundtrack operation, not a documented promise of beat synchronization or lip sync. Turbo audio input documentation.

For the theater reveal, prepare the exact musical excerpt you want before attaching it. Supplying a full song and expecting the chorus to be selected automatically can put the introduction beneath your chorus shot. Keep the full-length recording as the master audio track in your editor, then place the finished visual at its intended song timestamp.

If your chosen interface does not expose a soundtrack input, add the original song during editing and mute unwanted clip audio. Listen to the final export: an extra generated sound bed can remain even when the visuals look right.

For a visible singer who must articulate the recorded words, use a workflow designed for that task; our Suno music video lip-sync guide explains the preparation and review involved. A wooden figure moving to music requires a different level of synchronization from a close-up singing performance.

Turn the usable shot into part of the song

Review each result in its actual position in the track. For these examples, check whether the curtain reveal, upward glance, or final step finishes before the next cut. Trim away unnecessary settling time only if the action still makes sense.

Then view it at the intended posting size with the lyrics visible. The miniature theater frame can create an attractive border, but it can also leave very little room for readable text on a phone. Adjust the crop or lyric placement before replacing an otherwise useful shot.

Keep the accepted clip, its input images, prompt, and settings together. If the next chorus needs a variation, change a specific event—such as the figure turning toward the audience—while preserving the set description.

For planning the rest of a track, follow our song-to-music-video workflow. LyricVideoMaker offers a song-upload and editable lyric-video workflow. The external fal controls described here should not be assumed to appear in LyricVideoMaker's interface.

Give your song a visual story

Bring your audio, work on the lyrics, and build a video around your song.

Create a lyric video

Keep reading