For connected music video shots, decide what your Grok Imagine reference images must preserve before writing the motion prompt. A reference can guide a person's appearance or an object; a pinned frame controls an endpoint. Confusing those jobs can leave you with a recognizable performer in a composition that will not cut into your sequence.
This guide develops an imaginary record-shop scene from a small reference pack. The prompts are original planning examples, not tested presets or claims of guaranteed character consistency. The technical controls below refer to the xAI API documentation checked on September 16, 2026. A consumer app or third-party model picker may expose fewer controls.
Choose between an appearance reference and a fixed frame
Start with the failure you need to prevent. Is the jacket changing? Is the opening composition wrong? Or is an otherwise usable clip missing one detail?
The official API separates these operations:
| Your input | What it controls |
|---|---|
| A single image for image-to-video | The still that will be animated |
| Reference images | Visual elements that guide a generated scene, without fixing its opening composition |
First and last frame pins on grok-imagine-video-1.5 |
Exact endpoints, with generated movement between them |
| An existing video for editing | Changes to footage you already have |
See the official image-to-video, reference-to-video, and video editing documentation.
For example, a portrait is useful when the performer should reappear in a new location. A finished wide shot is useful when the clip must begin with the performer standing at a particular shelf. Those are different production decisions, even if both involve uploading a picture.
Build a small Grok Imagine reference pack with clear jobs
Prepare only the images needed to resolve this shot's important uncertainties. For our record-shop example:
- Performer: an unobstructed view of an adult character wearing a mustard cardigan and round glasses.
- Location: a quiet record shop with pale wooden shelves and a blue checkout counter.
- Prop: a plain coral record sleeve with a small cream circle, without lettering.
These are suggested reference roles, not required upload categories. Use material you created or have permission to use. Avoid a character image containing several equally prominent people: you would be asking the next stage to decide which person matters.
Check the pack for contradictions before generating anything. If the portrait has black glasses but a second image shows gold frames, decide which appearance you want. If the shop image is daytime and the brief says midnight, decide whether the reference supplies the layout, lighting, or both. Put that decision into the brief.
Keep the source files and their order beside your shot notes. Reordering images without updating their references can turn an apparently small prompt revision into a different request.
Write the relationship between the images
The official reference examples use image tags such as <IMAGE_1> in the prompt. Follow your interface's reference mapping; merely typing a tag does not attach a file. xAI reference examples.
Here is an original brief for the three-image pack above, assuming the interface maps the uploads in that order:
Use the adult performer from <IMAGE_1> in the record shop from <IMAGE_2>.
The performer wears the mustard cardigan and round glasses shown in the portrait.
The coral record sleeve from <IMAGE_3> rests on the blue counter.
A medium shot from the customer's side of the counter. The performer slides
one hand beneath the sleeve and lifts it to chest height, keeping its front
facing the camera. The other hand stays resting on the counter.
The camera remains stationary. The performer looks down at the record.
Use the shop reference for its layout, with soft afternoon light.
The important editorial choice is the relationship: one person, one prop, one location. The prompt also leaves some things alone. It does not ask the performer to lift the record, turn around, dance, walk outside, and reveal a new street within the same clip.
If the result is crowded, remove a reference that has no essential job before adding more instructions. If the prop must carry readable lyrics or a song title, add that text during editing. The plain sleeve in this example keeps typography out of the continuity test.
Use first and last frames when the cut needs an exact endpoint
Suppose the next shot is a close-up of the sleeve. You may want the preceding shot to end with that sleeve already filling a similar part of the frame. Draw or prepare those two compositions first, then judge whether the physical transition makes sense.
In the 1.5 API, image plus last_frame supplies both pins. The classic grok-imagine-video model rejects that combination. The documentation currently caps reference-to-video at 720p and 15 seconds on 1.5; do not carry a text-to-video resolution assumption into this workflow. First and last frame controls.
A minimal illustrative REST body looks like this. Replace both URLs with accessible images before using it; this is a request example, not an executed generation:
{
"model": "grok-imagine-video-1.5",
"image": { "url": "https://example.com/shop-opening.jpg" },
"last_frame": { "url": "https://example.com/sleeve-ending.jpg" },
"prompt": "The performer raises the record sleeve while the camera moves closer, settling on the supplied closing composition.",
"duration": 8,
"aspect_ratio": "16:9",
"resolution": "720p"
}
The eight seconds are an example setting. Choose the length around the edit you actually need. For a four-second passage, a long approach followed by a late reveal may leave you with no useful cut, even if the generated shot looks attractive on its own.
Compare the planned endpoints before sending the request:
- Does the sleeve begin in a place the performer can plausibly reach?
- Is the same hand holding it in both images?
- Can the camera move between the compositions without crossing a wall or counter?
- Is the lighting change intentional?
A pinned ending does not remove the need to review the movement leading into it. Watch the middle of the clip for a changing sleeve shape, a disappearing hand, or an abrupt reframing.
Diagnose continuity errors before regenerating
Review the candidate beside the previous approved shot, with the intended song playing. A detail that seems small in isolation can become obvious at the cut.
| What failed | What to inspect first | A focused next attempt |
|---|---|---|
| Wrong clothing or glasses | Conflicting details in the reference pack | Keep one unambiguous appearance reference |
| Correct person, wrong opening position | Whether the image was used as guidance or as an endpoint | Choose a fixed opening composition where the interface supports it |
| Sleeve changes shape during the lift | Hand contact and how much movement the shot requests | Shorten the action or begin with the sleeve already held |
| The cut jumps across the counter | Screen direction and camera position in neighboring shots | Keep both viewpoints on the same side for this sequence |
| Ending is correct but movement looks forced | Distance between the prepared compositions | Bring the endpoints closer or split the transition into two shots |
These are editing diagnoses, not promises that a particular wording will fix the model. Save the accepted source clip before trying an alternative. Record which reference or instruction changed so you can compare meaningful variations.
If a clip already has the movement you want, investigate video editing instead of rebuilding the whole scene. xAI documents a separate editing route with an input limit of 8.7 seconds and output capped at 720p; it preserves the input's duration and aspect ratio rather than accepting replacement values for those settings. Video editing constraints.
Keep the recording and the visual continuity checks separate
For this example, the performer's action is handling a record, not singing. That makes it possible to evaluate visual continuity without also judging whether mouth movements match a vocal.
Use your finished song as the timeline reference. Place each accepted clip against the actual lyric or instrumental passage, trim it, and listen for any unwanted audio from the clip. Muting a clip's audio in your editor gives you a direct way to protect the original song mix.
If the performer must visibly sing, treat synchronization as its own task. Our Suno music video lip-sync guide covers the shot and audio preparation involved. A convincing face reference alone does not establish that the mouth matches your recording.
For broader composition and camera planning, use the FLUX 3 shot-planning guide as a separate reference; its model-specific controls should not be copied into a Grok request. If your goal is a video built around your song and readable lyrics, you can also start from LyricVideoMaker. The Grok workflow described here is external and is not a claim of a built-in Grok integration.
Questions before your next generation
Should I fill every reference slot?
Only when every image has a necessary, nonconflicting role. Start with the smallest pack that explains the shot, then add a reference to address a specific missing detail. More files are not a substitute for deciding what each one contributes.
Why does the clip recognize my character but start somewhere else?
Check whether you supplied appearance guidance or a fixed starting frame. Then inspect the controls your interface actually sends. A sentence asking for a particular opening should not be treated as proof that a first-frame pin is active.
Can a prompt guarantee matching shots?
No. Keep the references, model choice, and approved clips organized, and review each new clip at its intended cut. For the record-shop example, get the single sleeve lift working before commissioning the rest of the sequence. That gives the next shot a concrete visual reference to match.
