MiniMax H3 Keyframe Control Guide: Add Image and Audio References at Any Frame

Use MiniMax H3 keyframe control to extract a useful moment from an existing video and reuse it as a visual anchor for a new generation. This workflow shows how to combine an extracted frame with optional audio while treating timing as generation guidance rather than frame-exact editing.
JoeyUpdated

Quick Answer

MiniMax H3 keyframe control can be used to place visual and audio guidance at selected moments in a new video workflow. But "at any frame" does not mean the generated result will match an exact frame or timestamp perfectly.

In my test, I started with an existing 6-second, 768P base video. I extracted one useful frame from 3.77 seconds, saved it as an Element called @keyframe_reference, and reused it to guide a new H3 generation.

The extracted frame worked well on its own. Carrie, her outfit, and the living room stayed consistent without separate character or scene Elements.

I then added an audio reference generated with MiniMax Speech 2.5 HD. In that second test, the image reference shaped the visual moment while the audio influenced the later dialogue. The result was good, but the speech started earlier than the timing written in my prompt.

So I would treat keyframe timing as generation guidance, not frame-exact editing.

What You Need Before You Start

This workflow starts with one approved source video.

I already had a reusable base clip from an earlier H3 workflow. It showed Carrie standing in an indoor living room, moving near the sofa, turning, and looking toward the camera.

The original clip was 6 seconds at 768P. I reused it here instead of creating another source video from scratch.

For this workflow, you need:

  • one approved source video
  • one useful frame from that video
  • MiniMax H3 for the new generation
  • an optional audio reference if the next shot needs dialogue

How to Use MiniMax H3 Keyframe Control

Step 1: Choose a Useful Moment From Your Existing Video

Do not extract a frame just because it sits in the middle of the video.

Choose a moment that already contains useful visual information. The character should be clear, the pose should be usable, and the scene should contain enough context for the next generation.

In my base video, I selected a frame from about 3.77 seconds. Carrie was clearly visible, the room composition was stable, and the frame contained enough information to reuse.

The important distinction is this:

You are not editing that frame inside the original video. You are turning it into a reusable visual asset for a new generation.

the Key Frame Extraction node with the source video and the selected 3.77s frame

Step 2: Extract the Frame and Save It as a Reusable Reference

I extracted the selected frame and saved it as an Element named:

@keyframe_reference

My canvas path was:

VID_01_BASE → Key Frame Extraction → @keyframe_reference

Saving the frame as an Element made the workflow easier to reuse. I could reference the approved visual moment directly in another generation branch instead of reopening the original video and finding the timestamp again.

Step 3: Use the Extracted Frame as the Visual Anchor for a New Video

For the first generation branch, I used only @keyframe_reference.

I did not add a separate character reference or scene reference. I wanted to find out whether the extracted frame itself contained enough information to guide the next clip.

I used this prompt:

Use @keyframe_reference as the visual guide for a key moment in the video.

Create a realistic 6-second video in the same indoor living room. Keep the woman, outfit, hairstyle, room layout, furniture, and lighting consistent with the reference.

From 0.0s to about 3.5s, show Carrie moving naturally into position. At about 3.7s, match the overall pose, framing, and scene composition of @keyframe_reference. After that, let her shift slightly and turn toward the camera with a calm natural expression.

Keep the motion smooth and believable. Use subtle camera movement only.

In this test, the extracted frame was enough.

Carrie remained recognizable. Her outfit and hairstyle stayed stable, and the living room still looked like the same environment. The motion around the guided moment also remained smooth.

I did not need extra @CHAR or @SCENE Elements for this branch.

However, the generated moment was not an exact copy of the extracted frame. H3 built new motion around the reference instead of reproducing it like a pasted still image.

That makes the extracted frame better understood as a visual anchor than as an exact frame replacement.

side-by-side comparison of @keyframe_reference at 3.77s and the closest generated moment from the image-only branch

Step 4: Add an Audio Reference for the Dialogue Layer

Once the image-only branch worked, I kept the same visual reference and added audio.

I generated the spoken reference with MiniMax Speech 2.5 HD using this line:

"I’ll be right there. Just give me one second."

The purpose of this branch was not to compare models. I wanted to see what changed when I kept the same visual anchor but added a spoken reference.

I used this prompt:

Use @keyframe_reference as the visual guide for a key moment in the video, and use the provided audio reference for the spoken line.

Create a realistic 6-second video in the same indoor living room. Keep the same woman, outfit, hairstyle, room layout, furniture, and lighting consistent with the visual reference.

From 0.0s to about 3.5s, show Carrie moving naturally into position in the room with smooth, believable body motion.

At about 3.7s, match the overall pose, framing, and scene composition of @keyframe_reference. This is the main guided visual moment.

From about 4.2s to 5.8s, have Carrie look toward the camera and deliver the spoken line naturally: "I’ll be right there. Just give me one second."

Keep the expression calm and natural. Let the mouth movement and body language support the speech without becoming exaggerated. Use subtle camera movement only.

This version also worked well.

The visual continuity remained strong after the audio was added. Carrie still looked like the same character, and the room remained recognizable. Her mouth, face, and body movement also supported the spoken line naturally enough for the shot to feel coherent.

In this test, the image and audio references influenced different parts of the same generated clip. The extracted frame shaped the visual beat, while the audio shaped the speaking performance.

That is a useful workflow distinction, but it comes from this test rather than a guarantee that every generation will behave the same way.

Astorie branch showing @keyframe_reference and the MiniMax Speech 2.5 HD audio reference feeding the new video generation

Step 5: Use Timing as Guidance, Then Check the Actual Result

The second branch also revealed the most important limitation in this workflow.

I asked for the visual moment around 3.7 seconds and the spoken line from about 4.2 seconds onward.

The dialogue appeared earlier than the exact timing written in the prompt.

The result was still usable. The sequence made sense, and the speech felt natural. But it did not behave like frame-accurate timeline editing.

That changed how I would use the workflow.

Instead of expecting an exact timestamp, I would:

  • use timing language to organize the shot
  • leave enough space around important actions or dialogue
  • review where those moments actually appear in the generated result

If the timing is already acceptable, there is no reason to rerun the clip only because it missed the requested second.

When to Use a Keyframe Instead of Another Reference

Different reference types solve different production problems.

Goal

Better Control

Keep the same character identity across generations

Character or image reference

Reuse movement from an existing clip

Motion transfer

Reuse one strong visual moment from an existing video

Keyframe

Add spoken performance or sound guidance

Audio reference

Change part of an already generated video

Video editing or inpainting workflow

The simplest decision rule is:

Use a keyframe when the moment itself is worth reusing.

If you only need the same person, use a character reference.

If the movement is what matters, motion transfer is the more relevant workflow.

If one frame already contains the pose, composition, and scene state you want to carry forward, extracting that frame gives you a reusable visual starting point.

Common Keyframe Control Problems and Fixes

The Generated Moment Is Similar but Not Identical

A generated keyframe moment may resemble the reference without matching it exactly.

That happened in my first branch. The pose and composition moved toward the extracted frame, but the result still looked like a newly generated video.

Judge the result by the production details that matter:

  • pose
  • framing
  • subject identity
  • scene continuity

Do not assume that success requires pixel-level matching.

Character or Scene Details Start to Drift

An extracted frame may not always contain enough information to preserve every detail.

If identity starts to drift, add a dedicated character reference. If the environment changes too much, add a scene reference.

I did not need either one in this test. The extracted frame was already strong enough for the result I wanted.

That is why I would start with the minimum useful reference set and add more inputs only when a clear problem remains.

Dialogue Starts Earlier or Later Than Expected

This happened in my audio branch.

The prompt placed the dialogue later in the clip, but the generated speech began earlier than requested.

The practical fix is to give dialogue more room rather than treating the timestamp as an exact edit point.

Keep the line short enough for the clip, leave space around it, and check the actual generated timing before deciding whether another run is necessary.

Too Many References Make the Workflow Hard to Diagnose

Adding every available reference can make a generation harder to understand.

If you use a character reference, scene reference, keyframe, audio file, and several timing instructions at once, it becomes difficult to tell which input solved the problem.

My workflow stayed simple:

keyframe only → evaluate → add audio

That was enough for this project.

Add another reference only when you can name the specific problem it needs to solve.

The Keyframe Workflow I Would Use Again

After both branches, this is the workflow I would reuse:

approved video → choose a strong moment → extract the frame → save it as a reusable Element → generate a new clip around that visual anchor → add audio if the shot needs dialogue → review continuity and timing

The first branch showed that one extracted frame could already carry enough character and scene information for my project.

The second showed that adding audio could extend the same workflow without requiring me to rebuild the visual setup.

The main limitation was timing. The requested seconds helped structure the prompt, but the actual result still needed to be checked after generation.

So the most useful way to think about keyframe control is not "perfect control of one exact frame."

It is this:

You can turn a proven moment from one video into a reusable asset for the next one.

FAQ

Can I use a frame from an existing AI video as a new H3 reference?

Yes. That is what I did in this workflow.

I extracted a frame at about 3.77 seconds from an existing base video, saved it as @keyframe_reference, and reused it in a new generation.

Does keyframe control guarantee an exact match at an exact second?

No.

In my test, the generated visual moment was similar to the extracted reference but not identical. The audio branch also began speaking earlier than the exact timing written in the prompt.

Treat the timing as guidance, then check the generated result.

Do I still need character and scene references after extracting a keyframe?

Not always.

I did not need them in this test because the extracted frame already contained enough character and scene information.

If your character or environment starts to drift, adding a dedicated reference can be the next step.

Can I combine image and audio references in the same H3 workflow?

I did so successfully in this test.

The extracted frame acted as the visual anchor, while the MiniMax Speech 2.5 HD reference influenced the dialogue. The combination worked well in this branch, but the timing still needed to be reviewed after generation.

Ready to try it on the canvas?

Open Astorie and fan your prompt across every frontier model in one workflow.

This website uses cookies

Analytics and marketing tags are on by default in your region — you can turn them off here at any time. We also use basic cookies to keep Astorie secure and remember preferences.

Read more