MiniMax H3 Can Now Inpaint Audio Too: How Video & Audio Editing Works

Edit dialogue in an existing AI video by deciding what is already approved and what actually needs to change. This workflow compares prompt regeneration with using an approved source video and replacement audio to show how a narrower revision can preserve the visual performance while updating the voice and mouth movement.
JoeyUpdated

Quick Answer

If an existing AI video only needs a dialogue change, you do not always need to rebuild the whole clip.

In Astorie, the workflow I tested is not a timeline-based audio masking interface. I used an existing video and replacement audio as inputs to a new video generation.

I compared two ways to change one spoken line in the same six-second character video.

First, I changed the dialogue in the original prompt and generated another version. The new voice worked, and the lip movement matched the new line. However, Carrie's body movement also changed.

Then I kept the approved video, created the replacement audio separately, and connected both to a new video node. The voice and mouth movement changed. In my visual comparison, I did not notice meaningful changes to the body motion, room, or camera movement.

The practical rule is:

Regenerate when the performance is still open to change. Use the approved video as an anchor when you only want to revise the speaking layer.

What Video and Audio Editing Means in This Workflow

Changing one sentence inside a video prompt sounds like a small edit.

In practice, it can reopen much more than the sentence itself.

A generated performance includes dialogue, facial expression, body movement, timing, framing, and camera motion. If you run the generation again, some of those choices may change even when most of the prompt stays the same.

That makes editing a question of scope.

Before generating a revision, decide what is already approved and what you are willing to change.

Prompt Regeneration Reopens the Performance

My original video showed Carrie in an indoor living room.

She started near the sofa, glanced to the side, turned her upper body, took a few steps, and looked toward the camera. Near the end, she smiled and said:

"Let's begin."

The clip was six seconds long at 768P.

For the first revision, I kept the same character references, scene reference, shot direction, and camera instructions. I changed the spoken line to:

"Ready when you are."

The dialogue changed successfully. Carrie's mouth movement also matched the new line.

But her body movement changed as well.

That was important because I was not trying to redesign the performance. I only wanted to update one sentence.

In this test, changing the prompt worked more like generating a new performance variant than making a narrow dialogue revision.

Source Video + New Audio Narrows the Revision

For the second path, I stopped rebuilding the shot from its original prompt.

Instead, I kept the approved video and created the new dialogue as a separate audio input. I connected both to another video node.

The prompt focused on preservation:

Use the connected video as the primary visual reference and the connected audio as the new dialogue reference.

Keep the same character, outfit, hairstyle, room, lighting, camera movement, body motion, shot composition, and overall timing from the original video.

Replace the original spoken line with the dialogue from the connected audio.

Adjust Carrie's mouth movement and subtle facial performance so they match the new audio naturally. Keep the rest of her movement as close as possible to the original video.

Do not redesign the scene or introduce new camera movement. Preserve the approved visual structure and only change what is necessary to make the new dialogue feel naturally synchronized with the character.

The result was much narrower.

The voice changed, and the mouth movement updated with it. In my visual comparison, I did not notice meaningful changes to Carrie's body motion, the room, or the camera movement.

That made this path more useful once the visual shot was already approved.

How to Edit Video and Audio With MiniMax H3

Step 1: Start With an Approved Source Video

Start with the video you actually want to preserve.

For this workflow, I reused the same six-second Carrie clip from my earlier H3 project.

The useful part was not that every detail was perfect. It was that the visual choices unrelated to the dialogue were already acceptable.

I wanted to preserve:

  • Carrie's appearance
  • her outfit and hairstyle
  • the indoor environment
  • the camera push-in
  • the basic body movement
  • the overall shot structure

Only the final spoken line needed revision.

This creates a clear success condition.

If the dialogue changes but unrelated parts of the shot also shift, the revision has reopened more variables than necessary.

Step 2: Separate Approved and Open Variables

The original dialogue was:

"Let's begin."

The revised line was:

"Ready when you are."

The two lines are similar in length, so the comparison does not depend on forcing a much longer sentence into the same speaking moment.

Before editing, I divided the shot into two groups.

Already approved:

  • character
  • environment
  • body motion
  • camera movement
  • shot structure

Open to revision:

  • spoken line
  • mouth movement connected to the new line
  • small facial adjustments needed for the delivery

This is a useful editing habit beyond this one project.

Do not begin with:

How do I regenerate this video?

Begin with:

What am I actually willing to let change?

Step 3: Revise the Original Prompt

The first method was the simplest.

I reused the original prompt and changed the dialogue while keeping the rest of the instructions close to the base version.

Use @CHAR_01_HERO, @CHAR_FULL, and @SCENE_01_MED as references.

Create a realistic 6-second video set in the same indoor living room environment.

Show Carrie starting near the sofa in a relaxed standing pose. At first, she glances slightly to the side. Then she turns her upper body, takes one or two natural steps forward, and looks toward the camera. Near the end, she gives a small friendly smile and says one short line: "Ready when you are."

Keep the character identity, outfit, hairstyle, and facial features consistent with the references. Keep the room layout, furniture, lighting, and atmosphere consistent with the scene reference.

Use natural body motion and clear pose changes. Keep the movement simple and believable. Add a subtle camera push-in so the shot feels slightly more dynamic, but keep the camera movement controlled and smooth.

Keep the overall shot structure, pacing, and visual setup as close as possible to the approved base version. The main intended change is the final spoken line.

What Happened in My Test

The new dialogue worked.

Carrie said "Ready when you are," and her lip movement matched the new speech.

However, I also noticed a change in her body movement compared with the original clip.

The wording was almost unchanged, but the result was still a new generation. The performance had been reopened.

That does not make this workflow useless.

I would still use it when I like the general scene and character direction but do not need to preserve the exact movement from the previous version.

If the shot is already approved, however, I would use a narrower revision path.

Step 4: Keep the Video and Add Replacement Audio

For the second method, I treated the original clip as the approved visual asset.

I generated the replacement line separately:

"Ready when you are."

Then I connected both inputs to the next video node:

approved video + replacement audio → new video

Astorie canvas showing the approved video and replacement audio connected to the new video node

I did not describe Carrie's entire performance again.

I did not repeat where she should stand, how she should walk, or how the camera should move. Those decisions were already represented by the source video.

Instead, I asked the new generation to preserve the existing shot and update the speaking performance needed for the replacement audio.

What Happened in My Test

This produced the narrower revision I wanted.

The voice changed.

Carrie's mouth movement also changed to match the new dialogue.

In my visual comparison, I did not notice meaningful changes to the body motion, room, or camera movement.

The result preserved the visual structure much better than the prompt-regeneration path.

[Insert image: matching frames from the original and video-plus-audio versions, showing the same body position, scene, and framing]

Step 5: Compare the Two Editing Paths

The same dialogue change produced two different kinds of revision.

Element

Original

Prompt Regeneration

Video + Replacement Audio

Dialogue

"Let's begin."

Changed

Changed

Mouth movement

Original

Updated for new line

Updated for new line

Body motion

Baseline

Visibly changed

No meaningful visible change observed

Room

Baseline

Compare against new generation

No meaningful visible change observed

Camera movement

Baseline

Compare against new generation

No meaningful visible change observed

Best use

Approved baseline

New performance variant

Narrow dialogue revision

three-way visual comparison aligned to the same point in the clip

The choice depends on what is already approved.

If the whole performance can still change, prompt regeneration is reasonable.

If the visual shot is already approved and only the dialogue needs revision, keeping the source video gives the workflow a stronger visual anchor.

Why Audio Editing Is More Than Replacing a Voice Track

Changing dialogue in a visible talking shot has two parts.

The first is the sound.

The second is the visible performance connected to that sound.

If you simply place a new voice track over an existing video, the words may no longer match the original mouth movement.

That can be acceptable for narration or an off-screen speaker.

It becomes more noticeable when the character is visible and speaking toward the camera.

For visible dialogue, the useful middle ground is to keep the approved visual structure while allowing the speaking performance to update.

That is what happened in my second workflow.

The replacement audio defined the new dialogue. The source video anchored the approved shot. The new generation only needed to reconcile the speaking performance with those inputs.

This creates three useful editing levels.

Audio Only

Use this when the sound does not need to match a visible speaker.

Examples include narration or other off-screen speech.

Audio + Speaking Performance

Use this when the dialogue changes and the character's mouth or facial timing also needs to change.

This was the editing scope I needed for the Carrie clip.

Full Audiovisual Regeneration

Use this when dialogue, acting, body motion, or the shot itself can all change.

This gives the generation more freedom, but it also reopens more decisions.

Common Mistakes When Editing H3 Video and Audio

Rewriting One Line and Expecting a Surgical Edit

A small prompt change does not guarantee a small output change.

In my first test, the new dialogue and lip sync worked, but the body movement also shifted.

If visual preservation matters, reuse the approved video instead of relying only on the original written instructions.

Re-describing an Approved Video From Scratch

Once the source video is connected, there is little value in rewriting every visual detail.

You do not need to repeatedly describe the room, walking path, framing, and camera movement when those decisions are already visible in the source.

Keep the new prompt focused on the revision.

Treating Voice Replacement and Visible Dialogue Editing as the Same Task

A replacement audio file solves the sound problem.

It does not automatically solve the visible performance problem.

When the speaker is on screen, check whether the mouth movement and facial timing work with the new dialogue.

Changing Too Many Variables at Once

If you change the line, scene, character, camera, and motion in the same revision, it becomes difficult to understand what caused the final result.

For this project, I kept one approved source clip and changed only the dialogue workflow.

That made the difference between the two paths easy to evaluate.

A Simple Rule for Future H3 Edits

Preserve what already works. Reopen only what needs to change.

For this workflow, I think about the inputs like this:

Prompt = reopen the generation

Use it when you are still designing the performance.

Source video = anchor the approved visual structure

Use it when you want the new result to stay close to an existing shot.

New audio = define the revised speaking performance

Use it when dialogue needs to change and the visible delivery needs to follow it.

The editing workflow should follow the scope of the revision.

FAQ

Is this workflow the same as timeline-based audio inpainting?

No.

The Astorie workflow tested here did not use a timeline mask to select a specific audio region. I used an approved source video and replacement audio as inputs to a new video generation.

The useful production idea is still selective revision: keep the approved visual information in the workflow and reopen only the part that needs a new performance.

Can I change dialogue in an existing H3 video?

One practical method is to keep the existing video, create replacement audio, and use both as inputs for the next video generation.

In my test, the voice and mouth movement changed. I did not notice meaningful visual changes to the body motion, room, or camera movement.

Will changing the dialogue prompt preserve the original motion?

Not necessarily.

In my prompt-regeneration test, the dialogue and lip sync both updated, but Carrie's body movement also changed.

If you need to protect an approved performance, the existing video can provide a stronger visual anchor than the written prompt alone.

Why not just replace the audio track?

That may be enough when the new sound does not need to match a visible speaker.

For on-camera dialogue, the new words also need to work with mouth movement and facial timing.

Using the video and replacement audio together allows the speaking performance to be revised along with the sound.

Should I regenerate the whole video when one line changes?

It depends on which parts are still open to change.

If the performance is still flexible, regenerating from the prompt can create a new coordinated version.

If the visual shot is already approved, start from that video and limit the revision to the dialogue-related performance.

Ready to try it on the canvas?

Open Astorie and fan your prompt across every frontier model in one workflow.

This website uses cookies

Analytics and marketing tags are on by default in your region — you can turn them off here at any time. We also use basic cookies to keep Astorie secure and remember preferences.

Read more