MiniMax H3 Review: Is Multimodal Control the Real Upgrade?

This MiniMax H3 review tests whether its biggest upgrade is better video quality or greater production control. I compared H3 with Hailuo 2.3 and Seedance 2.5, then tested multiple image references, audio input, first and last frame control, and 2K generation to see where the extra cost and waiting time actually pay off.
JoeyUpdated

Quick Answer

MiniMax H3 felt like a meaningful upgrade in our tests, but not because every video suddenly looked better. The bigger change was control.

H3 handled multiple references well, kept different reference roles separate, added speech without breaking the visual setup, and produced a believable transition between fixed first and last frames.

In our 2K test, the improvement went beyond visible sharpness: the character's face also stayed more stable during motion than in the 768P run.

The trade-off is cost and waiting time. H3 often took several minutes to finish, and the 2K version of one test took 11 minutes.

For workflows built around references, controlled shots, and multimodal inputs, that extra control can be useful. For simple image-to-video generation, the difference may be harder to justify.

What Is MiniMax H3?

MiniMax H3 is a video generation model that can work with several kinds of creative input in one workflow.

In Astorie, I used two H3 modes during this review.

Omni let me generate from a prompt and add image or audio references with @ references.

First / Last Frame let me define both ends of a shot, then ask H3 to generate the transition between them.

In my current Astorie workflow, H3 can generate videos from 4 to 15 seconds at 768P or 2K.

I did not try to test every possible H3 feature. Instead, I focused on the changes that affected an actual production workflow.

How I Tested MiniMax H3

I used Astorie for all of the video tests.

Where possible, I kept the same source assets and prompt between models. I also recorded generation time, credit cost, and whether the first result was already usable.

These are individual workflow tests, not a large benchmark. Video generation can vary from run to run, so the results below should be read as practical examples rather than universal rankings.

I tested six questions:

Test

What I wanted to learn

H3 vs Hailuo 2.3

What changed from the previous generation?

Three image references

Can H3 assign different jobs to different references?

Images + audio

Can H3 add another modality without breaking the visual setup?

First / Last Frame

Can it build a plausible transition between two controlled states?

768P vs 2K

Does higher resolution improve usable quality in this test?

H3 vs Seedance 2.5

How does H3 compare with another current video model?

MiniMax H3 vs Hailuo 2.3: What Actually Improved?

I started with the simplest generation comparison.

I created one 16:9 source image in GPT Image 2.5 Sunburst. It showed a woman standing on a train platform with a suitcase.

Then I gave Hailuo 2.3 and H3 the same motion prompt:

The woman looks toward the arriving train, then picks up the handle of her suitcase and starts walking along the platform. The camera slowly tracks backward in front of her while keeping her centered. A gust of wind moves her coat and hair as the train passes behind her. Near the end, she briefly turns her head toward the train while continuing to walk. Natural body movement, realistic fabric motion, continuous spatial relationships.

For a fairer baseline, both videos were six seconds long.

Hailuo 2.3 ran at 720P and used 27 credits. It finished in 1 minute 45 seconds.

H3 ran at 768P and used 84 credits. It took 2 minutes 25 seconds.

Hailuo 2.3
H3

The Difference Was Not Just Motion Quality

Hailuo 2.3 did create movement, but the performance was hard to read.

The woman kept turning her head from side to side. The motion existed, but it did not clearly communicate that she was watching the arriving train.

H3 organized the same instructions more clearly.

The woman's attention toward the train was easier to understand, and the sequence had a more natural rhythm. The result felt more like a directed performance instead of a collection of movements.

A model can generate movement without clearly expressing why the character is moving.

In this test, H3 did a better job of turning sequential instructions into readable performance.

The cost was also clear. It used more than three times the credits of Hailuo 2.3 and took 40 seconds longer.

How Good Is MiniMax H3 at Multimodal References?

This was the most useful part of the review.

Instead of giving H3 one starting image, I gave it three separate image references with different jobs.

I generated the references with GPT Image 2.5 Flare.

Each image cost 10 credits and took about 13 seconds to generate.

The references were:

  • Image A: character identity
  • Image B: outfit
  • Image C: bookstore environment

I deliberately made the outfit reference less useful as an identity reference and kept the bookstore reference free of people.

That made it easier to see whether H3 understood what each image was supposed to contribute.

Three Images, Three Different Jobs

I used this prompt:

Use Image A as the identity reference for the main character. Use Image B as the outfit reference. Use Image C as the environment reference. Create a cinematic 8-second 16:9 video of the woman from Image A standing inside the bookstore from Image C while wearing the outfit style from Image B. She first looks around the bookstore, then lightly runs her fingers along a row of books on a nearby shelf, and then takes a few slow steps forward through the space. The camera moves gently sideways to reveal more of the bookstore while keeping her as the main subject. Keep her face recognizable from Image A, keep the outfit details from Image B clear, and preserve the overall environment and atmosphere of Image C.

The H3 video used 112 credits and took 2 minutes 56 seconds.

The first result was already basically usable.

More importantly, the three references did not collapse into each other.

The woman still looked like the character from Image A. Her clothing came from Image B. The bookstore clearly came from Image C.

H3 did not simply preserve several references at once. It assigned different jobs to them.

That was a more meaningful result than simply confirming that multiple images could be uploaded.

3 images as reference generated H3 video

What Happens When You Add an Audio Reference?

I kept the same three image references and added a fourth input.

I generated a short voice clip with MiniMax Speech 2.5 HD.

The line was:

I think I could stay here for hours.

The audio took 2 seconds to generate and cost 10 credits.

I then asked H3 to place that speech inside the bookstore scene.

The video still cost 112 credits, but generation time increased to 4 minutes 56 seconds.

The woman's movement was not identical to the previous version, which is expected when another input changes the generation.

What mattered more was that the visual setup stayed intact.

The character did not noticeably deform. The bookstore stayed consistent. The outfit remained recognizable.

The speech also appeared in the scene with matching mouth movement.

Adding the audio changed the performance, but it did not break the visual reference structure.

That made the multimodal workflow feel more practical. I did not have to choose between keeping the visual setup stable and adding speech.

add audio and generate new video on canvas

Does First and Last Frame Control Actually Work?

For the next test, I wanted to see whether H3 could build a believable action between two controlled endpoints.

My first attempt at preparing the images exposed an important workflow issue.

I generated the first and last frames independently, but the bookstore background changed too much between them.

That would have made the video test meaningless. H3 would have been asked to reconcile two different environments instead of simply completing an action.

So I changed the workflow.

I generated the first frame first, then used that image as the reference for the last frame.

The character, outfit, bookstore, lighting, and composition stayed the same. Only the action changed.

In the first frame, the woman was looking at the bookshelf.

In the last frame, she was holding a book and starting to walk away.

The first frame cost 15 credits and took 16 seconds to generate.

The last frame cost another 15 credits and took 43 seconds.

H3 Built a Plausible Middle

I then asked H3 to generate an eight-second transition between the two frames at 768P.

The video cost 112 credits and took 3 minutes 51 seconds.

The start and end frames were both preserved well.

More importantly, H3 did not simply morph between two poses. It created a logical middle action in which the woman reached for the book, picked it up, and moved into the ending position.

That made the transition feel like an actual event.

There was one clear weakness.

The character's face became less stable during the middle of the motion.

The overall shot structure worked better than the local facial consistency.

canvas map showing the nodes and the middle shot

Is MiniMax H3 2K Worth It?

I reran the same first-and-last-frame task at 2K.

Nothing else about the task changed.

The 2K version cost 176 credits and took 11 minutes.

That is a large jump from the 768P version, which cost 112 credits and took 3 minutes 51 seconds.

The 2K result had cleaner detail.

In this run, it was also stronger in the area that had been weakest at 768P: facial stability during motion.

That does not prove that higher resolution will always improve facial stability. Video generation can vary between runs.

What it does show is that the 2K result in this test was better in more than visible sharpness.

I still would not use 2K automatically for every draft.

Eleven minutes is a long wait for one short clip, and the credit cost was also much higher.

For early exploration, 768P makes more sense.

For a selected shot where face detail or final output quality matters, 2K may be worth testing before export.

MiniMax H3 vs Seedance 2.5

I also ran the same three-reference bookstore task through Seedance 2.5.

The prompt, character reference, outfit reference, and environment reference stayed the same.

H3 generated the eight-second 768P version for 112 credits in 2 minutes 56 seconds.

Seedance 2.5 generated the comparable 720P result for 264 credits in 3 minutes 21 seconds.

The Seedance result looked visually plausible at first.

The problem appeared during the movement.

Partway through the video, the woman began moving backward in a way that did not fit the intended action.

The prompt described her moving forward through the bookstore. The generated movement weakened that spatial logic.

H3 handled the relationship between the person and the environment more clearly in this test.

The character moved through the bookstore in a way that better matched the intended sequence.

This does not mean H3 will outperform Seedance 2.5 on every task.

For this multi-reference scene, H3 produced the more coherent character-environment relationship while also using fewer credits.

H3
Seedance 2.5

MiniMax H3 Cost and Generation Time in My Tests

The cost difference between workflows was large enough that it deserves its own section.

Test

Credits

Generation time

Hailuo 2.3 baseline

27

1m45s

H3 baseline

84

2m25s

H3 three-image reference

112

2m56s

H3 images + audio

112

4m56s

H3 First / Last Frame 768P

112

3m51s

H3 First / Last Frame 2K

176

11m

Seedance 2.5 comparison

264

3m21s

The reference assets also had their own cost.

The three Flare images used for the multimodal test cost 30 credits total.

The short Speech 2.5 HD audio reference cost another 10 credits.

For the first-and-last-frame workflow, the two frame assets cost 30 credits total.

This is why I would not judge H3 only by the credit number shown on one model node.

The more useful question is how much it costs to reach a usable result.

Several H3 tests produced a usable first result. That can matter more than a cheaper generation that needs repeated rerolls.

I did not run enough repeated generations to calculate a reliable average cost per usable clip, so I would treat that as a workflow principle rather than a benchmark result.

MiniMax H3 Limitations

H3 was not completely stable in every test.

The clearest issue was facial consistency during motion at 768P. The first and last frames stayed strong, while the face drifted during the middle of the transition.

Multimodal generation could also be slow. Adding audio increased the generation time of the bookstore test from 2 minutes 56 seconds to 4 minutes 56 seconds.

The 2K run looked better, but the same short video took 11 minutes to finish.

These tests also used individual outputs. A different generation can behave differently, so the examples show workflow tendencies rather than guaranteed results.

Who Is MiniMax H3 Best For?

H3 made the most sense when the task involved several kinds of control at once.

It is especially useful when a project needs a specific character, outfit, environment, voice, or fixed shot endpoint to survive into the final video.

It also worked well when the shot itself needed structure rather than just motion.

The first-and-last-frame test was a good example. H3 did not simply animate an image. It had to infer the missing action between two controlled states.

H3 is less compelling when the task is only a simple image-to-video clip and speed or cost matters more than reference control.

For that kind of job, the extra capabilities may not justify the additional credits and waiting time.

MiniMax H3 FAQ

Can MiniMax H3 Use Multiple Reference Images?

Yes.

In my Astorie test, I used three references at once: one for identity, one for clothing, and one for the environment.

The first result kept those roles separate well enough to be usable without another generation.

Can MiniMax H3 Use an Audio Reference?

Yes.

In my test, I added a voice clip generated with MiniMax Speech 2.5 HD to the same three-reference bookstore workflow.

H3 incorporated the speech with matching mouth movement while keeping the character, outfit, and environment recognizable.

Is MiniMax H3 2K Better Than 768P?

The 2K result was stronger in my first-and-last-frame test, but it also cost more and took much longer.

It had cleaner detail, and the character's face stayed more stable during motion than in the 768P run.

The 768P video took 3 minutes 51 seconds. The 2K version took 11 minutes.

Because these were individual generations, I would treat the facial-stability difference as an observation from this test rather than a guaranteed effect of 2K.

Is MiniMax H3 Better Than Hailuo 2.3?

They fit different levels of workflow complexity.

Hailuo 2.3 was much cheaper and faster in my baseline test.

H3 produced a clearer performance from the same motion prompt, and its larger advantage appeared once I started using multiple references, audio, and controlled first and last frames.

For a simple image-to-video task, Hailuo 2.3 can still be the more efficient option.

Final Verdict: Is MiniMax H3 Worth Using?

MiniMax H3 felt less like a faster version of Hailuo and more like a broader production model.

Its strongest advantage in my tests was the ability to combine different kinds of control in one workflow.

Multiple references could carry different information. Audio could be added without breaking the visual setup. First and last frames could define the structure of a shot. In the 2K run, both visible detail and facial stability were stronger than in the 768P version.

The trade-off was consistent: more control often meant more credits and longer waits.

For workflows that depend on references and controlled shot design, the upgrade is meaningful.

For a simple image-to-video clip, the extra cost may matter more than the new control.

Ready to try it on the canvas?

Open Astorie and fan your prompt across every frontier model in one workflow.

This website uses cookies

Analytics and marketing tags are on by default in your region — you can turn them off here at any time. We also use basic cookies to keep Astorie secure and remember preferences.

Read more