MiniMax H3 Prompt Guide: How to Write Prompts That Direct the Shot

Write better MiniMax H3 prompts by separating stable reference information from the action, camera movement, dialogue, sound, ending state, and constraints that define the shot. This workflow shows how to turn a scene description into a clear sequence of events the model can follow.
JoeyUpdated

Quick Answer

A useful MiniMax H3 prompt should read more like a short production brief than a list of visual adjectives.

If you already have character and scene references, let those references carry the stable visual information. Use the prompt mainly to explain what changes over time.

A practical structure is:

References → Opening state → Action sequence → Camera movement → Dialogue and sound → Ending state → Constraints

The key is to describe what the viewer should actually see happen.

Instead of writing:

Carrie moves naturally through the room with cinematic motion.

Write the sequence:

Carrie starts beside the sofa, turns toward the side of the room, takes two steps forward, looks back toward the camera, speaks, and stops.

That gives the shot a clear beginning, progression, and ending.

What Makes a Good MiniMax H3 Prompt?

A video prompt should describe change, not only appearance.

A sentence like:

Carrie stands in a warm living room.

defines a visual state. It does not say what should happen after the video starts.

For a short video, think in time.

Ask:

  • Where does the subject start?
  • What happens first?
  • What changes next?
  • How does the camera respond?
  • When does dialogue happen?
  • Where does the motion end?

This turns the prompt from a scene description into shot direction.

How to Write a MiniMax H3 Prompt Step by Step

Step 1: Let References Handle Stable Visual Details

Start by deciding which information is already defined by your references.

In my Astorie workflow, I used three existing Elements:

  • @CHAR_01_HERO for Carrie's facial identity
  • @CHAR_FULL for her full-body appearance
  • @SCENE_01_MED for the living room

Because those references already showed the character and environment, I did not need to describe every visual detail again.

This creates a useful division:

References define what should stay recognizable. The prompt defines what needs to happen next.

That keeps the prompt focused.

@CHAR_01_HERO, @CHAR_FULL, and @SCENE_01_MED reference Elements

Step 2: Define the Opening State

Give the action a clear starting point.

For this project, I wrote:

At the beginning, Carrie stands beside the sofa in a relaxed pose.

This does not need to be complicated.

The opening state only needs to establish the subject's position and condition before the motion begins.

Without that starting point, instructions such as "walk to the side" have less spatial context.

Step 3: Write Actions as a Timeline

Avoid grouping several actions into one vague instruction.

I wanted Carrie to move across the room, turn back, speak, and then stop.

So I wrote those actions in order:

She turns slightly toward the side of the room, takes two natural steps forward, then slows down and looks back toward the camera.

The useful pattern is:

Start → Transition → Key action → End

Temporal words help make that sequence clearer:

  • at the beginning
  • then
  • as she...
  • near the end

This is more useful than adding extra adjectives such as "dynamic," "cinematic," or "natural."

Those words may describe a feeling, but they do not explain what should physically happen.

Step 4: Give the Camera One Main Job

Camera instructions should support the subject's movement.

For this shot, I used:

Use one subtle lateral tracking movement that follows Carrie as she walks.

The important part is not the term "tracking."

The instruction explains the relationship between the camera and Carrie.

She moves toward the side of the room, and the camera follows that movement.

Avoid stacking several camera commands into one short shot.

For example:

tracking shot, pan, zoom, dolly, dynamic cinematic camera

does not create a clear camera plan.

For a simple shot, one primary camera behavior is usually easier to direct.

Step 5: Attach Dialogue to the Action

Do not treat dialogue as a completely separate script line when timing matters.

Instead of writing:

Carrie says, "I'll be right there."

I connected the line to the moment when she turns back:

As she looks back, she says clearly: "I'll be right there."

This tells the prompt both what she says and when she says it.

For this test, I intentionally kept the dialogue direction simple.

I did not add emotion tags, vocal performance cues, pauses, or microexpressions. I wanted to test the structure of the shot rather than acting control.

Step 6: Separate Scene Sound From Music

Treat environmental sound and music as separate decisions.

For this scene, I requested:

Include quiet indoor room ambience, soft footsteps, and subtle clothing movement.

I then added:

Do not add background music.

This makes the audio intent easier to read.

The first instruction describes sounds that belong inside the scene. The second controls whether music should sit over the scene.

You do not need to fill every video prompt with sound instructions. Add them when audio matters to the finished shot.

Step 7: Give the Video an Ending State

Many prompts explain how an action starts but never say where it should finish.

For this video, I added:

Near the end, Carrie stops moving and holds a natural standing pose while looking toward the camera.

This gives the movement a destination.

That matters in short clips because the action should not feel like it is still unfolding when the video ends.

A useful prompt should direct not only how motion begins, but where it settles.

Step 8: Protect Only the Details That Matter

Finish with constraints for details that should remain stable.

For this test, I protected:

  • character identity
  • outfit
  • hairstyle
  • facial features
  • furniture placement
  • lighting
  • room appearance

I did not add a long list of generic quality words.

Constraints are most useful when they protect something specific.

"Keep her hairstyle unchanged" is actionable.

"Make everything perfect and high quality" is not.

The Full MiniMax H3 Prompt I Used

Here is the complete prompt from my test:

Use @CHAR_01_HERO, @CHAR_FULL, and @SCENE_01_MED as references.
Create a realistic 6-second single-shot video in the same indoor living room.
Keep Carrie’s facial identity, hairstyle, outfit, body appearance, and the room layout consistent with the references.
At the beginning, Carrie stands beside the sofa in a relaxed pose.
She turns slightly toward the side of the room, takes two natural steps forward, then slows down and looks back toward the camera.
As she looks back, she says clearly: “I’ll be right there.”
Use one subtle lateral tracking movement that follows Carrie as she walks. Keep the camera movement slow, smooth, and controlled. Do not add extra zooms, pans, or dramatic camera changes.
Include quiet indoor room ambience, soft footsteps, and subtle clothing movement.
Do not add background music or other non-diegetic music.
Near the end, Carrie stops moving and holds a natural standing pose while looking toward the camera.
Keep the movement realistic and restrained. Do not change her identity, outfit, hairstyle, facial features, furniture placement, lighting, or overall room appearance.

The prompt is longer than a simple scene description, but every section has a job.

It defines the references, starting position, action order, camera behavior, dialogue timing, sound, ending state, and details that should stay unchanged.

full prompt inside the MiniMax H3 node in Astorie

What Happened in My First Generation

The first generation completed the main shot without another prompt revision.

Carrie walked toward one side of the room. The camera followed her movement, and she said:

“I’ll be right there.”

The action sequence and camera direction were close enough to the intended shot that I did not need to regenerate it.

There was one clear weakness.

Her voice was flat, and her facial expression stayed close to a simple smile. The performance felt mechanical even though the physical instructions were completed.

I did not treat that as a failure of the prompt structure.

This test was designed to control the shot, not to create a nuanced emotional performance.

That distinction matters.

Prompt Structure and Acting Direction Are Different Jobs

A structurally clear prompt can tell the model what happens without fully controlling how the character performs it.

That is exactly what happened in my test.

The prompt successfully directed:

  • walking
  • camera tracking
  • looking back
  • dialogue timing
  • the final stopping position

But I had not told Carrie how the line should feel.

There were no cues for hesitation, urgency, warmth, sadness, tension, or any other specific performance.

So the simple smile and flat delivery were not surprising.

If acting quality matters to the shot, add performance direction as a separate layer.

That may include details such as a pause before speaking, a change in facial tension, a softer vocal delivery, or another observable performance cue.

Do not assume that a well-structured action prompt will automatically create nuanced acting.

When Should You Rewrite the Prompt?

Do not regenerate a video just because the first result is imperfect.

First ask:

Did the missing behavior affect the goal of this shot?

For my test, the goal was to see whether the prompt structure could direct the action, camera, dialogue, and ending.

It did.

The acting was mechanical, but improving the performance would have answered a different question. I therefore stopped after the first usable generation.

Revise the prompt when an important instruction fails, such as:

  • a required action is missing
  • the actions happen in the wrong order
  • the camera does not follow the intended movement
  • dialogue happens at the wrong moment
  • the subject does not reach the intended ending position
  • an unwanted element changes the meaning of the shot

A minor issue does not always justify another generation.

Only change the prompt when the change helps you reach the actual production goal.

Common MiniMax H3 Prompt Mistakes

Repeating Everything in the References

If your references already establish the subject and environment, repeating every visual detail can make the prompt unnecessarily dense.

Use that space for motion, timing, camera direction, and other information the references cannot show.

Listing Actions Without an Order

Writing:

turn, walk, look back, speak

identifies the actions but does not clearly organize them.

Write the progression instead:

She turns, takes two steps, then looks back and speaks.

The order is part of the direction.

Giving the Camera Too Many Instructions

More camera vocabulary does not automatically create better movement.

A short shot often needs one clear camera relationship.

Explain what the camera should do in relation to the subject.

Separating Dialogue From Its Moment

If a line belongs to a specific action, connect them.

"As she looks back, she says..." is clearer than placing the dialogue somewhere else in the prompt with no timing context.

Forgetting How the Shot Ends

A video prompt needs an ending as well as a beginning.

Tell the subject where to stop, what pose to hold, or what framing should remain at the end.

Using Emotion Words Without Performance Direction

Words such as "sad," "happy," or "dramatic" describe broad emotional categories.

They do not necessarily explain what the viewer should see or hear.

When emotional performance matters, describe the specific behavior you want rather than relying only on a label.

FAQ

How Long Should a MiniMax H3 Prompt Be?

A prompt should be long enough to make the important relationships clear.

You do not need to maximize word count.

Include the information needed to direct the references, action timeline, camera, dialogue or sound, ending state, and important constraints.

Remove details that do not change the shot.

Should I Describe Everything in My Reference Images Again?

Usually, no.

If the reference already defines the character's appearance or the environment, use the prompt for information the image cannot show.

Motion, timing, dialogue, camera behavior, and ending position are good examples.

Should I Use Emotion Tags in Every MiniMax H3 Prompt?

No.

Use emotion or performance cues when acting quality matters to the result.

For a simple movement or camera-control shot, adding detailed emotional direction may be unnecessary.

In my test, I intentionally left those cues out because the goal was to test shot structure.

Can I Put Multiple Shots in One Prompt?

You can describe more complex sequences, but the prompt becomes harder to organize as the number of shot changes increases.

For this guide, I used a single continuous shot so the relationship between subject movement, camera behavior, dialogue, and the ending remained easy to inspect.

Why Did MiniMax H3 Follow My Actions but the Acting Still Feel Flat?

Action structure and performance direction solve different problems.

In my test, Carrie completed the requested movement and dialogue, but her voice stayed flat and her smile felt mechanical.

The prompt explained what she should do. It did not contain detailed instructions for how she should emotionally perform the line.

If performance matters, add that direction deliberately rather than expecting it to emerge from the action structure alone.

Ready to try it on the canvas?

Open Astorie and fan your prompt across every frontier model in one workflow.

This website uses cookies

Analytics and marketing tags are on by default in your region — you can turn them off here at any time. We also use basic cookies to keep Astorie secure and remember preferences.

Read more