MiniMax H3 Emotion & Acting Guide: Dialogue Tags, Microexpressions, Pauses and Better AI Performances

Create more expressive MiniMax H3 performances by directing dialogue delivery, pauses, breathing, microexpressions, and emotional progression together. This workflow shows how to separate spoken dialogue from acting direction and build emotional changes around specific performance beats.
JoeyUpdated

Quick Answer

Better AI acting is not just about adding words like sad, angry, or crying to a prompt.

In my MiniMax H3 workflow, the strongest performance came from directing four layers together:

dialogue delivery + pauses and breath + microexpressions + emotional progression

There is one important distinction when writing the prompt.

For spoken dialogue, I use the H3 dialogue structure like this:

S1 says softly:
<d>[English] But I wasn't.</d>

The delivery, breathing, facial acting, and other performance instructions stay outside the <d> block.

During my Astorie tests, I also experimented with free-form markers such as <pause>, <breath>, and <softer>. They affected some generations, but I treat them as tested workflow notation, not documented MiniMax H3 dialogue syntax.

A safer reusable structure is:

She avoids looking into the lens.

S1 says:
<d>[English] I told you I was fine.</d>

She pauses and inhales shakily.

Her lower eyelids tighten as she slowly looks up.
Her voice becomes softer and less stable.

S1 says:
<d>[English] But I wasn't.</d>

Her breath catches.
She blinks rapidly as she tries to stop the tears.
Her chin trembles.

Her voice becomes shaky and tearful as she says:
<d>[English] I just... I just didn't want you to see me fall apart.</d>

I reached this structure after testing the same short dialogue several ways. Plain dialogue produced natural breathing and small body movements, but the acting felt flat. More deliberate pauses improved the rhythm. Specific microexpressions made the visual performance much stronger.

The final combined version had the strongest emotional impact. The voice also improved, although H3 still rendered the crying more cleanly than the broken pitch, partial voice loss, and uncontrolled sobbing I originally had in mind.

What Makes an AI Performance Feel Emotional?

A believable performance needs observable behavior, not just a mood label.

Instead of writing:

She is sad and emotional.

describe what that emotion does to the face, voice, and breathing:

Her lower eyelids tighten slightly.
Her inner eyebrows lift.
Her lips press together between phrases.
Her voice becomes less steady.
Her breath catches before the next sentence.

This gives the model something concrete to perform.

Think in Performance Beats, Not One Long Mood

A character who stays equally sad for the whole clip can still feel mechanical.

For my test scene, I used:

"I told you I was fine. But I wasn't. I just... I just didn't want you to see me fall apart."

The dialogue already contains an emotional progression.

"I told you I was fine."
She is still trying to stay composed.

"But I wasn't."
She admits the truth.

"I just... I just..."
Her control starts to fail.

"fall apart."
The emotion becomes harder to hide.

That structure gave me a clear timeline for voice, gaze, breath, and facial changes.

How to Use Dialogue Tags and Pauses in MiniMax H3

It helps to separate two things that can look similar inside a prompt:

dialogue syntax tells H3 what the character says.

acting direction tells H3 how the character behaves while saying it.

Keep Spoken Words Inside the Dialogue Block

A clean dialogue block looks like this:

S1 says softly:
<d>[English] But I wasn't.</d>

The exact spoken words stay inside <d>.

The delivery description stays outside:

Her voice is quiet and unstable.
She hesitates before admitting the truth.

S1 says:
<d>[English] But I wasn't.</d>

I would not write:

<d>[English, crying] But I wasn't.</d>

for the reusable template in this guide.

Instead, I would write what the crying should sound like before the dialogue:

Her voice becomes shaky and tearful.
Her breath is uneven and the final word loses strength.

S1 says:
<d>[English] But I wasn't.</d>

That keeps dialogue content and performance direction separate.

Free-Form Acting Markers Are Different

In my Astorie workflow, I also tested notation such as:

<pause>
<breath>
<inhale>
<catches breath>
<sniff>
<softer>

Some of these changed the generated performance.

For example, after adding pause and breathing cues, I got more varied timing, a visible intake of breath, wetter eyes, and a small gaze change.

However, I do not treat those markers as documented MiniMax H3 dialogue syntax.

They are better understood here as free-form acting notation I tested inside this workflow.

If you want the most portable prompt structure, write the same ideas in plain language:

She pauses.

She inhales audibly before continuing.

Her breath catches.

She gives a small sniff.

Her voice becomes softer.

Put Pauses at Emotional Turning Points

A pause works best when the character has a reason not to continue immediately.

For example:

S1 says:
<d>[English] I told you I was fine.</d>

She pauses and takes a shaky breath.

S1 says:
<d>[English] But I wasn't.</d>

Her breath catches before she tries to continue.

S1 says:
<d>[English] I just...</d>

In my baseline generation, H3 already inserted natural pauses. The problem was that they felt almost equal in length and emotional weight.

Once the pauses were tied to hesitation and admission, the performance gained more shape.

plain dialogue baseline
the version with pause, breath, and softer delivery cues

How to Prompt Microexpressions

Microexpressions made the biggest visual difference in my test.

The useful shift was moving from broad emotional language to specific facial behavior.

Instead of:

She becomes more vulnerable.

I used directions such as:

Her gaze becomes less steady.
Her lower eyelids tighten.
The inner ends of her eyebrows lift slightly.
Her lips press together before she continues.

Describe Specific Parts of the Face

A short set of precise cues is enough.

Emotional moment

Specific acting direction

Holding back tears

Glossy eyes, slight lower-eyelid tension

Hesitation

Gaze wavers, slower blink

Trying not to cry

Rapid blinking, lips pressed together

Emotion breaking through

Chin quiver, lower lip tremble

Avoidance

Gaze drops, head turns slightly away

Losing control

Tears spill, breathing becomes irregular

"Sad" gives context.

"Her chin trembles as she tries to finish the sentence" gives the model an action.

Keep the Microexpressions Moving

In my stronger generations, the emotional change started before the character actually cried.

Her eyes became glossy first. Then her gaze became less steady. Later, she blinked rapidly while trying to hold back tears.

Near the emotional peak, her chin trembled and tears finally appeared.

That gradual build felt much more convincing than a sudden switch from neutral to crying.

Tie Facial Changes to Specific Lines

I got better results when I attached facial cues directly to dialogue beats.

On "I told you I was fine":
She keeps her gaze lowered and tries to stay composed.

Before "But I wasn't":
She slowly raises her eyes.
Her lower eyelids tighten.

On "I just... I just...":
She blinks rapidly two or three times.
Her chin begins to tremble.

On "fall apart":
Her gaze drops and tears run down her cheeks.

This made the face develop with the dialogue instead of performing one fixed expression.

How to Direct Vocal Emotion, Crying, and Voice Breaks

The face and voice did not improve at the same rate in my tests.

At one point, the visual acting was already convincing. The eyes, gaze, mouth, and tears were working.

The voice still sounded too controlled.

That meant I needed to direct vocal behavior separately.

Describe the Sound of a Breaking Voice

Instead of relying on a label like:

crying

I found it more useful to describe the physical sound I wanted:

Her pitch becomes unstable.
Her throat sounds tight.
The ends of some words tremble.
Her voice cracks on the final phrase.
Her breath catches between words.

Other useful vocal descriptions include:

  • shaky breath
  • trembling vowels
  • words losing volume
  • sobs interrupting airflow
  • brief loss of voice

For example:

Her voice becomes softer and less stable.
The pitch wavers from word to word.
A small sob interrupts the airflow without making the dialogue unintelligible.

S1 says:
<d>[English] I just... I just didn't want you to see me fall apart.</d>

In my final generation, the vocal delivery did improve.

The second "I just" became noticeably lighter because of the crying tone, and the voice had more variation than the baseline.

It still did not fully match the sound I originally wanted. I expected stronger pitch breaks, brief loss of voice, and more obvious sobbing.

That became one of the clearest findings from the project: strong visual crying does not automatically produce equally strong vocal breakdown.

Use Repetition as Part of the Performance

This small change worked well:

I just... I just...

The repetition created a natural emotional beat.

It gave the character a place to lose breath, blink rapidly, hesitate, and try again.

That made the dialogue itself part of the acting direction.

Build the Performance as a Timeline

The most useful prompt structure was a short performance timeline.

You do not need to reuse one long prompt. A template is more practical.

Reusable MiniMax H3 Acting Template

[Character + fixed framing]

[Emotional context]

At the beginning:
[initial voice state]
[initial gaze]
[subtle facial tension]

S1 says:
<d>[English] [line 1]</d>

She pauses and inhales.

Before [line 2]:
[voice change]
[gaze change]
[specific microexpression]

S1 says:
<d>[English] [line 2]</d>

Her breath catches.

At the emotional turning point:
[rapid blink]
[chin or lip tremble]
[voice crack]
[tears]

Her voice becomes [specific vocal change] as she says:
<d>[English] [final line]</d>

The exact cues will change with the scene.

The useful structure stays the same:

What changes in the voice?
What changes in the face?
When does the change happen?
What line triggers it?

A Practical MiniMax H3 Acting Workflow in Astorie

For this project, I used a MiniMax H3 video node inside Astorie.

I first created a realistic chest-up portrait of a woman against a white background. My first image already looked close to tears.

That made it a poor control image.

If the starting portrait already contains strong sadness, it becomes harder to tell whether later emotion came from the video prompt or the source frame.

I created a second portrait with a calmer expression and saved it as @actress_calm.

Keep the Shot Controlled

I kept the scene simple:

  • one character
  • frontal chest-up framing
  • plain white background
  • locked camera
  • 10-second clip
  • the same short dialogue

This kept the focus on acting instead of camera movement or scene complexity.

I then branched several MiniMax H3 nodes from the same @actress_calm reference.

The shared upstream image made the comparisons easier to read because the character stayed consistent while the acting direction changed.

What Changed in My Test

The main differences appeared when I changed the type of direction.

Prompt direction

What I observed

Plain dialogue

Natural breathing, shoulder motion, and small head movements, but flat delivery, even pauses, and almost no facial progression

Added pause and breathing direction

More varied timing, a visible intake of breath, some tears, and a small gaze change

Specific microexpressions

Much stronger visual acting, with clearer changes in the eyes, mouth, chin, and gaze

Combined acting and vocal direction

Strongest overall performance and better vocal variation, but the crying voice remained cleaner than true sobbing or partial voice loss

The strongest improvement came from treating the scene as a timed performance rather than adding more emotion adjectives.

Common Problems That Can Break the Performance

Story Context Can Accidentally Add Another Character

One of my failed generations unexpectedly changed into an over-the-shoulder composition.

The prompt said that the woman struggled to face "the person she was speaking to."

The model turned that narrative context into a visible second person.

I changed the direction to:

She avoids looking into the lens.

and made the camera her only addressee.

That kept the emotional meaning without inviting a second character into the frame.

failed generation showing the unexpected foreground person or over-the-shoulder composition

The Background Can Drift During an Emotional Scene

Another generation changed the background even though I wanted the same white studio setup.

For a controlled acting shot, I found it useful to make the visual constraints explicit:

Keep the same frontal chest-up composition.
The background remains plain white.
No camera movement.
No second person.
No reframing.

This keeps scene changes from interfering with the performance.

Tested Acting Notation Can Be Mistaken for Official Syntax

If I write:

<pause>
<inhale>
<softer>

inside an example without explanation, it is easy to read those markers as part of the official H3 dialogue format.

That is not how I use them in this guide.

The reusable dialogue structure remains:

S1 says:
<d>[English] Exact spoken words.</d>

Anything describing breath, pauses, vocal instability, facial acting, or emotional delivery can be written as normal direction around that dialogue block.

That also makes the prompt easier to understand and adapt.

MiniMax H3 Emotion and Acting Prompt Checklist

Before generating an emotional dialogue scene, check that the prompt covers:

  • exact spoken words inside the dialogue block
  • delivery and acting direction outside the dialogue block
  • pauses or breaths with a clear emotional purpose
  • physical microexpressions
  • facial changes tied to dialogue beats
  • an emotional progression from start to finish
  • separate direction for vocal and visual emotion
  • stable framing when composition matters
  • no accidental suggestion of extra characters

You do not need every cue in every scene.

Use the ones that support the emotional arc you actually want.

FAQ

What is the difference between H3 dialogue syntax and acting cues?

Use the dialogue block for the words that should actually be spoken:

S1 says:
<d>[English] But I wasn't.</d>

Write delivery, breathing, and facial behavior outside that block:

She pauses.
Her breath becomes shaky.
Her voice softens before the next line.

In my Astorie tests, I also experimented with free-form markers such as <pause> and <breath>, but I do not present them as documented MiniMax H3 dialogue syntax.

How do I make a MiniMax H3 character cry more naturally?

Build the crying gradually.

Start with smaller cues such as glossy eyes, shallow breathing, lower-eyelid tension, and unstable eye contact. Then introduce rapid blinking, voice breaks, chin tremble, tears, and sobbing at the emotional peak.

Should I use emotion words or detailed acting directions?

Use emotion words to establish context.

Use specific acting directions to control what the audience can actually see and hear.

Ready to try it on the canvas?

Open Astorie and fan your prompt across every frontier model in one workflow.

This website uses cookies

Analytics and marketing tags are on by default in your region — you can turn them off here at any time. We also use basic cookies to keep Astorie secure and remember preferences.

Read more