How to Create Consistent AI Characters with MiniMax H3 — Without a Complex ComfyUI Workflow
Quick Answer
The easiest way to keep an AI character consistent is to build the character as a reusable asset before generating complex video shots.
In my test, I first defined a clear identity system for one recurring mini-drama protagonist. I generated one hero image and one multi-view character sheet, cropped the approved sheet into reusable character views, and then used only the references needed for each MiniMax H3 shot.
This workflow did not require building a custom ComfyUI node graph. I created the character assets first, saved the approved views as reusable references in Astorie, and called the relevant references from the canvas when the shot needed them.
The workflow was:
define identity anchors → generate hero image → generate multi-view sheet → crop reusable views → validate the character in simple H3 shots → add view-specific references only when needed → increase cinematic complexity
The biggest lesson was simple:
Consistent characters start with character design, not with a longer video prompt.
What Makes an AI Character Easier to Keep Consistent?
A character prompt should do more than describe someone who looks interesting.
For repeated video production, I need to know which details define the character and which details are allowed to change.
I used a fictional urban mystery protagonist named Mara Lin for this test.
Identity Anchors
I first defined a small group of details that should make Mara recognizable across different images and shots.
Her core identity included:
- a softly angular oval face
- a jet-black asymmetrical bob
- a deep side part
- a small beauty mark near the tail of her left eyebrow
- a silver crescent-shaped ear cuff on her left ear
The goal was not to add as many details as possible.
It was to create a few features that could still be recognized when the pose, framing, or camera angle changed.
The hairstyle gave Mara a strong overall silhouette. The beauty mark and ear cuff gave me smaller details I could check in closer views.
Separate Identity From Wardrobe
I also defined Mara's wardrobe separately from her facial identity.
For this version of the character, she wore:
- a muted burgundy cropped leather jacket
- a cream ribbed knit top
- charcoal straight-leg trousers
- black ankle boots
- a slim black crossbody bag
This distinction matters.
A character's face and hair answer:
Who is this person?
The wardrobe answers:
Which version of this character am I using in this scene?
Expression, pose, lighting, camera angle, and environment belong to another layer. They can change without automatically turning Mara into a different character.
Decide What Must Not Change
Before generating anything, I wrote down the details I planned to check later.
Layer | Example | Keep Fixed? |
Facial identity | Face shape and features | Yes |
Hair | Asymmetrical black bob | Yes |
Facial detail | Beauty mark | Yes |
Accessory | Crescent ear cuff | Yes |
Wardrobe | Burgundy jacket and cream top | Yes for this version |
Expression | Calm, afraid, suspicious | No |
Pose | Standing, walking, turning | No |
Camera angle | Front, 3/4, side | No |
This made later QA more useful.
Instead of asking whether a result merely "looked similar," I could check which layer had changed.
Build the Character Reference Pack
I did not start with H3 video generation.
I first created a compact reference pack that could be reused later.
All character images in this workflow were generated with GPT Image 2.5 Sunburst at 2K.
Step 1: Generate One Hero Image
The hero image became the clearest identity reference.
I kept the composition simple. Mara was shown clearly enough to see her face, hairstyle, ear cuff, beauty mark, and upper-body wardrobe.
I avoided dramatic poses, strong perspective, and complicated backgrounds.
The image had one job:
Show a clean, approved version of the character before motion is introduced.
For the A/B test, I generated two versions.
Version A used a normal descriptive prompt. It described a stylish woman with black bob-length hair, a burgundy jacket, dark trousers, and a modern urban look.
Version B used a structured character definition. It separated facial identity, hairstyle, signature details, wardrobe, and elements that should stay unchanged.
Both were usable.
The difference became clearer when I generated the multi-view sheets.
Step 2: Generate a Multi-View Character Sheet
I generated one sheet for each version.
I used a 16:9, 2K canvas because I planned to crop each view into a separate asset later.
The sheet contained four views:
- front
- three-quarter
- left side profile
- full body
I also left large gaps between the figures.
That mattered more than I expected. A crowded sheet may show more poses, but it becomes harder to crop cleanly.
A sheet
B sheet
Step 3: Choose the Production-Ready Version
A was visually coherent, but it was still a more generic character definition.
B gave me a clearer identity system to carry into production.
This is what I actually checked:
Check | Version A | Version B |
Distinctive hair silhouette | Present, but more generic | Clear asymmetrical bob |
Beauty mark | Not defined in the A prompt | Visible in the approved B assets |
Crescent ear cuff | Not defined in the A prompt | Visible in the approved B assets |
Burgundy wardrobe | Present | Present and clearly defined |
Cross-angle recognizability | Usable | Stronger as one designed character |
I therefore selected B for the rest of the workflow.
The result does not prove that a longer prompt always creates a more consistent character.
It supports a narrower conclusion:
A structured character design gave this project clearer visual features to preserve and evaluate.
Step 4: Crop the Sheet Into Reusable Character Assets
I cropped the approved B sheet into separate character views and saved them as reusable Elements in Astorie.
The final set was:
- @MARA_HERO
- @MARA_FRONT
- @MARA_34
- @MARA_SIDE
- @MARA_FULL
One approved sheet became a reusable character library.
I did not need to regenerate five separate character images from scratch.
How to Use the Character Pack in MiniMax H3
Once the character pack was ready, I moved into video generation.
I did not use every reference in every shot.
Instead, I assigned a practical role to each reference inside this workflow.
That made later troubleshooting easier.
Step 1: Start With the Minimum Useful References
The first video was a simple validation shot.
In this workflow, I used:
- @MARA_HERO primarily for facial identity, hair, beauty mark, and ear cuff
- @MARA_FULL primarily for wardrobe and overall body appearance
- one corridor reference for the environment
The scene was a quiet upscale corridor at night.
Mara looked at her phone, reacted to a sound, and turned slightly.
I kept the movement restrained on purpose.
The goal was not to create the most cinematic shot possible. I wanted to see whether the approved static character could enter video production without obvious identity drift.
It did.
Mara remained recognizable, and the main hairstyle and wardrobe details stayed stable.
The useful rule from this step was:
Start with the smallest reference set that is sufficient for the shot.
Step 2: Add a View-Specific Reference When the Shot Needs It
The second shot introduced more motion and a clearer three-quarter angle.
Mara walked through the corridor at a calm but purposeful pace.
For this shot, I kept the core references and added:
@MARA_34
In this workflow, I used that image because it showed Mara from the angle I was now asking the shot to maintain.
The result remained consistent while she walked.
Her face, hairstyle, outfit, and overall silhouette still matched the approved character pack.
This gave me a second production rule:
Keep the core identity references, then add a matching view when a new camera angle needs information the existing references do not show as clearly.
I did not need to load the entire character library into every prompt.
Step 3: Diagnose the Problem Before Adding More References
The third test failed the first time.
I asked Mara to remain in side profile, turn her head, rotate her shoulders, and finish in an over-the-shoulder pose.
The result produced an obvious visual jump.
One moment showed a clean side profile. The next shifted abruptly into a much more back-facing pose, almost like a hidden cut or body reset.
Importantly, Mara was still recognizable.
The problem was not primarily identity drift.
The motion instructions were conflicting.
I had asked the shot to preserve one orientation while also moving into another.
Instead of adding another reference, I rewrote the choreography as one clear sequence:
face the camera → realize something → turn the whole body continuously → walk toward the far end of the corridor
That version worked.
The turn remained continuous, and Mara stayed consistent through the rotation.
This became the most useful diagnostic lesson from the project:
A broken shot is not always a character-reference problem. Sometimes the character is stable and the choreography is not.
Use a Simple Diagnostic Before You Change the Prompt
After these tests, I found it useful to separate consistency problems into three categories.
Identity Problem
Symptom: The character starts looking like a different person even in familiar angles.
Check the core identity design and the references you are using.
Do not start by changing the camera or adding more action.
Reference-Angle Problem
Symptom: The character looks stable from the front but begins drifting when the camera moves into a less-supported angle.
Add a reference that clearly shows that view.
This is where a three-quarter or side reference can become useful.
Motion Problem
Symptom: The character still looks correct, but the body suddenly jumps, rotates unnaturally, or resets between poses.
Rewrite the choreography first.
Adding more character references will not fix conflicting motion instructions.
This diagnostic saved me from treating every bad generation as the same type of consistency failure.
Move From Consistency Testing to a Real Mini-Drama Shot
The first few videos were intentionally conservative.
That made identity easier to judge, but it also made the clips feel slightly stiff.
There was limited acting, limited camera movement, and little dramatic escalation.
That was useful during validation.
It was not the final production style I wanted.
Once Mara remained stable across simpler shots, I increased the complexity.
Build Suspense With a Timed Multi-Shot Prompt
For the final showcase, I created a 10-second urban suspense sequence.
Instead of asking for one long action, I divided the prompt into timed shots.
The sequence was:
0.0–2.3s — overhead phone close-up
The camera looks down at Mara and the phone in her hands.
She reads something alarming. Her expression tightens, and the hand holding the phone begins to tremble.
2.3–4.3s — extreme eye close-up
The video cuts tightly to her eyes.
The phone glow reflects in them. Her pupils tremble slightly, and the fear becomes more visible.
4.3–6.7s — medium reaction shot
A sudden sound comes from deeper in the corridor.
Mara reacts and turns toward it.
At the same moment, the corridor lights go out.
6.7–9.4s — empty corridor
The visual attention moves away from Mara.
The camera faces the empty hallway and gradually pulls away.
Footsteps echo from deeper in the corridor. A heartbeat builds underneath them while the lights begin flickering irregularly.
9.4–10.0s — final beat
The empty corridor holds briefly before the image falls to black.
This test had a different purpose from the earlier validation shots.
The character was already approved.
That meant I could spend more of the prompt on:
- shot rhythm
- facial performance
- lighting changes
- sound layers
- framing
- suspense escalation
This order made troubleshooting much easier.
If I had started with the most complicated sequence, I would have had more variables to diagnose at once.
What I Learned From This Workflow
Design the Character Before Designing the Shots
The most important consistency work happened before video generation.
The hero image and multi-view sheet established the character first.
The shots came later.
Give Each Reference One Practical Role
In this workflow, I used the hero image mainly for facial identity and the full-body view mainly for wardrobe and overall appearance.
The three-quarter and side views became useful when those angles entered the shot.
These were production choices for this project, not fixed rules about how H3 must interpret every image.
More References Are Not the Default Fix
The failed turning shot proved this clearly.
Mara was still recognizable.
The problem was the movement.
The fix was a cleaner temporal sequence, not another character image.
Validate First, Then Add Cinematic Complexity
Simple shots helped me judge identity.
The final multi-shot sequence tested whether the approved character could survive a more cinematic production workflow.
Those are different stages.
Keeping them separate made the process easier to control.
Common Problems and Fixes
Problem | Likely Cause | Fix |
Character looks generic across references | Character brief has weak identity anchors | Define a recognizable hairstyle, face structure, and signature details |
Sheet is difficult to crop | Too many views or not enough spacing | Use a wide sheet with fewer views and large gaps |
Character drifts at a new angle | Current references do not show that angle clearly | Add a matching view reference |
Turn creates a strange jump | Body-direction instructions conflict | Rewrite the movement as one clear start-to-end sequence |
Character is stable but video feels stiff | Validation prompt is intentionally conservative | Add acting, camera, lighting, sound, and pacing after identity is approved |
Every prompt becomes overloaded with references | References have no clear purpose | Use only the references needed for the current shot |
FAQ
Do I need ComfyUI to keep a character consistent with MiniMax H3?
Not for the workflow I used here.
I created the character assets first, stored the approved views as reusable Astorie Elements, and called those references directly from later H3 shots.
That was enough for this character-building and short-form production workflow.
This article does not cover advanced custom training, LoRA building, or specialized ComfyUI control systems. Those may still be useful when a project needs a different level of customization.
How many character reference images should I prepare?
I created one hero image plus four reusable views:
- front
- three-quarter
- side
- full body
That gave me a small library to choose from.
I did not use all five references in every H3 shot.
Should I use every character reference in every shot?
No.
My first validation shot used the hero and full-body references.
I added the three-quarter reference only when the next shot moved into a three-quarter walking angle.
Use a reference because it solves a specific production problem, not because it exists in the library.
Is a character sheet better than generating each angle separately?
For this project, the sheet was more efficient.
I generated one wide multi-view sheet, checked whether the views still looked like the same person, and then cropped the approved views into reusable assets.
The large spacing between figures made that workflow much easier.
Can I change a character's clothes without losing identity?
Yes, but I would treat the new outfit as a new wardrobe version of the same character.
Keep the core facial and hairstyle identity stable, then define the new clothes separately.
Do not use clothing alone as the test of whether the character is still the same person.
Ready to try it on the canvas?
Open Astorie and fan your prompt across every frontier model in one workflow.