MiniMax H3 Depth Control Guide: What Changes When AI Video Gets More Spatial Control
Quick Answer
MiniMax H3 Depth Control changes AI video workflows by adding a stronger guide for spatial structure.
A motion prompt can describe creative intent:
- a car accelerates
- a camera follows the subject
- the shot feels fast and cinematic
However, the model still needs to interpret the spatial relationship between the subject, camera, and environment.
Depth reference provides another layer of guidance. It can help preserve:
- subject position in space
- camera movement
- distance changes
- scene structure
In my Formula-style racing shot test, prompt-only generation recreated the feeling of a racing video but did not consistently match the original camera path and vehicle movement. The depth-only workflow produced the closest result to the source clip, preserving the starting position, tracking movement, and turn direction more closely.
The test also showed that adding more motion instructions does not always improve control. When the prompt introduced movement details that conflicted with the depth reference, the output became less faithful.
The biggest change is not creating more motion. It is giving AI video generation a clearer spatial structure to follow.
Why Motion Prompts Alone Are Not Always Enough
Text prompts are good at describing creative decisions.
For example:
Create a fast Formula racing shot with a low tracking camera.
The model understands the intended style.
However, a real camera shot contains many spatial decisions that are difficult to describe with words:
- where the subject starts
- how far it is from the camera
- when it enters the foreground
- how the camera rotates
- how the environment changes during movement
A racing shot is not only about the car moving.
It is also about the relationship between:
- the car
- the camera
- the track
- the surrounding environment
When generating from text alone, the model needs to interpret these relationships itself.
This means a prompt can describe the idea of a shot, but it may not reproduce the exact spatial structure.
Depth Control vs Other AI Video Controls
Different inputs solve different control problems.
Control Method | Best For |
Prompt | Creative direction, style, action description |
Image reference | Appearance and identity consistency |
Pose reference | Subject body position and gesture |
Depth reference | Spatial relationships, camera movement, and scene structure |
The difference is simple:
Prompt tells the model what you want to see.
Depth reference helps show how the scene exists in space.
For example:
A prompt can say:
A race car drives quickly past the camera.
A depth reference can provide:
- where the car starts
- how it approaches
- how the camera follows
- how the distance changes
This makes depth particularly useful for shots where spatial continuity matters.
Creating a Depth Control Test With a Racing Shot
To understand what depth control changes, I created the same Formula-style racing shot with three different workflows.
The goal was not to compare models.
It was to compare different ways of controlling the same generation process.
Test | Input | Question |
Motion prompt baseline | Car references + scene reference + detailed prompt | Can text recreate the shot structure? |
Depth only | Car references + scene reference + depth video | Can depth preserve spatial structure? |
Depth + prompt | Car references + scene reference + depth video + motion instructions | Do extra instructions improve or conflict? |
Preparing the Assets
Car Reference
I created a reusable Formula-style race car reference set:
- hero view
- front view
- three-quarter view
- side view
These references kept the vehicle design consistent across generations.
Track Scene Reference
I created a racing circuit reference containing:
- track layout
- barriers
- fencing
- grandstands
- broadcast-style composition
This kept the environment consistent.
Depth Video Reference
I converted a real racing clip into a depth video reference.
The original clip contained several spatial changes:
- the car approaching from a distance
- the camera following the vehicle
- the car changing position through the turn
This provided a structural reference for the shot.
Test 1: Motion Prompts Can Describe a Shot, but Not Always Preserve It
The first workflow used:
- car references
- track scene reference
- detailed motion prompt
- no depth reference
The prompt described:
- the car approaching the camera
- acceleration through the track
- camera tracking movement
- transition from front three-quarter view to side view
The result captured the general feeling of a racing broadcast.
The car looked dynamic. The scene had the right atmosphere. The camera movement felt energetic.
However, the shot structure was less consistent.
During generation, I observed:
- the car entered from a different direction
- the vehicle position changed during the sequence
- the camera movement did not fully match the source clip
The model understood the idea of a racing shot, but it created its own interpretation of the space.
This shows the limitation of motion prompts:
A prompt can describe movement intent, but it does not always preserve exact spatial relationships.
Test 2: Depth Reference Preserved the Racing Shot Structure More Closely
The second workflow used:
- the same car references
- the same track scene reference
- depth video reference
- no additional motion instructions
In this racing shot test, the depth-only workflow produced the closest match to the original clip.
The generated video better preserved:
- the car’s starting position
- the movement direction
- camera tracking behavior
- the turning motion
The biggest improvement appeared during the corner.
The source clip contained a subtle camera adjustment and vehicle drift during the turn. The depth-guided result maintained this spatial feeling instead of creating a completely different camera interpretation.
This does not mean depth control improves every type of video.
In this workflow, depth was effective because the main challenge was preserving the relationship between:
- the car
- the camera
- the track
The result may vary depending on the source reference and the type of shot.
Test 3: More Motion Instructions Can Create Conflicts
The third workflow combined:
- car references
- track scene reference
- depth video
- additional motion instructions
At first, this seemed like the most controlled setup.
However, the result was less faithful than the depth-only version.
The generated shot changed important parts of the original movement:
- the starting position shifted
- the car turned in a different direction
- the camera path became less consistent
The issue was not a lack of information.
The issue was conflicting information.
The depth reference already contained:
- movement direction
- camera behavior
- spatial progression
The additional prompt introduced another interpretation of the same movement.
The result showed an important workflow lesson:
More control inputs do not always create more control.
When a reference already defines the shot structure, prompts should add creative details instead of rewriting the movement.
[Insert video: motion prompt + depth]
When Should You Use Depth Control?
Depth control is most useful when spatial relationships are central to the shot.
Vehicle Shots
Vehicles depend on:
- movement paths
- camera tracking
- distance changes
A small change in position can make the entire shot feel different.
Tracking Shots
For shots where the camera follows a moving subject, depth can help preserve:
- subject placement
- camera direction
- scene progression
Complex Scene Movement
Depth can be useful when a subject moves through a physical environment and the camera relationship matters.
Depth control is less important for simple scenes.
For example:
- static product shots
- fixed camera compositions
- simple dialogue scenes
In these cases, a detailed prompt or image reference may already provide enough guidance.
FAQ
Does depth control replace motion prompts?
No.
They solve different problems.
Motion prompts define creative intent.
Depth references provide spatial structure.
A strong workflow uses each input for the information it provides best.
Is depth control always better than prompts?
No.
In this racing shot test, depth reference produced a closer spatial match because the main challenge was preserving camera movement and subject position.
For other types of shots, prompts may be enough.
Should I add detailed motion instructions with a depth reference?
Only when they add missing creative information.
If the depth reference already defines:
- movement direction
- camera path
- spatial structure
then repeating those details can create conflicts.
Use prompts to refine the shot, not replace the reference.
What is the biggest change from depth control?
The biggest change is moving from:
Describe the shot and let AI interpret the space.
to:
Provide spatial structure and let AI build within that structure.
For complex AI video workflows, depth control changes the process from generating a scene to designing a more controlled production workflow.
Ready to try it on the canvas?
Open Astorie and fan your prompt across every frontier model in one workflow.