MiniMax H3 Image Generation and Editing Guide

Separate H3 research capabilities from public static-image endpoints, then follow concrete image-01, still-editing and video-frame workflows.
ClaraUpdated

For MiniMax H3 image generation and editing, distinguish the research model from the public product route: MiniMax’s documented still-image API uses image-01, while the H3 routes verified here generate video. H3 research includes text-to-image and image-to-image tasks, and H3 video can yield extracted stills. This article has not verified a local H3 still-generation or editing workflow. [1–3]

In Astorie, keep the task tied to the node’s actual model: edit the illustration on an image node, approve it, then connect that still to H3 when motion is needed. The supplied project used an image model for color changes and H3 for animation. That separation helps avoid paying for a video when the deliverable is simply a corrected picture.

MiniMax H3 image generation and editing by task

Requested result

Documented route

What is being verified

New still from a description

MiniMax image-01 text-to-image

Native still-image generation

New still guided by one subject image

MiniMax image-01 subject-reference path

New image with subject characteristics

Change a chosen region in a current illustration

A tool with documented masks/inpainting or a manual composite

Exact preserved regions must be checked

Move an approved image

H3 first/last-frame video path

Timed video anchored by source image

Guide a video with mixed examples

H3 reference video path

Video informed by referenced media

Select a still from finished video

Frame extraction

H3-derived frame, not a verified still endpoint

The distinction is about how the result is created, not whether its saved filename ends in PNG or JPEG. A native still request creates an image directly. A video-frame export saves an instant from a generated sequence. A localized edit changes a defined region of an existing image. Each can produce a still file, but they have different controls and evidence requirements.

MiniMax’s image-01 guide uses a text prompt plus aspect ratio for text-to-image, and its subject_reference array for the image-based route. The documented subject-reference example accepts one reference image per request and returns a new image. That is useful for recreating a person in another scene, but it does not promise that an arbitrary local patch will be changed while all other pixels stay identical. MiniMax lists image-01 at $0.0035 per generated image in its public pay-as-you-go table; host subscriptions and outputs can be different. [2, 4]

Generate a still with MiniMax image-01

The following is an implementation recipe from MiniMax’s current guide, not a new hands-on image-01 test. Obtain an API key for the public MiniMax API, keep it in an environment variable, and submit a POST request to https://api.minimax.io/v1/image_generation with Bearer authentication. [2]

Body for a text-to-image request:
{
  "model": "image-01",
  "prompt": "A watercolor illustration of a scholar reading beside a canal, medium shot, warm paper texture",
  "aspect_ratio": "16:9",
  "response_format": "base64"
}

On a successful response, decode each entry in data.image_base64 and save the image. The official example writes JPEG files. Before animating, check the subject, composition and small details at the intended display size. Treat an unsatisfactory pose as an image-stage problem rather than expecting a later video request to repair the illustration automatically.

Use one reference when identity matters

Add a subject_reference array containing one object with type set to character and image_file set to the reference image URL. The documented path supports one reference image per request. Use a clear subject image, then describe the new scene or pose in the prompt. This requests a new image guided by the subject, not a pixel-locked replacement of a selected patch. [2]

Additional field for the subject-reference request:
"subject_reference": [
  {"type": "character", "image_file": "YOUR_REFERENCE_IMAGE_URL"}
]

The URL in this example is a field to replace with your own accessible image, not a submitted test asset. Keep the same image-01 model, aspect ratio and response handling as above. Compare the returned face, hairstyle, clothing and proportions with the source before approving the result.

Edit a specific object before sending the image to H3

For a book-cover recolor, retain a copy of the approved original, select the image-editing node, attach that original and describe the one desired change. Compare the result against the original at equal scale. If unchanged pixels are a delivery requirement, use a masking or compositing process and verify the protected area; a convincing visual match is not a pixel-level guarantee.

In the available Astorie project, GPT Image 2.5 changed a lantern from red to blue and, in a later 2K branch, the scholar’s book from ochre to red. The tester estimated about thirty seconds per edit and saw little other visible difference. Those are useful observations about that image model in this one illustrated scene. They neither prove a pixel-perfect unchanged background nor demonstrate a MiniMax image-01 or H3 still-edit command.

Figure 1. The lantern-color change is on an image node labeled GPT Image 2.5, not on an H3 still-image endpoint.

In this example, the benefit of an image-stage correction is that the approved color becomes part of the source used by the video model. The handoff adds one review step and one asset version, but avoids trying to enforce a still-image correction through a motion prompt. Keep the original and edited versions separately so a rejected animation does not erase the approved illustration.

Animate the approved still with H3

Add or select the still in Astorie, connect a Video node, choose MiniMax H3 and verify its appearance in the Reference Image area. The supplied example displays Omni, 16:9, 768p and eight seconds. Write the desired motion and sound, review the quoted credits, generate, then download the video. A visible reference thumbnail confirms attachment; it does not expose the API role used internally.

If calling the API directly, use first-frame input when the opening picture must be anchored, or reference generation when the image should guide appearance. MiniMax lists 4–15 seconds at 768p or 2K for H3. The image stays a separate approved asset; the new deliverable is a timed sequence. [1]

Figure 2. Existing image connected to an H3 video node; this screenshot concerns a timed video request.

The project owner supplied this workflow’s screenshots and output. Its value here is the verified handoff from an existing still to a video node. The recorded image edits belong to the displayed image model, and the video belongs to H3; neither is presented as an image-01 test.

Extract a usable still from an H3 video

Use the original downloaded MP4, choose the desired moment and export that frame at the video’s native dimensions. For example, the following FFmpeg command extracts one frame at 3.2 seconds; replace the input filename with your own file. It is an extraction recipe, not a new model generation.

ffmpeg -i input.mp4 -ss 00:00:03.200 -frames:v 1 h3-frame.png

Check the frame for motion blur, deformed fingers, half-blinks and transient texture changes. A frame that looks acceptable during playback may be unsuitable as a cover image. The two supplied H3 MP4s are 1344 × 768, so extracting them does not produce native 2K or 4K detail. Increasing a saved image’s dimensions afterward is a separate resize or enhancement step.

A useful caption is “Frame extracted at 00:03.200 from an H3-generated video.” The image is genuinely H3-derived, while that wording makes clear how it was obtained. Save the source MP4 and chosen time alongside it. If the frame is edited later, record that additional model or editing step too.

Choose the route by the final deliverable

Deliverable

Practical route

Acceptance check

New illustration

image-01 text or subject reference

Identity, composition and image detail

Exact local color correction

Mask-capable editor or composite

Requested area changes; protected area matches

Animated illustration

Approved still → H3 video

Continuity, action and soundtrack

Poster moment from an existing clip

H3 MP4 → selected frame

Sharpness and artifacts at that instant

Keep a satisfactory still in its existing image workflow. Move to H3 when motion, shot progression or a soundtrack is part of the brief; those benefits come with temporal and audio review. Use frame extraction when the desired composition already exists in the video, rather than paying for another generation that may change it.

FAQ

Does H3 research mention image tasks? 

Yes. MiniMax lists text-to-image and image-to-image reference/editing in its pretraining description. That is distinct from the public API examples verified here. [3]

Can MiniMax image-01 edit an existing picture? 

The documented subject-reference path creates a new still guided by one clear subject image. The cited guide does not document an arbitrary masked inpainting guarantee. [2]

Is an H3 video frame an H3-derived image? 

Yes. It is a frame from an H3 video, not proof of a separately exposed still-image generation endpoint.

Sources

Official specifications and prices checked September 27, 2026. Image-01 and extraction instructions are documented recipes, not claimed additional tests.

1. MiniMax H3 video route

2. MiniMax image-01 text and subject-reference guide

3. MiniMax H3 research and pretraining tasks

4. MiniMax image-01 and H3 public pricing


Ready to try it on the canvas?

Open Astorie and fan your prompt across every frontier model in one workflow.

This website uses cookies

Analytics and marketing tags are on by default in your region — you can turn them off here at any time. We also use basic cookies to keep Astorie secure and remember preferences.

Read more