Skip to main content

MiniMax Hailuo 3 integration

MiniMax Hailuo 3 generates video from a prompt, from keyframes, or from a mixed set of reference images, videos and audios that you cite directly in the prompt.
MiniMax Hailuo 3 is an AI video generation API served by MiniMax’s Video Generation V2 model. It produces videos from a text prompt alone, from a first-frame image (optionally with a last-frame image_end), or from reference assets that steer subject, motion and sound. Durations run from 4 to 15 seconds. Resolution is selected at generation time: there is a dedicated 768p and 2K generate endpoint. Each resolution has its own list and status endpoints.

Key capabilities

  • Three generation modes: Text-to-video, keyframe image-to-video, and reference-guided generation
  • Mixed references: Up to 9 reference_images, 3 reference_videos and 3 reference_audios in a single request
  • Cite references in the prompt: Address them as @Image1, @Video1, @Audio1, matching their 1-indexed array order
  • First-to-last-frame transitions: Provide both image and image_end to generate a transition between the two
  • Durations: Any integer from 4 to 15 seconds (default 5) — Hailuo 3 accepts 4-second clips, unlike Hailuo 3 Max
  • Seven aspect ratios, plus adaptive to infer the ratio from the visual references
  • Optional AIGC watermark: aigc_watermark stamps a visible provider watermark
  • Async processing: Webhook notification or polling for task completion

Use cases

  • Character consistency: Supply reference images of a subject and cite them as @Image1 so the same character carries across shots
  • Motion transfer: Give a reference_video and describe the new scene to reuse its movement
  • Audio-driven scenes: Pair a reference image with a reference_audio to synchronise action to a sound
  • Product animation: Animate a still with a first-frame image, or transition between two product shots
  • Social media content: Generate vertical social_story_9_16 clips for TikTok, Instagram Reels and YouTube Shorts

Generation modes

Keyframes and references are mutually exclusive. A request that mixes image or image_end with any reference_* field is rejected.
Reference videos are billed. Their combined duration is added to the output duration, capped at 15 seconds. A 5-second generation with one 10-second reference video is billed as 15 seconds, not 5. Reference images and reference audios add nothing to the price.

API Operations

Generate a video using the 768p or 2k endpoint, then track it with the list and status endpoints for that same resolution. Generation returns a task ID for async polling or webhook notification.

POST /v1/ai/video/minimax-h3-768p

Generate a 768p video

POST /v1/ai/video/minimax-h3-2k

Generate a 2K video

GET /v1/ai/video/minimax-h3-768p

List all 768p tasks

GET /v1/ai/video/minimax-h3-2k

List all 2K tasks

GET /v1/ai/video/minimax-h3-768p/{task-id}

Get a 768p task by ID

GET /v1/ai/video/minimax-h3-2k/{task-id}

Get a 2K task by ID

Endpoint structure

Parameters

Resolution is determined by the generate endpoint you call (-768p or -2k) and is not a body parameter. The request body is identical for both endpoints.

Frequently Asked Questions

MiniMax Hailuo 3 is a video generation model served by MiniMax’s Video Generation V2 API. You submit a text prompt — optionally with keyframes or reference assets — to a resolution-specific generate endpoint and receive a task ID immediately. Poll the GET status endpoint or supply a webhook_url to be notified when the task completes, then download the MP4 from the returned URL.
Reference them by modality and 1-indexed position in their array: the first entry of reference_images is @Image1, the second @Image2, and so on; likewise @Video1…@Video3 and @Audio1…@Audio3. For example, "@Image1 walks past @Image2 while @Audio1 plays". A reference you never cite still counts toward the limits and, for videos, toward the price.
No. image and image_end belong to the keyframe mode, and reference_images, reference_videos and reference_audios belong to the reference mode. The provider rejects a request that mixes the two families. Pick one mode per request.
The combined duration of your reference_videos is added to the output duration, capped at 15 seconds, and the total is what you are charged. A 5-second generation with one 10-second reference video is billed as 15 seconds. Reference images and reference audios are free — only video references add to the price.
adaptive asks the model to infer the output ratio from your visual references, so it only means something when the request carries reference_images. A text-only request has nothing to infer from and the provider rejects adaptive outright — send an explicit ratio, or leave the field out and get widescreen_16_9. The field is ignored entirely in keyframe mode, where the output always follows the aspect ratio of the supplied image.
No. Audio cannot be the only reference: at least one reference_images or reference_videos entry is required alongside it.
Hailuo 3 is served by MiniMax directly and accepts mixed references — images, videos and audios — in the same request as keyframes and text. Hailuo 3 Max is fal’s optimised build: it adds seed, prompt_expansion_mode and enable_safety_checker, its minimum duration is 5 seconds instead of 4, and reference-guided generation lives on its own dedicated endpoints.
Because they are served by different providers, and each enforces its own limit. Hailuo 3 goes to MiniMax directly, which rejects anything past 7,000 characters with content[0].text too long (2013). Hailuo 3 Max goes through fal, which accepts up to 50,000. Neither cap is set by this API — raising the Hailuo 3 limit would only move the rejection upstream.In practice the ceiling is well above the useful range either way: MiniMax’s own guidance is to keep the prompt under roughly 500 Chinese characters or 1,000 English words for the model to follow it reliably.
Resolution is selected by the generate endpoint, not by a body parameter. Call POST /v1/ai/video/minimax-h3-768p for 768p or POST /v1/ai/video/minimax-h3-2k for 2K. Listing and status endpoints are per resolution, so use the pair that matches the endpoint you generated with.
Poll the GET task endpoint or provide a webhook_url to be notified on completion. The result MP4 is delivered via a signed URL that stays valid for 30 minutes — download it promptly, or re-request the task before the link expires.
Rate limits depend on your subscription tier. See the Rate Limits page for current limits by plan.
Pricing varies by resolution and billed duration, and reference videos add their combined duration to that total. See the Pricing page for current rates and credit costs.

Best practices

  • Cite every reference: A reference the prompt never mentions still counts toward the limits, and video references still cost — only send what you actually use
  • Keep reference videos short: Their duration is added to your bill, so trim them to the segment that carries the motion you want
  • Mind the 15-second cap: Combined reference duration above 15 seconds is trimmed by the provider, so the extra footage is neither used nor useful
  • Prompt writing: Be specific about scene, subjects, motion, camera movement, lighting and style; place @Image1-style citations where the subject appears in the sentence
  • Aspect ratio: Use adaptive only with reference images; in keyframe mode the setting is ignored and the first frame decides
  • Duration selection: Start at 4 seconds for fast iteration, then raise it for final output
  • Production integration: Use webhooks instead of polling for scalable applications
  • Result retrieval: Download the output promptly — the delivery URL is valid for 30 minutes
  • MiniMax Hailuo 3 Max: fal’s optimised Hailuo 3 build, with seed control, prompt expansion and dedicated reference-to-video endpoints
  • MiniMax Hailuo 2.3: The previous Hailuo generation, at 768p and 1080p
  • Seedance 2.0 Pro: Alternative high-quality video model with native audio