> ## Documentation Index
> Fetch the complete documentation index at: https://docs.magnific.com/llms.txt
> Use this file to discover all available pages before exploring further.

# MiniMax Hailuo 3 API

> Generate AI videos with MiniMax Hailuo 3 at 768p or 2K. Text-to-video, first and last keyframes, and reference-guided generation from up to 9 images, 3 videos and 3 audios, with durations from 4 to 15 seconds.

<Card title="MiniMax Hailuo 3 integration" icon="video">
  MiniMax Hailuo 3 generates video from a prompt, from keyframes, or from a mixed set of reference images, videos and audios that you cite directly in the prompt.
</Card>

MiniMax Hailuo 3 is an AI video generation API served by MiniMax's Video Generation V2 model. It produces videos from a text `prompt` alone, from a first-frame `image` (optionally with a last-frame `image_end`), or from reference assets that steer subject, motion and sound. Durations run from `4` to `15` seconds.

Resolution is selected at generation time: there is a dedicated **768p** and **2K** generate endpoint. Each resolution has its own list and status endpoints.

### Key capabilities

* **Three generation modes**: Text-to-video, keyframe image-to-video, and reference-guided generation
* **Mixed references**: Up to **9** `reference_images`, **3** `reference_videos` and **3** `reference_audios` in a single request
* **Cite references in the prompt**: Address them as `@Image1`, `@Video1`, `@Audio1`, matching their 1-indexed array order
* **First-to-last-frame transitions**: Provide both `image` and `image_end` to generate a transition between the two
* **Durations**: Any integer from `4` to `15` seconds (default `5`) — Hailuo 3 accepts 4-second clips, unlike Hailuo 3 Max
* **Seven aspect ratios**, plus `adaptive` to infer the ratio from the visual references
* **Optional AIGC watermark**: `aigc_watermark` stamps a visible provider watermark
* **Async processing**: Webhook notification or polling for task completion

### Use cases

* **Character consistency**: Supply reference images of a subject and cite them as `@Image1` so the same character carries across shots
* **Motion transfer**: Give a `reference_video` and describe the new scene to reuse its movement
* **Audio-driven scenes**: Pair a reference image with a `reference_audio` to synchronise action to a sound
* **Product animation**: Animate a still with a first-frame `image`, or transition between two product shots
* **Social media content**: Generate vertical `social_story_9_16` clips for TikTok, Instagram Reels and YouTube Shorts

### Generation modes

Keyframes and references are **mutually exclusive**. A request that mixes `image` or `image_end` with any `reference_*` field is rejected.

| Mode | Required input | Optional input | Output behavior |
| - | - | - | - |
| **Text-to-Video** | `prompt` | `aspect_ratio`, `duration`, `aigc_watermark` | Video generated entirely from the text prompt |
| **Image-to-Video** | `prompt` + `image` | `image_end`, `duration`, `aigc_watermark` | Video starts from the first-frame image; the image sets the aspect ratio |
| **Reference-to-Video** | `prompt` + at least one of `reference_images` / `reference_videos` | the other reference fields, `aspect_ratio`, `duration` | References steer subject, motion and sound; cite them in the prompt |

<Warning>
  **Reference videos are billed.** Their combined duration is added to the output duration, capped at 15 seconds. A 5-second generation with one 10-second reference video is billed as 15 seconds, not 5. Reference images and reference audios add nothing to the price.
</Warning>

## API Operations

Generate a video using the `768p` or `2k` endpoint, then track it with the list and status endpoints for that same resolution. Generation returns a task ID for async polling or webhook notification.

<div className="my-11">
  <Columns cols={2}>
    <Card title="POST /v1/ai/video/minimax-h3-768p" icon="video" href="/api-reference/video/minimax-h3/generate-768p">
      Generate a 768p video
    </Card>

    <Card title="POST /v1/ai/video/minimax-h3-2k" icon="video" href="/api-reference/video/minimax-h3/generate-2k">
      Generate a 2K video
    </Card>

    <Card title="GET /v1/ai/video/minimax-h3-768p" icon="list" href="/api-reference/video/minimax-h3/tasks-768p">
      List all 768p tasks
    </Card>

    <Card title="GET /v1/ai/video/minimax-h3-2k" icon="list" href="/api-reference/video/minimax-h3/tasks-2k">
      List all 2K tasks
    </Card>

    <Card title="GET /v1/ai/video/minimax-h3-768p/{task-id}" icon="magnifying-glass" href="/api-reference/video/minimax-h3/task-by-id-768p">
      Get a 768p task by ID
    </Card>

    <Card title="GET /v1/ai/video/minimax-h3-2k/{task-id}" icon="magnifying-glass" href="/api-reference/video/minimax-h3/task-by-id-2k">
      Get a 2K task by ID
    </Card>
  </Columns>
</div>

### Endpoint structure

| Operation | Endpoint |
| - | - |
| **Generate 768p** | `POST /v1/ai/video/minimax-h3-768p` |
| **Generate 2K** | `POST /v1/ai/video/minimax-h3-2k` |
| **List 768p tasks** | `GET /v1/ai/video/minimax-h3-768p` |
| **List 2K tasks** | `GET /v1/ai/video/minimax-h3-2k` |
| **Get 768p task** | `GET /v1/ai/video/minimax-h3-768p/{task-id}` |
| **Get 2K task** | `GET /v1/ai/video/minimax-h3-2k/{task-id}` |

### Parameters

Resolution is determined by the generate endpoint you call (`-768p` or `-2k`) and is not a body parameter. The request body is identical for both endpoints.

| Parameter | Type | Required | Default | Description |
| - | - | - | - | - |
| `prompt` | `string` | Yes | - | Text description of the video. Chinese and English are supported. Up to 7000 characters; for best instruction following keep it under \~500 Chinese characters or \~1,000 English words |
| `image` | `string` | No | - | First-frame image, as a public URL or Base64 string. Mutually exclusive with every `reference_*` field |
| `image_end` | `string` | No | - | Last-frame image, as a public URL or Base64 string. Requires `image`. Mutually exclusive with every `reference_*` field |
| `reference_images` | `array` | No | - | Up to **9** reference images (URL or Base64). Cite as `@Image1`…`@Image9`. JPG, JPEG, PNG, WEBP, HEIC or HEIF; at most 30 MB; 256–5760 px per side; aspect ratio 0.4–2.5 |
| `reference_videos` | `array` | No | - | Up to **3** reference videos (public HTTPS URL or a `upl_vid_...` upload file ID). Cite as `@Video1`…`@Video3`. MP4 or MOV (H.264/H.265); 2–15 s each; 23.976–60 fps; at most 50 MB; 256–5760 px per side; combined duration at most 15 s. **Billed** |
| `reference_audios` | `array` | No | - | Up to **3** reference audios (public HTTPS URL). Cite as `@Audio1`…`@Audio3`. WAV or MP3; 2–15 s each; at most 15 MB; combined duration at most 15 s. Requires at least one reference image or reference video alongside it |
| `duration` | `integer` | No | `5` | Video length in seconds: any integer from `4` to `15` |
| `aspect_ratio` | `string` | No | `widescreen_16_9` | Output ratio: `adaptive`, `film_horizontal_21_9`, `widescreen_16_9`, `classic_4_3`, `square_1_1`, `traditional_3_4`, `social_story_9_16` |
| `aigc_watermark` | `boolean` | No | `false` | Stamps a visible AIGC watermark on the generated video |
| `webhook_url` | `string` | No | - | URL for async status notifications when the task completes |

## Frequently Asked Questions

<AccordionGroup>
  <Accordion title="What is MiniMax Hailuo 3 and how does it work?">
    MiniMax Hailuo 3 is a video generation model served by MiniMax's Video Generation V2 API. You submit a text `prompt` — optionally with keyframes or reference assets — to a resolution-specific generate endpoint and receive a task ID immediately. Poll the GET status endpoint or supply a `webhook_url` to be notified when the task completes, then download the MP4 from the returned URL.
  </Accordion>

  <Accordion title="How do I cite references in the prompt?">
    Reference them by modality and 1-indexed position in their array: the first entry of `reference_images` is `@Image1`, the second `@Image2`, and so on; likewise `@Video1`…`@Video3` and `@Audio1`…`@Audio3`. For example, `"@Image1 walks past @Image2 while @Audio1 plays"`. A reference you never cite still counts toward the limits and, for videos, toward the price.
  </Accordion>

  <Accordion title="Can I combine keyframes and references in the same request?">
    No. `image` and `image_end` belong to the keyframe mode, and `reference_images`, `reference_videos` and `reference_audios` belong to the reference mode. The provider rejects a request that mixes the two families. Pick one mode per request.
  </Accordion>

  <Accordion title="How are reference videos billed?">
    The combined duration of your `reference_videos` is added to the output duration, capped at 15 seconds, and the total is what you are charged. A 5-second generation with one 10-second reference video is billed as 15 seconds. Reference images and reference audios are free — only video references add to the price.
  </Accordion>

  <Accordion title="When should I use the adaptive aspect ratio?">
    `adaptive` asks the model to infer the output ratio from your visual references, so it only means something when the request carries `reference_images`. A text-only request has nothing to infer from and the provider rejects `adaptive` outright — send an explicit ratio, or leave the field out and get `widescreen_16_9`. The field is ignored entirely in keyframe mode, where the output always follows the aspect ratio of the supplied image.
  </Accordion>

  <Accordion title="Can I send a reference audio on its own?">
    No. Audio cannot be the only reference: at least one `reference_images` or `reference_videos` entry is required alongside it.
  </Accordion>

  <Accordion title="What is the difference between Hailuo 3 and Hailuo 3 Max?">
    Hailuo 3 is served by MiniMax directly and accepts mixed references — images, videos and audios — in the same request as keyframes and text. [Hailuo 3 Max](/api-reference/video/minimax-h3-max/overview) is fal's optimised build: it adds `seed`, `prompt_expansion_mode` and `enable_safety_checker`, its minimum duration is `5` seconds instead of `4`, and reference-guided generation lives on its own dedicated endpoints.
  </Accordion>

  <Accordion title="Why does Hailuo 3 accept 7,000 characters when Hailuo 3 Max accepts 50,000?">
    Because they are served by different providers, and each enforces its own limit. Hailuo 3 goes to MiniMax directly, which rejects anything past 7,000 characters with `content[0].text too long (2013)`. [Hailuo 3 Max](/api-reference/video/minimax-h3-max/overview) goes through fal, which accepts up to 50,000. Neither cap is set by this API — raising the Hailuo 3 limit would only move the rejection upstream.

    In practice the ceiling is well above the useful range either way: MiniMax's own guidance is to keep the prompt under roughly 500 Chinese characters or 1,000 English words for the model to follow it reliably.
  </Accordion>

  <Accordion title="How do I choose 768p or 2K?">
    Resolution is selected by the generate endpoint, not by a body parameter. Call `POST /v1/ai/video/minimax-h3-768p` for 768p or `POST /v1/ai/video/minimax-h3-2k` for 2K. Listing and status endpoints are per resolution, so use the pair that matches the endpoint you generated with.
  </Accordion>

  <Accordion title="How long does generation take and how do I get the result?">
    Poll the GET task endpoint or provide a `webhook_url` to be notified on completion. The result MP4 is delivered via a signed URL that stays valid for **30 minutes** — download it promptly, or re-request the task before the link expires.
  </Accordion>

  <Accordion title="What are the rate limits for Hailuo 3?">
    Rate limits depend on your subscription tier. See the [Rate Limits](/ratelimits) page for current limits by plan.
  </Accordion>

  <Accordion title="How much does Hailuo 3 cost?">
    Pricing varies by resolution and billed duration, and reference videos add their combined duration to that total. See the [Pricing](/pricing) page for current rates and credit costs.
  </Accordion>
</AccordionGroup>

## Best practices

* **Cite every reference**: A reference the prompt never mentions still counts toward the limits, and video references still cost — only send what you actually use
* **Keep reference videos short**: Their duration is added to your bill, so trim them to the segment that carries the motion you want
* **Mind the 15-second cap**: Combined reference duration above 15 seconds is trimmed by the provider, so the extra footage is neither used nor useful
* **Prompt writing**: Be specific about scene, subjects, motion, camera movement, lighting and style; place `@Image1`-style citations where the subject appears in the sentence
* **Aspect ratio**: Use `adaptive` only with reference images; in keyframe mode the setting is ignored and the first frame decides
* **Duration selection**: Start at `4` seconds for fast iteration, then raise it for final output
* **Production integration**: Use webhooks instead of polling for scalable applications
* **Result retrieval**: Download the output promptly — the delivery URL is valid for 30 minutes

## Related APIs

* **[MiniMax Hailuo 3 Max](/api-reference/video/minimax-h3-max/overview)**: fal's optimised Hailuo 3 build, with seed control, prompt expansion and dedicated reference-to-video endpoints
* **[MiniMax Hailuo 2.3](/api-reference/image-to-video/minimax-hailuo-2-3-768p/post-minimax-hailuo-2-3-768p)**: The previous Hailuo generation, at 768p and 1080p
* **[Seedance 2.0 Pro](/api-reference/video/seedance-2-pro/overview)**: Alternative high-quality video model with native audio


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.