Skip to main content

Overview

PixVerse V5.5 Text to Video is a text-to-video API that generates a video from a text prompt with native synchronized audio and multi-clip output. Compared to PixVerse V5, version 5.5 adds generate_audio_switch for background music, sound effects, or dialogue, generate_multi_clip_switch for dynamic camera changes within one generation, thinking_type prompt reasoning, and a 10-second duration option.

Key capabilities

  • Native synchronized audio: enable generate_audio_switch to produce audio (music, SFX, or dialogue) together with the video
  • Multi-clip with dynamic cameras: enable generate_multi_clip_switch for multi-clip output with camera changes in a single generation
  • Prompt reasoning (thinking_type): choose enabled (default), disabled, or auto to control whether the model rewrites the prompt before rendering
  • Duration (5, 8, or 10 seconds): 5 (default), 8, or 10; 8s costs double, 10s is available up to 720p, and 1080p is limited to 5 or 8 seconds
  • Resolutions: 360p, 540p, 720p, 1080p
  • Aspect ratios: widescreen_16_9 (default), classic_4_3, square_1_1, traditional_3_4, social_story_9_16
  • Visual styles: anime, 3d_animation, clay, cyberpunk, comic
  • Async processing: poll the task endpoint or receive a webhook notification on completion

Use cases

  • Marketing and ads: short promotional clips with synchronized audio generated from a brief
  • Social content: vertical social_story_9_16 clips with sound for TikTok, Instagram Reels, and YouTube Shorts
  • Story sequences: multi-clip output with camera changes for teasers and narrative shorts
  • Stylized shorts: anime, cyberpunk, or clay looks for art and concept work
  • Prototyping: iterate on prompt and audio variations before production

POST /v1/ai/text-to-video/pixverse-v5-5

Generate a video from a text prompt with PixVerse V5.5

GET /v1/ai/text-to-video/pixverse-v5-5/{task-id}

Get task status and result by ID

GET /v1/ai/text-to-video/pixverse-v5-5

List all PixVerse V5.5 text-to-video tasks

Parameters

Frequently Asked Questions

PixVerse V5.5 adds native synchronized audio (generate_audio_switch), multi-clip output (generate_multi_clip_switch), prompt reasoning (thinking_type), and a 10-second duration option. PixVerse V5 supports only 5 or 8 second durations with no audio or multi-clip.
Set generate_audio_switch to true and PixVerse V5.5 produces synchronized audio (background music, sound effects, or dialogue) together with the video in a single request. No separate audio call is required.
When set to true, PixVerse V5.5 produces multi-clip output with dynamic camera changes inside a single generation, simulating cuts and camera moves without stitching multiple requests.
PixVerse V5.5 supports 5 (default), 8, or 10 second durations and resolutions of 360p, 540p, 720p, and 1080p. 8-second videos cost double, 10-second videos are available up to 720p, and 1080p is limited to 5 or 8 seconds.
thinking_type controls prompt reasoning. enabled (default) rewrites the prompt automatically for better results, disabled uses the prompt exactly as written, and auto lets the model decide based on the input.
Rate limits and pricing depend on your subscription tier. See Rate Limits and the Pricing page for current values.

Best practices

  • Audio: enable generate_audio_switch only when you want the model to author audio; if you have your own track, leave it false and mix externally
  • Duration selection: use 5 seconds for most shots to reduce cost; request 10s only at 720p or lower
  • Prompt reasoning: leave thinking_type as enabled for general prompts; switch to disabled when you need literal prompt adherence
  • Multi-clip: enable generate_multi_clip_switch for narrative sequences that benefit from camera changes
  • Production integration: use webhook_url instead of polling for scalable workflows
  • Error handling: implement retry with exponential backoff for 503 responses