AI Video & Image Model Profiles (2026): segment length, references, audio and list prices
Specs and list prices of 13 models — Seedance, MiniMax H3, Wan 3.0, Kling, Veo, Sora, Runway Gen-4.5, Vidu, PixVerse V6, SeedDream, Qwen-Image — and which platforms expose them.
Segment length and price
Video models: seconds per generation
Published per-second prices (API list + platform prices)
Video model spec table
| Seedance 2.0 / 2.5 | MiniMax Hailuo H3 | Wan 3.0 / 3.0 Prime | Kling 3.0 / Omni | Veo 3.1 | Sora | Runway Gen-4.5 / Aleph | Vidu | PixVerse V6 | LTX-2 | |
|---|---|---|---|---|---|---|---|---|---|---|
| Segment length | 4–15 s (2.0); longer on 2.5 | 4–15 s | 2–30 s | 5–10 s (extendable) | ≈ 8 s (extendable) | Up to ≈ 20 s (per site) | 5–10 s | 4–8 s | 5–10 s (extendable) | Per site |
| Resolution | 480p / 720p (2.0); higher on 2.5 | 768p / 2K | 480p / 720p / 1080p | Up to 1080p | Up to 1080p / 4K (partial) | Up to 1080p | Up to 4K via upscale | Up to 1080p | Up to 1080p | Per site |
| Audio | Native audio (generate_audio) | Native speech + ambience; ≤3 reference clips, ≤15 s total | Native audio; dialogue delimited by quotes, 'no dialogue' as control phrase | Lip-sync; native audio depends on version | Native audio and dialogue | With audio | Audio; custom voices (Pro+) | Limited | Lip-sync | Per site |
| References | First/last frame, reference image/video/audio, multimodal | ≤9 reference images (surcharge beyond 5); first-frame and references are mutually exclusive | Reference image/video/audio, free | Multi-element / subject reference | Ingredients references, first/last frame | Image/video reference, Cameo | Character reference, image-to-video | Multi-subject reference | Character consistency references | Casting references |
| Intra-segment cuts | Executes multi-shot prompts within one generation | 3–6 cuts per segment executed reliably | Multi-shot prompts; execution varies by version | Single-shot oriented | Single-shot; Flow handles continuation | Storyboard tool for multi-shot | Single-shot | Single-shot | Chained on Canvas | Controlled by LTX Studio's storyboard layer |
| List price | Token-billed: 2.0 ≈ ¥28 per 1M tokens, 2.5 ≈ ¥42 (+50%) | API list: CN ¥0.8/s (2K), ¥0.5/s (768p); international $0.13 / $0.08 | API ¥1.2/s (1080p), ¥0.6 (720p), ¥0.3 (480p) | Membership credits; prices not captured (visible on Runway/Luma pages) | Google AI Pro/Ultra credits; API per second ≈ $0.35–0.75/s | Bundled with ChatGPT Free/Plus/Pro (page returned 403) | 60 credits per 5 s ≈ 720/min; Standard 625 credits/mo at $15 | Subscription credits (30 days) / purchased (2 years); prices not captured | API $4.80/min (per site) | Compute credits within LTX Studio subscription |
| Where to use | Volcengine Ark API, Dreamina, and platforms integrating it (SceneMixer, PixVerse, Runway, Luma) | MiniMax API (China / international platforms), Hailuo app, PixVerse, SceneMixer | Model Studio API (workspace endpoint), Tongyi site, wan.video, SceneMixer | Kling China/global sites, API, Runway, Luma, Morph Studio | Flow, Gemini API, Vertex AI, Runway, invideo | sora.com, iOS; API per OpenAI | Runway web / API | vidu.com / vidu.cn, API | PixVerse web/iOS/Android/API/CLI | LTX Studio |
All models
Seedance 2.0 / 2.5
ByteDance's video model, available via Volcengine Ark API and Dreamina; 2.5 shipped July 2026 as Pro/Lite/Turbo.
MiniMax Hailuo H3
MiniMax's video model: multiple references, reference audio, native speech and intra-segment cuts — a common drama-segment engine.
Wan 3.0 / 3.0 Prime
Alibaba's Wan 3.0 video model: up to 30 s per segment, free reference media, audio; Prime is the high-speed tier.
Kling 3.0 / Omni
Kuaishou's Kling video model, known for motion and physics, multi-element reference, integrated by Runway and Luma.
Veo 3.1
Google's video model with native audio and dialogue and strong realism; via Flow, Gemini API/Vertex, Runway and invideo.
Sora
OpenAI's video model with audio, Storyboard and Remix; quota via ChatGPT plans.
Runway Gen-4.5 / Aleph
Runway's in-house generation (Gen-4.5) and editing (Aleph) models.
Vidu
Shengshu's video model, built around reference-to-video and multi-subject consistency.
PixVerse V6
PixVerse's flagship in-house model; the platform also offers Seedance 2.5 and MiniMax H3.
LTX-2
Lightricks' in-house video model powering LTX Studio.
SeedDream 5.0
ByteDance's image model (Pro / Lite) with 2K output and strong understanding of Chinese costume terms — a common choice for character sheets and scene plates.
Qwen-Image 3.0 / 3.0 Pro
Alibaba's Qwen-Image 3.0 family: text-to-image and image editing in one model, pixel area up to 2048×2048.
Gemini 3 Pro Image / 3.1 Flash Image
Google's image models (Nano Banana family), strong on English prompts and modern settings.