Chinese AI Video Model Comparison

Share:

Chinese AI video model comparison: Seedance, Kling, Hailuo, Vidu — duration, resolution, access. Provisional. In your browser.

RT-AI-068 · AI Tools

Chinese AI Video Model Comparison

One row per MODEL VERSION, not per vendor. A Chinese lab routinely sells several versions at once with incompatible limits, so collapsing a vendor into a single row would print a number that is true of nothing it ships.

Every figure here is read off the vendor's own API reference or price list and carries the link and the date it was checked. Where the vendor publishes nothing, the cell says so. Where nobody has looked yet, it shows a dash. Nothing is guessed.

39 39 model versions from 9 vendors, read live from our AI directory.

How old this data is The stalest claim on this page was checked 31 days ago. Anything past 120 days is treated as stale and fails tools:check-video-comparison.
How to read this table
  • 15 s A value: stated in the vendor's own documentation. Open "Evidence & full specs" to check it yourself.
  • Not published Not published: we looked and the vendor does not say. That is not the same as "no".
  • A dash: nobody has researched that field yet.
  • * A star: the value carries a qualification — "1080P at 6 seconds only" and the like. The qualification lives in that row's "Evidence & full specs". Read it before you use the number.
  • A triangle: the value is coupled to another figure on the same row. The rule that couples them stays on screen under the row and is never folded away.

On a narrow screen the grid shows the headline specs. Every remaining figure is in each row's "Evidence & full specs" — nothing is dropped.

Chinese AI video models compared version by version
Model version Max clip Fixed lengths Max res. Frame rate Native audio Modes Price (global) Price (CN) Public API Open weights
Kling 3.0 Kuaishou Current 15 s Not researched yet 4K qualified Not published qualified Yes qualified T2V · I2V · FL2V · REF2V qualified $0.084 / s qualified Not researched yet Yes Not researched yet
Verified 30 days ago
Evidence & full specs (5 sources)
Max clip
15 s
Max resolution
4K resolution enum 720p | 1080p | 4k; 4K bills at $0.42/s against $0.084/s at 720p
Frame rate
Not published Kling publishes no frame rate — its own Video Capability Map lists duration and resolution and no fps, and no endpoint takes an fps parameter
Native audio
Yes opt-in: settings.audio enum native | off, default off
Modes
T2V · I2V · FL2V · REF2V ref2v runs through the Element library rather than raw reference images
Price (global)
$0.084 / s 720p without native audio. 1080p $0.112/s, 4K $0.42/s; with native audio 720p $0.126/s, 1080p $0.168/s. Billed in Units, 1 Unit = $0.14
Public API
Yes
Min clip
3 s
Watermark
optional options.watermark_info.enabled defaults false; when true a watermarked copy is returned alongside the clean one
Availability
global, single Singapore API endpoint api-singapore.klingai.com is the only regional endpoint documented on the international open platform
Negative prompt
No the current endpoint has no negative_prompt field — positive and negative descriptions both go in the prompt. The superseded legacy endpoint did expose one
Max input images
2 first_frame + last_frame; last-frame-only is not supported. Up to 3 Elements may additionally be referenced

Sources (5)

Directory entry: Kling
Kling 3.0 Omni Kuaishou Current 15 s Not researched yet 4K Not researched yet Yes qualified T2V · I2V · FL2V · REF2V · V2V qualified $0.084 / s qualified Not researched yet Yes Not researched yet
Verified 30 days ago
Evidence & full specs (4 sources)
Max clip
15 s
Max resolution
4K
Native audio
Yes settings.audio enum native | off, default off
Modes
T2V · I2V · FL2V · REF2V · V2V video reference and multi-image to video are marked supported on Omni and O1 only
Price (global)
$0.084 / s 720p, no video input, no native audio. With native audio 720p $0.112/s, 1080p $0.14/s; with video input 720p $0.126/s, 1080p $0.168/s; 4K $0.42/s throughout
Public API
Yes
Min clip
3 s
Watermark
optional
Negative prompt
No

Sources (4)

Directory entry: Kling
Kling 3.0 Turbo Kuaishou Current 15 s Not researched yet 1080p qualified Not researched yet Yes qualified T2V · I2V qualified $0.112 / s qualified Not researched yet Yes Not researched yet
Verified 30 days ago
Evidence & full specs (4 sources)
Max clip
15 s
Max resolution
1080p enum is 720p | 1080p only — no 4K on Turbo, unlike Kling 3.0
Native audio
Yes always on — the endpoint exposes no audio toggle and the price list carries only a 'with native audio' row
Modes
T2V · I2V the capability map shows no first/last frame and no Element control for Turbo
Price (global)
$0.112 / s 720p with native audio (0.8 Units/s); 1080p $0.14/s
Public API
Yes
Min clip
3 s
Watermark
optional
Negative prompt
No

Sources (4)

Directory entry: Kling
Kling O1 Kuaishou Current 10 s Not researched yet 1080p Not researched yet No qualified T2V · I2V · FL2V · REF2V · V2V qualified $0.084 / s qualified Not researched yet Yes Not researched yet
Verified 30 days ago
Evidence & full specs (3 sources)
Max clip
10 s
Max resolution
1080p
Native audio
No the audio enum is original | off — 'original' RETAINS an input video's sound; the model synthesises none, and the price list has no audio dimension for O1
Modes
T2V · I2V · FL2V · REF2V · V2V content types: prompt, first_frame, last_frame, refer_image, feature_video, base_video, element
Price (global)
$0.084 / s 720p, no video input; 1080p $0.112/s. With video input 720p $0.126/s, 1080p $0.168/s
Public API
Yes
Min clip
3 s integers 3-10; with only a first frame and no other reference, only 5 or 10 are accepted
Watermark
optional
Negative prompt
No

Sources (3)

Directory entry: Kling
Kling 2.6 Kuaishou Superseded 10 s 5 / 10 s qualified 1080p Not researched yet Yes qualified T2V · I2V · FL2V · S2V qualified $0.042 / s qualified Not researched yet Yes Not researched yet
Verified 30 days ago
Evidence & full specs (4 sources)
Max clip
10 s
Fixed lengths
5 / 10 s the API duration enum is exactly {5, 10} on both t2v and i2v. Kling's own Video Capability Map advertises 3~10s for this model; that range is NOT reachable through the documented API
Max resolution
1080p
Native audio
Yes audio=native FORCES 1080p — no 720p rate exists for the audio tiers
Modes
T2V · I2V · FL2V · S2V s2v is voice-reference input (at most 2 voices) and requires audio != off; first+last frame requires 1080p
Price (global)
$0.042 / s 720p silent; 1080p silent $0.07/s; with native audio (1080p only) $0.14/s; with voice control $0.168/s
Public API
Yes
Watermark
optional
Negative prompt
No

Sources (4)

Directory entry: Kling
Kling 2.5 Turbo Kuaishou Superseded 10 s 5 / 10 s 1080p Not researched yet No qualified T2V · I2V · FL2V $0.042 / s qualified Not researched yet Yes Not researched yet
Verified 30 days ago
Evidence & full specs (4 sources)
Max clip
10 s
Fixed lengths
5 / 10 s
Max resolution
1080p
Native audio
No no audio parameter on this endpoint; the price list carries only a 'no native audio' row
Modes
T2V · I2V · FL2V
Price (global)
$0.042 / s 720p; 1080p $0.07/s
Public API
Yes
Watermark
optional

Sources (4)

Directory entry: Kling
Dreamina Seedance 2.5 ByteDance Current 30 s qualified Not researched yet 1080p qualified 24 fps Yes qualified T2V · I2V · FL2V · REF2V · V2V · S2V qualified $0.231 / s qualified Not researched yet Yes qualified Not researched yet
Verified 31 days ago
Evidence & full specs (3 sources)
Max clip
30 s duration accepts [4, 30] or -1 for smart duration (the default). Video-EDITING tasks accept only -1 — you cannot specify an output length there
Max resolution
1080p 480p | 720p | 1080p, default 720p. NO 4K on 2.5 — only Seedance 2.0 offers it, so 'Seedance: up to 4K' is wrong for the current flagship. 1080p output is 10-bit H.265, which some players cannot decode
Frame rate
24 fps
Native audio
Yes generate_audio defaults to TRUE; output audio is always mono regardless of input channels
Modes
T2V · I2V · FL2V · REF2V · V2V · S2V omni reference takes 0-30 images, 0-10 videos and 0-10 audio clips; 2.5 is the only version accepting audio-ONLY input
Price (global)
$0.231 / s vendor's own worked example: 720p 16:9 5s with no video input. 480p $0.103/s, 1080p $0.569/s. Billing is per TOKEN underneath, so reference video raises the bill sharply — 1080p with 30s of input video reaches $11.907 for a 5s output
Public API
Yes gated: needs an account balance over USD 30, a USD 30+ AI Savings Plan, or a purchased Seedance resource pack
Min clip
4 s
Watermark
optional watermark boolean defaults FALSE; true stamps 'AI Generated' bottom-right. This is the BytePlus international endpoint — mainland AIGC labelling obligations attach to the Volcano Engine service instead
Availability
ap-southeast-1 only base URL ark.ap-southeast.bytepluses.com. Volcano Engine 火山引擎 is the separate mainland-China console for the same models, with different model IDs, pricing and labelling duties
Negative prompt
No no negative-prompt parameter anywhere in the request body; steering is positive-only
Max input images
30 omni reference accepts 1-30. First-frame i2v takes 1 and first+last takes 2, and those modes are MUTUALLY EXCLUSIVE with omni reference

Released 2026-06-28

Sources (3)

Directory entry: Seedance
Dreamina Seedance 2.0 ByteDance Current 15 s qualified Not researched yet 4K qualified 24 fps Yes T2V · I2V · FL2V · REF2V · V2V qualified $0.15 / s qualified Not researched yet Yes Not researched yet
Verified 31 days ago
Evidence & full specs (3 sources)
Max clip
15 s [4, 15] or -1
Max resolution
4K 480p | 720p | 1080p | 4k — the ONLY Seedance version offering 4k, and it is rate-limited far harder (15 RPM, concurrency 1) and encoded 10-bit H.265
Frame rate
24 fps
Native audio
Yes
Modes
T2V · I2V · FL2V · REF2V · V2V omni reference takes 0-9 images, 0-3 videos, 0-3 audio; audio-only input is NOT supported, unlike 2.5
Price (global)
$0.15 / s 720p 16:9 5s, no video input. 480p $0.07/s, 1080p $0.37/s, 4K $0.78/s
Public API
Yes
Min clip
4 s
Watermark
optional
Availability
ap-southeast-1 only
Negative prompt
No
Max input images
9

Released 2026-01-28

Sources (3)

Directory entry: Seedance
Dreamina Seedance 2.0 Fast ByteDance Current 15 s Not researched yet 720p qualified 24 fps Yes T2V · I2V · FL2V · REF2V · V2V $0.12 / s qualified Not researched yet Not researched yet Not researched yet
Verified 31 days ago
Evidence & full specs (3 sources)
Max clip
15 s
Max resolution
720p 480p | 720p only — no 1080p and no 4k
Frame rate
24 fps
Native audio
Yes
Modes
T2V · I2V · FL2V · REF2V · V2V
Price (global)
$0.12 / s 720p 16:9 5s, no video input; 480p $0.06/s. A limited-time 25% discount ran 7 Aug - 7 Sep 2026
Min clip
4 s
Max input images
9

Released 2026-01-28

Sources (3)

Directory entry: Seedance
Dreamina Seedance 2.0 Mini ByteDance Current 15 s Not researched yet 720p qualified 24 fps Yes T2V · I2V · FL2V · REF2V · V2V $0.08 / s qualified Not researched yet Not researched yet Not researched yet
Verified 31 days ago
Evidence & full specs (3 sources)
Max clip
15 s
Max resolution
720p 480p | 720p only
Frame rate
24 fps
Native audio
Yes
Modes
T2V · I2V · FL2V · REF2V · V2V
Price (global)
$0.08 / s 720p 16:9 5s, no video input; 480p $0.04/s. Cheapest Seedance tier; a 60% discount ran 7 Aug - 7 Sep 2026
Min clip
4 s

Released 2026-06-15

Sources (3)

Directory entry: Seedance
Seedance 1.5 Pro ByteDance Superseded 12 s qualified Not researched yet 1080p qualified 24 fps Yes T2V · I2V · FL2V qualified Not researched yet Not researched yet Yes qualified Not researched yet
Verified 31 days ago
Evidence & full specs (2 sources)
Max clip
12 s [4, 12] or -1
Max resolution
1080p 480p | 720p | 1080p, default 720p
Frame rate
24 fps
Native audio
Yes
Modes
T2V · I2V · FL2V no reference-to-video and no video input — those arrived with 2.0. Uniquely supports a cheap draft mode, re-run from the draft's task ID
Public API
Yes the only Seedance version also offered on the cheaper flex offline-inference tier
Min clip
4 s
Max input images
2 first_frame + last_frame

Released 2025-12-15

Sources (2)

Directory entry: Seedance
Seedance 1.0 Pro (and 1.0 Pro Fast) ByteDance Superseded 12 s qualified Not researched yet 1080p qualified 24 fps No qualified T2V · I2V · FL2V qualified Not researched yet Not researched yet Not researched yet Not researched yet
Verified 31 days ago
Evidence & full specs (2 sources)
Max clip
12 s [2, 12] — the only versions reaching down to 2s
Max resolution
1080p 480p | 720p | 1080p, and uniquely DEFAULTS to 1080p rather than 720p
Frame rate
24 fps
Native audio
No generate_audio is supported only on 2.5, the 2.0 series and 1.5 Pro — the 1.0 models are silent-only
Modes
T2V · I2V · FL2V 1.0 Pro Fast drops first+last frame — text-to-video and first-frame image-to-video only
Min clip
2 s

Released 2025-05-28

Sources (2)

Directory entry: Seedance
Dreamina app — Seedance 2.5 ByteDance Current 30 s qualified Not researched yet Not published qualified Not researched yet Yes T2V · I2V · REF2V · S2V qualified Not researched yet Not researched yet No qualified Not researched yet
Verified 31 days ago
Evidence & full specs (2 sources)
Max clip
30 s the app's standard mode, matching the API's single-generation cap. The page also advertises extension to 180 seconds in a beta long-video mode, which is stitching rather than one continuous generation
Max resolution
Not published unresolved between two vendor pages: the Dreamina product page claims 4K, while the ModelArk API caps dreamina-seedance-2-5 at 1080p and exposes 4k only on Seedance 2.0. Both are ByteDance's own surfaces, so the conflict IS the finding — recorded as unknown rather than picking a side
Native audio
Yes
Modes
T2V · I2V · REF2V · S2V the app accepts text prompts, scripts, reference photos, videos, music and style guides, plus a green-screen reference workflow
Public API
No Dreamina is a credit-billed consumer app and publishes no API of its own. The same models are exposed programmatically through BytePlus ModelArk — see the `seedance` entry
Watermark
free-tier watermark-free download is a paid upgrade
Availability
global Dreamina is the international front door; 即梦 at jimeng.jianying.com is the mainland-China one, with separate accounts

Sources (2)

Directory entry: Jimeng AI 即梦 (Dreamina)
Dreamina app — Seedance 2.0 ByteDance Current Not researched yet Not researched yet 4K qualified Not researched yet Yes qualified T2V · I2V · REF2V · S2V Not researched yet Not researched yet No Not researched yet
Verified 31 days ago
Evidence & full specs (2 sources)
Max resolution
4K unlike the 2.5 claim, this one IS corroborated by the API, where dreamina-seedance-2-0 accepts resolution=4k
Native audio
Yes synchronises narration, dialogue, sound effects and visuals
Modes
T2V · I2V · REF2V · S2V
Public API
No
Watermark
free-tier
Max input images
9 up to 12 reference resources — 9 photos, 3 videos, 3 audio clips — exactly matching the ModelArk limits for the 2.0 series

Sources (2)

Directory entry: Jimeng AI 即梦 (Dreamina)
即梦 Jimeng (mainland China app) ByteDance Current Not published qualified Not researched yet Not researched yet Not researched yet Not researched yet T2V · I2V · REF2V · S2V qualified Not researched yet Not researched yet No qualified Not researched yet
Verified 31 days ago
Evidence & full specs (1 sources)
Max clip
Not published the Jimeng and Dreamina generator UIs are behind a login wall and neither product publishes a specification page or help-centre article with the in-app duration picker; the model-level limits live on the `seedance` entry instead
Modes
T2V · I2V · REF2V · S2V the widget labels the mode 全能参考; 数字人 and 动作模仿 are separate tools
Public API
No consumer app; the mainland programmatic route is Volcano Engine 火山引擎, a separate product
Availability
CN mainland
Max input images
9 the generator widget states 上传最多12个参考素材 across 图/文/音/视频 — the same 12-slot budget Dreamina breaks down as 9 photos, 3 videos, 3 audio

Sources (1)

Directory entry: Jimeng AI 即梦 (Dreamina)
MiniMax H3 (Hailuo 3.0) MiniMax Current 15 s Not researched yet 2K qualified 24 fps Yes qualified T2V · I2V · FL2V · REF2V · REF2VA qualified $0.13 / s qualified ¥0.8 / s qualified Yes Yes qualified
Verified 31 days ago
Evidence & full specs (7 sources)
Max clip
15 s
Max resolution
2K directly selectable on the base create endpoint (resolution enum = 768P | 2K); a SEPARATE regeneration endpoint also upsells 768P→2K at its own price, but 2K is NOT gated behind it
Frame rate
24 fps
Native audio
Yes 32 kHz stereo, jointly generated with the video
Modes
T2V · I2V · FL2V · REF2V · REF2VA one create endpoint; mode is selected by content[].role (first_frame, last_frame, reference_image, reference_video, reference_audio). i2v/fl2v and ref2v are mutually exclusive in one request. Vendor names its two open checkpoints FL2VA and Ref2VA — the trailing 'A' means the output carries audio, so ref2va is the vendor's own framing of reference-to-video+audio
Price (global)
$0.13 / s international platform (platform.minimax.io), 2K output. 768P output is $0.08/second. Mainland platform (platform.minimaxi.com) bills ¥0.80/s (2K) and ¥0.50/s (768P). Input material billed separately: first 5 reference images free then $0.04 each; reference video billed at output rate by its own duration; reference audio free. A separate /video-generation-v2-regeneration ENDPOINT re-emits an existing 768P H3 clip at 2K for $0.05/s (mainland CNY 0.30/s); it reuses the MiniMax-H3 model id, so it is an additional route and a billing row rather than a distinct model — and 2K is NOT gated behind it
Price (mainland)
¥0.8 / s mainland platform platform.minimaxi.com, 2K output; 768P is CNY 0.50/s. The regeneration route is CNY 0.30/s. Recorded as a separate field rather than a converted figure: these are two independently-set price lists, not one price in two currencies
Public API
Yes
Open weights
Yes open-sourced 2026-08-03, after the 2026-07-31 hosted launch; two checkpoints published (MiniMax-H3-Base-FL2VA, MiniMax-H3-Base-Ref2VA)
Min clip
4 s
Licence
MiniMax H3 Community License Agreement users in the USA, EU, UK and South Korea must submit an application form before use
Watermark
optional API: aigc_watermark boolean, optional, DEFAULT false — so API output is unwatermarked unless requested. Documented on the mainland create endpoint and on the international regeneration endpoint, but NOT listed on the international create endpoint. The consumer product (hailuoai.video) is separate and watermarks free-tier downloads
Availability
global two separate self-serve platforms with different billing currencies: platform.minimax.io (international, USD) and platform.minimaxi.com (mainland China, CNY). Open weights carry a licence application requirement for USA/EU/UK/South Korea
Negative prompt
No
Max input images
9 reference mode: ≤9 images, ≤3 video clips, ≤3 audio clips, ≤12 files total; per-asset caps video ≤50 MB, image ≤30 MB, audio ≤15 MB, request body ≤64 MB

Released 2026-07-31

Sources (7)

Directory entry: Hailuo Video 海螺视频
MiniMax-Hailuo-2.3 MiniMax Superseded 10 s qualified 6 / 10 s combination limit 1080P combination limit Not published qualified Not published qualified T2V · I2V $0.28 / clip qualified ¥2 / clip qualified Yes No

▲ Combination limit 1080P is available on 6-second clips ONLY — choosing 10s forfeits 1080P and caps output at 768P

Verified 31 days ago
Evidence & full specs (4 sources)
Max clip
10 s 768P only
Fixed lengths
6 / 10 s 10s is 768P-only; 1080P is 6s-only
Max resolution
1080P 6-second clips only; 10-second clips top out at 768P (the default)
Frame rate
Not published vendor does not publish an output frame rate for the 2.x line on either the t2v or i2v reference page
Native audio
Not published vendor documents no audio output for this line — the request schema has no audio field and the response is video only — but it is nowhere stated as absent, so recorded as unpublished rather than false
Modes
T2V · I2V
Price (global)
$0.28 / clip international platform, 768P 6s. Also 768P 10s $0.56, 1080P 6s $0.49. Mainland: ¥2.00 / ¥4.00 / ¥3.50 per video respectively. Billed PER CLIP, not per second — unlike H3
Price (mainland)
¥2 / clip mainland platform, 768P 6s; 768P 10s CNY 4.00, 1080P 6s CNY 3.50. Recorded as a separate field rather than a converted figure: these are two independently-set price lists, not one price in two currencies
Public API
Yes
Open weights
No
Min clip
6 s
Negative prompt
No

Released 2025-10-28

Sources (4)

Directory entry: Hailuo Video 海螺视频
MiniMax-Hailuo-2.3-Fast MiniMax Superseded Not researched yet 6 / 10 s combination limit 1080P combination limit Not researched yet Not researched yet I2V qualified $0.19 / clip qualified ¥1.35 / clip qualified Yes Not researched yet

▲ Combination limit 1080P is available on 6-second clips ONLY — choosing 10s forfeits 1080P and caps output at 768P

Verified 31 days ago
Evidence & full specs (2 sources)
Fixed lengths
6 / 10 s 10s is 768P-only; 1080P is 6s-only
Max resolution
1080P 6-second clips only
Modes
I2V listed in the i2v model enum only — it is absent from the t2v model enum, so this variant appears to be image-to-video only
Price (global)
$0.19 / clip international platform, 768P 6s. Also 768P 10s $0.32, 1080P 6s $0.33. Mainland: ¥1.35 / ¥2.25 / ¥2.31 per video
Price (mainland)
¥1.35 / clip mainland platform, 768P 6s; 768P 10s CNY 2.25, 1080P 6s CNY 2.31. Recorded as a separate field rather than a converted figure: these are two independently-set price lists, not one price in two currencies
Public API
Yes
Negative prompt
No

Sources (2)

Directory entry: Hailuo Video 海螺视频
MiniMax-Hailuo-02 MiniMax Superseded 10 s 6 / 10 s combination limit 1080P combination limit Not researched yet Not researched yet T2V · I2V $0.1 / clip qualified ¥0.6 / clip qualified Yes No

▲ Combination limit 1080P is available on 6-second clips ONLY; 10-second clips are limited to 512P or 768P

Verified 31 days ago
Evidence & full specs (4 sources)
Max clip
10 s
Fixed lengths
6 / 10 s 10s available at 512P and 768P; 1080P is 6s-only
Max resolution
1080P 6-second clips only; this line uniquely also offers 512P
Modes
T2V · I2V
Price (global)
$0.1 / clip international platform, 512P 6s (cheapest tier). Also 512P 10s $0.15, 768P 6s $0.28, 768P 10s $0.56, 1080P 6s $0.49. Mainland: ¥0.60 / ¥1.00 / ¥2.00 / ¥4.00 / ¥3.50 per video
Price (mainland)
¥0.6 / clip mainland platform, 512P 6s; 512P 10s CNY 1.00, 768P 6s CNY 2.00, 768P 10s CNY 4.00, 1080P 6s CNY 3.50. Recorded as a separate field rather than a converted figure: these are two independently-set price lists, not one price in two currencies
Public API
Yes
Open weights
No
Min clip
6 s
Negative prompt
No

Released 2025-06-18

Sources (4)

Directory entry: Hailuo Video 海螺视频
T2V-01 / I2V-01 / S2V-01 (incl. -Director, -live) MiniMax Superseded 6 s 6 s qualified 720P qualified Not researched yet Not researched yet T2V · I2V · REF2V qualified Not researched yet Not researched yet Yes qualified No
Verified 31 days ago
Evidence & full specs (4 sources)
Max clip
6 s
Fixed lengths
6 s 6 seconds only; 10s not supported on this line
Max resolution
720P default and only value for this line
Modes
T2V · I2V · REF2V split across three model IDs and endpoints: T2V-01/-Director (t2v), I2V-01/-Director/-live (i2v), S2V-01 (subject reference, character face only)
Public API
Yes still present in the model enums on the t2v/i2v/s2v reference pages, but ABSENT from the Pay-as-you-go price list entirely — no published price on either platform
Open weights
No
Negative prompt
No
Max input images
1 S2V-01 only: subject_reference takes a single image (type 'character'); JPG/JPEG/PNG/WebP, <20 MB, shorter side >300px, aspect ratio 2:5 to 5:2

Sources (4)

Directory entry: Hailuo Video 海螺视频
Vidu Q3 Pro Shengshu Current 16 s Not researched yet 1080p qualified 24 fps Yes qualified T2V · I2V · FL2V $0.12 / s qualified Not researched yet Yes Not researched yet
Verified 30 days ago
Evidence & full specs (4 sources)
Max clip
16 s
Max resolution
1080p 540p | 720p | 1080p, default 720p
Frame rate
24 fps
Native audio
Yes audio boolean, Q3 only, defaults TRUE (dialogue plus sound effects). bgm is unavailable on Q3
Modes
T2V · I2V · FL2V
Price (global)
$0.12 / s 1080p PEAK rate (24 credits/s at $0.005/credit); 720p $0.10/s, 540p $0.045/s. Off-peak roughly halves each — Vidu publishes separate peak and off-peak rates, so any single figure needs the qualifier
Public API
Yes
Min clip
1 s free-form integer 1-16, default 5
Negative prompt
No documented parameters are prompt, duration, resolution, aspect_ratio, movement_amplitude, style, bgm, audio, seed
Max input images
1 image-to-video accepts exactly 1 image; first/last-frame mode takes 2

Sources (4)

Directory entry: Vidu 维度AI
Vidu Q3 Turbo Shengshu Current 16 s Not researched yet 1080p 24 fps Yes qualified T2V · I2V · FL2V · REF2V $0.065 / s qualified Not researched yet Yes Not researched yet
Verified 30 days ago
Evidence & full specs (5 sources)
Max clip
16 s
Max resolution
1080p
Frame rate
24 fps
Native audio
Yes audio defaults to TRUE on the Q3 series and emits dialogue plus sound effects; bgm is unavailable on q3. Two qualifiers the earlier scope had wrong: voice_id is documented as "not effective" on the Q3 series, and audio_type is accepted but its SPLITTING is documented as supporting "q2, q1, and 2.0 series models" only — so on Q3 Turbo you get combined audio and cannot isolate speech or effects. Both caveats are stated on image-to-video; the Subjects-Reference to Video body carries audio_type without them.
Modes
T2V · I2V · FL2V · REF2V
Price (global)
$0.065 / s 1080p peak (13 credits/s); 720p $0.055/s, 540p $0.035/s; off-peak roughly halves
Public API
Yes
Min clip
1 s 1-16 on t2v/i2v; the reference-to-video endpoint narrows this model to 3-16
Negative prompt
No
Max input images
7 reference-to-video accepts 1-7; the image-to-video endpoint accepts exactly 1

Sources (5)

Directory entry: Vidu 维度AI
Vidu Q3 Mix Shengshu Current 16 s Not researched yet 1080p qualified 24 fps Yes qualified REF2V qualified Not researched yet Not researched yet Yes Not researched yet
Verified 30 days ago
Evidence & full specs (3 sources)
Max clip
16 s
Max resolution
1080p 720p | 1080p only — unlike Q3 Pro and Turbo, this model offers no 540p
Frame rate
24 fps
Native audio
Yes The model emits synchronised audio — the endpoint's own model list describes viduq3-mix as supporting "simultaneous audio and video output" — but NOTHING PARAMETERISES IT. reference-to-video.md documents two endpoints: Subjects-Reference to Video carries audio and audio_type but its model enum excludes viduq3-mix, while Reference to Video accepts viduq3-mix and carries no audio field at all (only bgm, which q3 does not support). So on this model audio cannot be switched off, split, or voiced through the API.
Modes
REF2V Reference-to-video is the only mode this model accepts. The first/last-frame endpoint (start-end-to-video) does not list viduq3-mix among its accepted models, and pricing carries only reference2video rows for Q3-mix.
Public API
Yes
Min clip
1 s
Max input images
7

Sources (3)

Directory entry: Vidu 维度AI
Vidu Q2 (q2, q2-pro, q2-turbo) Shengshu Superseded 10 s Not researched yet 1080p 24 fps Yes qualified T2V · I2V · FL2V · REF2V · V2V qualified Not researched yet Not researched yet Yes Not researched yet
Verified 30 days ago
Evidence & full specs (5 sources)
Max clip
10 s
Max resolution
1080p
Frame rate
24 fps
Native audio
Yes the audio parameter defaults to FALSE on Q2 and Q1, against TRUE on Q3 — the same field, opposite default
Modes
T2V · I2V · FL2V · REF2V · V2V v2v on viduq2-pro only, which supports video reference, editing and replacement; with a video attached its reference budget drops from 7 images to 4
Public API
Yes
Min clip
1 s 1-10, default 5
Max input images
7 1-7 reference images; 1-4 when a reference video is also supplied

Sources (5)

Directory entry: Vidu 维度AI
Vidu Q1 Shengshu Superseded Not researched yet 5 s qualified 1080p qualified 24 fps Not researched yet T2V · I2V · FL2V · REF2V $0.4 / clip qualified Not researched yet Yes Not researched yet
Verified 30 days ago
Evidence & full specs (3 sources)
Fixed lengths
5 s 5 seconds is the only selectable duration on this model
Max resolution
1080p 1080p is the only option — Q1 cannot generate 540p or 720p
Frame rate
24 fps
Modes
T2V · I2V · FL2V · REF2V
Price (global)
$0.4 / clip 80 credits per 5-second 1080p clip; off-peak is half. Billed per clip, not per second, unlike Q2 and Q3
Public API
Yes

Sources (3)

Directory entry: Vidu 维度AI
Vidu 2.0 Shengshu Superseded Not researched yet 4 / 8 s 1080p qualified 32 fps qualified Not researched yet I2V · FL2V · REF2V qualified Not researched yet Not researched yet Yes Not researched yet
Verified 30 days ago
Evidence & full specs (3 sources)
Fixed lengths
4 / 8 s
Max resolution
1080p 1080p is reachable at 4s only, where image-to-video offers 360p (default), 720p and 1080p; the 8s option is 720p and nothing else. Pricing bills the 4s 1080p clip as a product ("Vidu 2.0 img2video 4S 1080P", 100 credits / $0.50), and the model map lists 360p, 720p, 1080p for the family.
Frame rate
32 fps the only Vidu model not at 24 fps
Modes
I2V · FL2V · REF2V First/last-frame is reachable: the start-end-to-video endpoint accepts vidu2.0, and pricing bills "Vidu 2.0 start-end2video" at four tiers (4s 360p/720p/1080p and 8s 720p). Text-to-video is genuinely absent — that endpoint accepts only viduq3-turbo, viduq3-pro, viduq2 and viduq1. The model map additionally marks Effects/Templates for this model, which is a preset gallery rather than a generation mode and has no code in this vocabulary.
Public API
Yes

Sources (3)

Directory entry: Vidu 维度AI
Wan 2.7 Alibaba Current 15 s qualified Not researched yet 1080p qualified 30 fps Yes qualified T2V · I2V · FL2V · REF2V · V2V · S2V qualified $0.15 / s qualified Not researched yet Yes No qualified
Verified 31 days ago
Evidence & full specs (5 sources)
Max clip
15 s t2v and i2v 2-15s; the r2v and videoedit endpoints cap at 10s
Max resolution
1080p 720P | 1080P, default 1080P on i2v. No 480P tier on 2.7, unlike Wan 2.5 and Wan 3.0
Frame rate
30 fps
Native audio
Yes generates dubbing when no audio is supplied; an audio_url may instead be passed as a lip-sync driver
Modes
T2V · I2V · FL2V · REF2V · V2V · S2V split across four model IDs — wan2.7-t2v, -i2v, -r2v, -videoedit. Request shapes differ from the 2.6 'legacy protocol' and are not interchangeable
Price (global)
$0.15 / s 1080P; 720P $0.10/s. Same rate across t2v, i2v, r2v and videoedit. 50s free quota, Singapore only. r2v and videoedit bill INPUT video duration in addition to output
Public API
Yes
Open weights
No the Wan-Video GitHub organisation publishes model repositories for 2.1 and 2.2 only — there is no 2.5, 2.6, 2.7 or 3.0 weights repository. 'Wan is open source' is true of the 2022-2.2 line and false of everything currently sold
Min clip
2 s
Negative prompt
Yes negative_prompt, maximum 500 characters
Max input images
2 first_frame + last_frame on wan2.7-i2v

Sources (5)

Directory entry: Wan 通义万相
Wan 3.0 Alibaba Preview 30 s Not researched yet 1080p qualified 30 fps Yes T2V · I2V · FL2V · REF2V qualified $0.2 / s qualified Not researched yet No qualified Not researched yet
Verified 31 days ago
Evidence & full specs (2 sources)
Max clip
30 s
Max resolution
1080p 480P | 720P | 1080P
Frame rate
30 fps
Native audio
Yes
Modes
T2V · I2V · FL2V · REF2V documented as all-in-one reference — images, videos and audio, plus docx/ppt/pdf files and web links as input
Price (global)
$0.2 / s 1080P; 720P $0.10/s, 480P $0.05/s. Billed on input plus output video duration combined, not output alone
Public API
No flagged Invitational Preview on the price list — the endpoint exists but access is gated, not self-serve. Do not present Wan 3.0 as generally available
Min clip
2 s
Availability
Singapore and China (Beijing) only narrower than the rest of the Wan line, which the same price page also lists in Frankfurt, Virginia and Tokyo. The 30s free quota applies in Singapore only

Sources (2)

Directory entry: Wan 通义万相
Wan 2.6 Alibaba Superseded 15 s qualified Not researched yet 1080p 30 fps Yes qualified T2V · I2V · REF2V $0.15 / s qualified Not researched yet Yes Not researched yet
Verified 31 days ago
Evidence & full specs (3 sources)
Max clip
15 s t2v/i2v 2-15s; r2v 2-10s. The US-scoped fork wan2.6-t2v-us accepts only 5 and 10 — same version number, different spec, purely by region
Max resolution
1080p
Frame rate
30 fps
Native audio
Yes wan2.6 and wan2.5 generate audio by default; on the -flash variants audio is a priced toggle and audio=false halves the rate
Modes
T2V · I2V · REF2V
Price (global)
$0.15 / s 1080P; 720P $0.10/s. The -flash variants are far cheaper: wan2.6-i2v-flash is $0.075/s at 1080p with audio, $0.0375/s silent
Public API
Yes
Min clip
2 s
Negative prompt
Yes

Sources (3)

Directory entry: Wan 通义万相
Wan 2.5 Alibaba Superseded 10 s 5 / 10 s qualified 1080p qualified 30 fps Yes T2V · I2V $0.15 / s qualified Not researched yet Not researched yet Not researched yet
Verified 31 days ago
Evidence & full specs (3 sources)
Max clip
10 s
Fixed lengths
5 / 10 s a fixed set — 2.5 cannot do free-form durations, unlike 2.6 and 2.7
Max resolution
1080p 480P | 720P | 1080P, default 1920x1080
Frame rate
30 fps
Native audio
Yes
Modes
T2V · I2V
Price (global)
$0.15 / s 1080P; 720P $0.10/s, 480P $0.05/s
Negative prompt
Yes

Sources (3)

Directory entry: Wan 通义万相
Wan 2.2 (open weights, and the hosted -plus / -flash endpoints) Alibaba Superseded Not researched yet 5 s qualified 1080p qualified 30 fps qualified No qualified T2V · I2V · FL2V · S2V qualified $0.1 / s qualified Not researched yet Not researched yet Yes qualified
Verified 31 days ago
Evidence & full specs (3 sources)
Fixed lengths
5 s the hosted t2v/i2v/kf2v endpoints are fixed at 5 seconds. The separate wan2.2-animate-move / -animate-mix character-animation endpoints run 2-30s at 15/25 fps
Max resolution
1080p hosted wan2.2-t2v-plus is 480P and 1080P (no 720P); -i2v-flash and -kf2v-flash add 720P. The OPEN checkpoint is a different configuration — its own repo documents 480P and 720P, with TI2V-5B at 720P
Frame rate
30 fps hosted API. The open TI2V-5B checkpoint is documented at 24 fps — the served model and the downloadable model are NOT the same configuration
Native audio
No 'No audio' on every hosted wan2.2 video endpoint; the open S2V-14B checkpoint is speech-DRIVEN (audio in), not audio-generating
Modes
T2V · I2V · FL2V · S2V s2v is open-weights only (S2V-14B); the hosted endpoints are t2v/i2v/kf2v plus the animate pair
Price (global)
$0.1 / s hosted wan2.2-t2v-plus at 1080P; 480P $0.02/s. wan2.2-i2v-flash is $0.036/s at 720P. Zero if you self-host the open weights
Open weights
Yes five checkpoints: T2V-A14B, I2V-A14B, TI2V-5B, S2V-14B, Animate-14B. This is the NEWEST open Wan release — 2.5 onward is API-only
Licence
Apache-2.0

Sources (3)

Directory entry: Wan 通义万相
HunyuanVideo 1.5 Tencent Current 10 s qualified Not researched yet 720p qualified 24 fps Not published qualified T2V · I2V Not researched yet Not researched yet Not published qualified Yes
Verified 31 days ago
Evidence & full specs (4 sources)
Max clip
10 s the README benchmarks 10-second 720p synthesis; the reference CLI's --video_length defaults to 121 frames (5s at 24fps) and takes a free-form frame count, so 10s is the longest length the vendor documents rather than a selectable preset
Max resolution
720p native checkpoints are 480p and 720p only; the 1080p figure people quote is a SEPARATE few-step super-resolution network applied afterwards, not a generation resolution
Frame rate
24 fps
Native audio
Not published no audio component is documented in the repo or the model card; the card lists video output only
Modes
T2V · I2V
Public API
Not published Tencent publishes no self-serve API tied to this checkpoint. Tencent Cloud sells a separate hosted 混元生视频 product, priced at 1.5 credits per 720p 5s call, but never states which model version backs it — so it cannot be attributed to HunyuanVideo 1.5
Open weights
Yes
Licence
Tencent Hunyuan Community License the licence text opens by excluding the EU, UK and South Korea from its Territory — 'open weights' and 'usable in the EU' are different claims here
Availability
global excl. EU/UK/KR licence-gated, not service-gated: the weights download anywhere, but the grant does not extend to those three territories
Negative prompt
Yes in the reference inference script (--negative_prompt); Tencent ships no first-party hosted API for this checkpoint

Released 2025-11-20

Sources (4)

Directory entry: Hunyuan Video 混元视频
HunyuanVideo (13B) Tencent Superseded 5 s qualified Not researched yet 720p qualified Not researched yet Not researched yet T2V qualified Not researched yet Not researched yet Not researched yet Yes
Verified 31 days ago
Evidence & full specs (2 sources)
Max clip
5 s 129 frames
Max resolution
720p 720x1280 / 1280x720 / 1104x832 / 832x1104 / 960x960; a 540p tier also exists
Modes
T2V the base 13B model is text-to-video; image-to-video ships as a separate checkpoint, HunyuanVideo-I2V. The 1.5 line lives in a SEPARATE repository, not a branch of this one
Open weights
Yes
Licence
Tencent Hunyuan Community License same EU/UK/South Korea territorial exclusion as 1.5

Released 2024-12-03

Sources (2)

Directory entry: Hunyuan Video 混元视频
CogVideoX-3 (hosted) Zhipu AI Current 10 s 5 / 10 s 4K qualified 60 fps qualified Yes qualified T2V · I2V · FL2V Not researched yet ¥1 / clip qualified Yes Not researched yet
Verified 31 days ago
Evidence & full specs (3 sources)
Max clip
10 s
Fixed lengths
5 / 10 s
Max resolution
4K the size enum is a fixed list — 1280x720, 720x1280, 1024x1024, 1920x1080, 1080x1920, 2048x1080, 3840x2160. 4K means that last member, not arbitrary resolution
Frame rate
60 fps enum 30 or 60, default 30
Native audio
Yes with_audio defaults to false, and the docs describe it as 生成 AI 音效 — sound effects, not dialogue or speech
Modes
T2V · I2V · FL2V
Price (mainland)
¥1 / clip 1 元/次 on the CN platform bigmodel.cn — flat per call regardless of duration or resolution. The international platform z.ai prices the SAME model id at $0.2/video
Public API
Yes
Min clip
5 s
Watermark
optional watermark_enabled defaults TRUE and applies both a visible AI mark and an implicit digital watermark, described as 符合政策要求. Turning it off requires a signed disclaimer through 个人中心-安全管理-去水印管理 — not freely switchable
Availability
global, on two separately-billed platforms bigmodel.cn for CN in CNY and docs.z.ai internationally in USD — same model id, different endpoint host and price
Negative prompt
No read against the full request schema — no negative-prompt parameter exists on any video model on the platform
Max input images
2 first/last-frame mode: first image is the first frame, second is the last; each at most 5MB

Sources (3)

Directory entry: Zhipu Qingying 智谱清影 (CogVideoX)
CogVideoX-2 (hosted) Zhipu AI Superseded Not researched yet Not researched yet 4K qualified 60 fps qualified Not researched yet T2V · I2V qualified Not researched yet ¥0.5 / clip Yes Not researched yet
Verified 30 days ago
Evidence & full specs (2 sources)
Max resolution
4K size enum 720x480, 1024x1024, 1280x960, 960x1280, 1920x1080, 1080x1920, 2048x1080, 3840x2160
Frame rate
60 fps enum 30 or 60
Modes
T2V · I2V single input image only — no first/last-frame mode, unlike CogVideoX-3
Price (mainland)
¥0.5 / clip
Public API
Yes
Max input images
1

Sources (2)

Directory entry: Zhipu Qingying 智谱清影 (CogVideoX)
CogVideoX1.5-5B (open checkpoint) Zhipu AI Superseded 10 s qualified Not researched yet 1360x768 qualified 16 fps Not researched yet T2V · I2V Not researched yet Not researched yet Not researched yet Yes qualified
Verified 31 days ago
Evidence & full specs (1 sources)
Max clip
10 s 5-10s; frame counts must satisfy 16N+1 with N at most 10
Max resolution
1360x768 the I2V variant accepts variable resolution with a minimum side of 768 and a maximum of 1360
Frame rate
16 fps
Modes
T2V · I2V
Open weights
Yes this is the NEWEST open checkpoint — there are no open weights for CogVideoX-2 or CogVideoX-3, so 'CogVideoX is open source' is true of the 2024 line and false of everything currently sold. The repo also MOVED, from THUDM/CogVideo to zai-org/CogVideo
Licence
CogVideoX LICENSE the 5B-class T2V and I2V transformer weights are under the bespoke CogVideoX LICENSE; only the smaller CogVideoX-2B is Apache-2.0

Released 2024-11-08

Sources (1)

Directory entry: Zhipu Qingying 智谱清影 (CogVideoX)
PixVerse C1 AIsphere Current 15 s Not researched yet 1080p qualified Not published qualified Yes qualified T2V · I2V · FL2V · REF2V qualified $0.01 / credit qualified Not researched yet Yes Not researched yet
Verified 31 days ago
Evidence & full specs (6 sources)
Max clip
15 s
Max resolution
1080p quality enum 360p, 540p, 720p, 1080p
Frame rate
Not published PixVerse exposes no fps parameter and publishes no output frame rate anywhere in the platform docs — duration, quality and aspect_ratio are the only output controls
Native audio
Yes opt-in via generate_audio_switch, default FALSE, and it raises the credit rate (1080p goes 19 to 24 credits/s). 'Has native audio' is true; 'generates audio by default' is not
Modes
T2V · I2V · FL2V · REF2V the vendor names them Text-to-video, Image-to-video, Transition and Fusion
Price (global)
$0.01 / credit $10 = 1,000 credits at the smallest pack. C1 bills per second: 360p 6 cr/s, 540p 8, 720p 10, 1080p 19 silent; 8/10/13/24 with audio. A 5s 1080p clip with audio is 120 credits, about $1.20
Public API
Yes
Min clip
1 s
Watermark
optional the API exposes a water_mark boolean. The consumer free-tier watermark is widely reported by third parties but appears on no PixVerse-owned page, so only the API parameter is recorded
Availability
global the vendor states the platform serves creators across 175 countries; no per-region model gating is documented
Negative prompt
Yes

Released 2026-04-07

Sources (6)

Directory entry: PixVerse
PixVerse V6 AIsphere Current 15 s Not researched yet 1080p Not researched yet Yes qualified T2V · I2V · FL2V · V2V · REF2V qualified $0.01 / credit qualified Not researched yet Yes Not researched yet
Verified 31 days ago
Evidence & full specs (3 sources)
Max clip
15 s
Max resolution
1080p
Native audio
Yes switchable across all five modes; generate_audio_switch defaults false
Modes
T2V · I2V · FL2V · V2V · REF2V vendor names: Text-to-Video, Image-to-Video, Transition, Video Extension, Reference-to-Video. Multi-clip is available on t2v and i2v only
Price (global)
$0.01 / credit V6 is 5/7/9/18 credits per second at 360p/540p/720p/1080p silent, 7/9/12/23 with audio; Fusion with VIDEO references costs double
Public API
Yes
Min clip
1 s

Sources (3)

Directory entry: PixVerse
PixVerse V5.6 AIsphere Superseded 10 s qualified 5 / 8 / 10 s qualified 1080p Not researched yet Yes qualified Not researched yet $0.01 / credit qualified Not researched yet Not researched yet Not researched yet
Verified 31 days ago
Evidence & full specs (2 sources)
Max clip
10 s 720p and below; 1080p caps at 8s
Fixed lengths
5 / 8 / 10 s 1080p cannot use 10s on v5.5/v5.6
Max resolution
1080p
Native audio
Yes generate_audio_switch exists on v5.5, v5.6, v6 and c1 only — NOT on v5 and earlier
Price (global)
$0.01 / credit V5.6 bills per CLIP, not per second — 5s costs 35/35/45/75 credits at 360p/540p/720p/1080p silent; 10s costs 77/77/99 with no 1080p. The billing MODEL changes at v6, which is what a flat duration column destroys

Sources (2)

Directory entry: PixVerse
Advertisement
After tool · AD-W1Responsive · Post-tool

How to use this comparison

Find the version, not the brand

Each row is one model version. Kling 3.0 and Kling 2.6 are different products with different limits, and so are Hailuo H3 and Hailuo 2.3. Work out which version you can actually call before you read a number. Superseded versions are hidden by default — tick the switch above the table to bring them in.

Narrow it with length and resolution

Max clip is the longest SINGLE generation, not a stitched sequence. Fixed lengths only carries a value where the vendor offers a fixed set rather than a free-form range. A star on a value means it is qualified — on several models the top resolution exists only at the shortest duration — and the qualification is in that row's "Evidence and full specs". A triangle means two figures on the row are coupled, and that rule stays on screen under the row rather than being folded away.

Check the billing unit before the number

Some models bill per second and some bill per finished clip, and those two figures are not comparable. The global and mainland platforms are separately priced lists, so they get separate columns and no conversion — turning $0.13/s into yuan to compare it with ¥0.80/s produces a number no vendor ever published.

Open the evidence and check us

Every row shows the date it was verified, and its "Evidence and full specs" holds the complete sheet for that version: every value in full, every qualification, and a link to the vendor page each figure came from — including the columns a narrow screen leaves out of the grid. Do not take this table on trust, open the link. If something does not match, that is our bug and we would like to hear about it.

Advertisement
After how-to · AD-W2Responsive

About the Chinese AI video model comparison

Why the rows are versions, not vendors

Flattening a vendor into one row is the commonest error in a comparison table of this kind, and the hardest to notice. It reads neatly, and every individual number in it has a source. The failure is in a question the table never answers: who exactly is this row about?

Kling 3.0 and Kling 2.6 are both on sale. The first takes any integer from 3 to 15 seconds; the second's API enum is exactly 5 or 10. Hailuo H3 and Hailuo 2.3 are both on sale. The first bills per second; the second bills per finished clip. Write a single max duration for "Kling", or a single per-second price for "Hailuo", and whichever number you pick is wrong for half the people reading it.

So the unit here is vendor plus version. The service that renders this table has no code path that merges fields across versions, and that is held in place by mutation testing rather than by good intentions: put the merging back and the suite has to go red.

An error somebody can catch beats a plausible figure nobody can falsify.

Why price is two columns and never converted

The same model routinely carries two entirely separate prices. MiniMax lists H3 at $0.13 per second on platform.minimax.io and ¥0.80 per second on platform.minimaxi.com. No official exchange rate connects those two figures — they are two independently-set price lists that happen to describe the same weights.

Collapsing them into one column would mean choosing a rate, and that rate would be ours rather than the vendor's. The page would then display a number that appears in no vendor document anywhere. Hence two columns, and no conversion, ever.

The billing unit deserves the same care. "$0.13 per second" and "$0.28 per clip" cannot be ranked against each other in one column, and dropping the word "clip" is a quiet way of misstating the price of half the models on this page.

Every model here is a version, not a vendor

01

A Chinese lab selling several model versions at once is the norm, not the exception — which is exactly why this table has a row per version rather than per vendor.

02

Max duration and max resolution often cannot be taken together: on several models the top resolution exists only at the shortest clip length.

03

The billing unit matters more than the number. Comparing a per-second price against a per-clip price in one column produces a wrong answer every time.

04

The global and mainland platforms are separately priced. There is no official exchange rate linking one vendor's two price lists.

05

Vendor marketing pages and API references contradict each other regularly. Where they do, this site records the API reference and files the marketing claim as a qualification.

06

"Not published" is not "no". Most Chinese video models have never published a frame rate — that is a finding, not a zero.

07

Open weights does not mean usable everywhere: some community licences carve out whole jurisdictions and require a separate written grant.

08

A model filed under "legacy" is usually the vendor's own shelving, not a sunset date. Whether it still answers is a question for the API reference.

09

Mainland Chinese generative services must label AIGC output. That is regulation, not a product choice, and should not be read as one.

10

Every figure on this page carries its source and the date it was read, and an automated gate watches those dates — a stale table fails a check instead of sitting here for a year.

Frequently asked questions

  • No. The page is a comparison table rendered on our own servers from our AI directory database. It generates no video, calls no model, and makes no request to any third party.
  • Because it sells several versions at once and their limits are incompatible. Kling 3.0 accepts any integer from 3 to 15 seconds while Kling 2.6's API enum is exactly 5 or 10; MiniMax H3 bills per second while its own 2.3 line bills per finished clip. Merged into one row, whatever number you write is false for one of them.
  • "Not published" means we read the vendor's own documentation and the vendor does not state it — a sourced finding, and you can open it to see which page was checked. A dash means nobody has researched that field yet. Neither one means "no" and neither one means zero.
  • Each one comes from the vendor's own API reference, price list or model card, and carries the link and the date it was read. Resellers are excluded on principle — a wrapper's published specs are its own product's, not the model maker's. Open the source and check us.
  • It follows the AI directory's own re-research cycle. Each row shows the date of its OLDEST verified claim, and anything past 120 days is flagged stale by an automated check. The date is computed from the data, not typed by hand, so it cannot drift away from what it describes.
  • Because they are two separately-set price lists, not one price in two currencies. MiniMax sells H3 at $0.13/s on its global platform and ¥0.80/s on its mainland platform; converting one into the other would put a number on the page that neither vendor page contains.
  • Because this table is for comparing the domestic video models Chinese-speaking users actually choose between, in one context. Sora, Runway, Veo and the rest have their own entries and their own spec sheets in our AI directory — they are simply not what this page is for.
  • Usually, yes. "Superseded" here follows the vendor's own shelving: some vendors file a model under a legacy heading while its API reference carries no deprecation notice at all and the endpoint answers normally. We do not invent a sunset date the vendor has not published.
  • No. The page takes no input and sends nothing anywhere. The table is rendered server-side and delivered as finished HTML.
  • Tell us. Every row links its sources precisely so you can falsify us. An error somebody can catch is far better than a plausible figure nobody can check — that second kind is what this rebuild existed to remove.

Related News

You may be interested in these recent stories from our newsroom.

View all news →

Method & sources

How it computes

Compares Chinese AI video models on published specifications and list pricing, from the same per-row-sourced card. This page is a PROJECTION of that data, which is why its freshness is gated rather than assumed.

What this tool implements

  • Prices are per-row provenance: each model carries its own vendor pricing URL and the date it was last verified, rather than one blanket citation for the table.
  • Whole table last verified 2 September 2026 against the vendors' own published pricing pages.
  • ⚠️ Published list prices only. Enterprise agreements, committed-use discounts and promotional rates are not visible to us and are frequently lower.
  • Cached-input and batch rates differ substantially from standard rates at several vendors, and are carried separately rather than averaged into one number.
  • ⚠️ LLM pricing moves faster than almost any other input on this site. A rate card is a snapshot with a date on it, and the date is shown so a stale figure is visible rather than silent.

What can make this go out of date

  • Six vendor pricing pages, checked by hand rather than scraped. Each row records its own verified_at date, so a row that has not been rechecked is identifiable instead of quietly inheriting the table's freshness.
Advertisement
Pre-footer · AD-W3 728 × 90