Generate high-quality video with ByteDance Doubao Seedance 2.5 and 2.0 Standard / Fast / Mini. Supports text-to-video, first frame / first & last frame and multimodal reference-to-video; 2.5 goes up to 30 s and the API accepts up to 30 images + 10 videos + 10 audio references, with optional synchronized audio.
SeeDance 2.5 supports 4–30 s; Smart mode usually picks ~10 s, roughly doubling the 5 s cost.
Automatically generate synchronized narration, sound effects and background music (on by default).
Additionally return the last frame image of the video in the result.
Add an 'AI-generated' watermark to the output video.
Gacha mode: submit multiple tasks with the same settings
Prices are estimated defaults; failed tasks are not charged. Final billing is by actual token usage.
No videos generated yet
SeeDance is ByteDance Doubao's video generation model family: 2.5 goes up to 30 s and 1080p, the API takes up to 30 reference images + 10 videos + 10 audio clips, and audio can be the sole reference; 2.0 comes in Standard / Fast / Mini. One endpoint, one content schema — roles distinguish first/last frame, reference images, reference video and reference audio.
Turn a prompt into a clip (4–30 s on 2.5, 4–15 s on 2.0), optionally with synchronized narration, sound effects and background music.
Provide one image as the first frame to control the opening and keep character and scene continuity.
Provide both first and last frames to precisely control the start and end; the last frame can relay into the next clip.
Freely combine multiple reference images + reference videos + reference audio to keep subject and style consistent; ingest assets first and reference them by asset:// ID.
ByteDance Doubao-grade quality with native audio generation, priced to match the official rate with ample concurrency.
SeeDance 2.5 plus 2.0 Standard, Fast and Mini cover every video capability through a single endpoint.
Seedance 2.5: 4–30 s, up to 1080p; the API accepts up to 30 reference images + 10 videos + 10 audio clips, audio may be the sole reference; first-frame tasks use the adaptive ratio.
Standard variant, better stability, supports first frame / first & last frame and multimodal reference, up to 1080p.
Fast variant, quicker generation, up to 720p (no 1080p).
Mini variant, faster and about half the price of Standard; 480p / 720p only.
Prompt-only generation (4–30 s on 2.5, 4–15 s on 2.0); duration can be -1 for the model's smart pick.
First frame ± last frame, or multiple reference images + reference videos + reference audio, preserving subject and style; this page caps inputs for demo purposes, the API accepts more.
16:9 / 4:3 / 1:1 / 3:4 / 9:16 / 21:9 / adaptive, 480p–1080p.
Use a dedicated APIYI SeeDance2-group token, configured separately in Settings.
Questions? Contact hi@apiyi.com for more help.
Generate ByteDance Doubao-grade video across four capabilities, up to 30 s on 2.5, with native synchronized audio and direct MP4 output.