Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model
About
Recent strides in video generation have paved the way for unified audio-visual generation. In this work, we present Seedance 1.5 pro, a foundational model engineered specifically for native, joint audio-video generation. Leveraging a dual-branch Diffusion Transformer architecture, the model integrates a cross-modal joint module with a specialized multi-stage data pipeline, achieving exceptional audio-visual synchronization and superior generation quality. To ensure practical utility, we implement meticulous post-training optimizations, including Supervised Fine-Tuning (SFT) on high-quality datasets and Reinforcement Learning from Human Feedback (RLHF) with multi-dimensional reward models. Furthermore, we introduce an acceleration framework that boosts inference speed by over 10X. Seedance 1.5 pro distinguishes itself through precise multilingual and dialect lip-syncing, dynamic cinematic camera control, and enhanced narrative coherence, positioning it as a robust engine for professional-grade content creation. Seedance 1.5 pro is now accessible on Volcano Engine at https://console.volcengine.com/ark/region:ark+cn-beijing/experience/vision?type=GenVideo.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Robotic Video Generation | R-Bench | Average Score58.4 | 44 | |
| Context Learning in Video Generation | CLVG-Bench | P.R. Score35.47 | 8 | |
| Image-to-Video Generation | SeedVideoBench Audio dimension 2.0 | Usability Rate (Quality & Expressiveness)93.99 | 6 | |
| Image-to-Video | SeedVideoBench 2.0 (overall) | Motion Quality2.53 | 6 | |
| Image-to-Video Generation | SeedVideoBench Video dimension 2.0 | Usability Rate (Motion Quality)51.44 | 6 | |
| Text-to-Video Generation | SeedVideoBench 2.0 | Motion Fidelity2.39 | 6 |