ByteDance launches Seedance 1.5 Pro with joint audio-video generation
The model generates video and audio through a single diffusion transformer rather than a separate pass, aiming for accurate lip-sync across eight languages.
- Models & capabilities
- Colour
ByteDance’s Seed team released Seedance 1.5 Pro, an update to its Seedance video model that generates video and synchronised audio together rather than as separate passes. The model uses a dual-branch diffusion transformer, one branch for video frames and one for audio waveforms, which ByteDance said gives close temporal and semantic alignment between the two and phoneme-level lip-sync accuracy across multiple languages and dialects. It also added autonomous cinematography features, such as dolly zooms and continuous tracking shots, generated without explicit camera-movement prompts. ByteDance acknowledged that motion stability in complex scenes still had room for improvement.
The model reached individual users through Doubao and Dreamina and enterprise users through the Volcano Engine API and BytePlus ModelArk, on a credit-based pricing system; a five-second 720p clip with audio costs roughly $0.65. It is positioned as ByteDance’s answer to Google’s Veo 3, OpenAI’s Sora 2 and Runway’s Gen-4 in the fast-moving contest over synchronised audio-video generation.