ByteDance’s Doubao team this week released Seed Audio 1.0, an audio generation model that has gone live in the Volcano Ark experience center for creator testing. The company is also inviting enterprise users to test its API. The model is designed to treat vocals, background music, sound effects, and ambient audio as parts of the same sound scene, producing a complete audio track of up to nearly two minutes in a single run and supporting more than 20 languages.
Doubao said individual users can access the model directly through the Volcano Ark experience center, while international developers can connect through BytePlus. Seed Audio 1.0 has also been integrated into the Doubao app itself.
Built for full-scene audio generation rather than standard TTS
The report says that traditional audio production usually breaks the process into separate steps. Voice-over, music, and effects are often created independently, then aligned on a timeline and manually assembled in editing software. Seed Audio 1.0 is aimed at removing that handoff process by generating the entire sound scene in one inference pass.
That puts it in a different category from conventional text-to-speech, or TTS, which focuses on turning text into spoken output. Seed Audio 1.0 instead maps speech, background music, effects, and ambient sound into one unified acoustic representation. Rather than generating a speaker’s voice and the surrounding sound bed separately, it builds the scene as one object.
Zero-shot reference support and multi-character prompting
According to the article, Seed Audio 1.0 supports zero-shot multimodal reference. Users do not need to retrain the model. By providing either a text sample or an audio sample, they can have the system imitate the timbre and style of the referenced voice.
A single prompt can set up multi-character dialogue, background music, and Foley-style effects at the same time. Even when the generated clip approaches two minutes in length, the voices of multiple characters remain stable instead of drifting during playback.
The model’s timeline control reaches 100-millisecond precision, allowing users to specify when dialogue begins and when effects enter a scene with tight timing. It supports more than 20 languages, including Chinese, English, and Japanese. The report also says timbre and style can be adjusted separately, so users do not need to swap to a different voice just to change the delivery style.
Targeted at dubbing, scoring, Foley, and mixing workflows
The capability set is positioned for film-style audio creation, covering dubbing, music scoring, Foley, and mixing. The report also extends those use cases to audiobooks, radio dramas, podcasts, short-form video, games, and interactive media. In practical terms, it pulls together work that has often been split across separate post-production roles and tools.
One representation space instead of separate pipelines
The article argues that older speech-generation workflows usually split tasks into separate lines: one for speech synthesis, one for music, and one for sound effects. Each model is trained and run independently, and the user is left to align everything afterward on the same timeline.
Seed Audio 1.0 flips that logic. It places all audio elements into the same representation space first and then generates them together. The advantage described in the report is that character voices, music, and effects are synchronized from the start rather than matched later through post-adjustment.
Comparison with ElevenLabs
BlockTempo compares the product with ElevenLabs Studio. Under that setup, speech, music, and sound effects are handled through three separate APIs, and users need to stitch and align the results themselves on a timeline. Doubao is trying to replace those steps with a single inference process.
On pricing, the article says ElevenLabs Pro costs $99 a month, while Doubao uses Volcano Engine’s token-based pricing model, which charges based on usage. The report says that in Chinese-language scenarios, the cost is clearly lower.
The article closes by framing voice-over, music, and sound effects as tasks that have typically required three roles, three software stacks, and three export steps. Seed Audio 1.0 is an attempt to compress those procedures into one model run. Whether it can truly replace professional specialization in areas such as dubbing, audiobooks, and podcasts will depend on how creators judge it after using the product.

