ByteDance launches Doubao Seed Audio 1.0 to generate full audio scenes in a single pass

ByteDance launches Doubao Seed Audio 1.0 to generate full audio scenes in a single pass

N
News Editor
2026-07-22 07:07:25
ByteDance’s Doubao team this week introduced Seed Audio 1.0, a new audio generation model that is now available for creator testing through the Volcano Ark experience center, with enterprise API access also opened for invitation-based trials. The model treats vocals, background music, sound effects, and ambient noise as parts of one unified audio scene rather than separate tracks, allowing it to generate a complete soundtrack of up to nearly two minutes in a single inference run. According to the report, Seed Audio 1.0 supports more than 20 languages, including Chinese, English, and Japanese. It also supports zero-shot multimodal reference, meaning users can provide text or audio samples and have the system imitate the referenced voice timbre and style without retraining. A single prompt can arrange multi-character dialogue, music, and Foley-style effects, while timing can be controlled at 100-millisecond precision. The model has been integrated into the Doubao app. Individual users can try it through Volcano Ark, while international developers can access it via BytePlus. The report contrasts Doubao’s one-step generation approach with ElevenLabs Studio, which separates speech, music, and sound effects into three APIs. It also notes that ElevenLabs Pro costs $99 per month, while Doubao uses Volcano Engine’s token-based pricing model.
ByteDanceDoubaoSeed Audio 1.0AI audio generationVolcano ArkBytePlusTTS

ByteDance’s Doubao team this week released Seed Audio 1.0, an audio generation model that has gone live in the Volcano Ark experience center for creator testing. The company is also inviting enterprise users to test its API. The model is designed to treat vocals, background music, sound effects, and ambient audio as parts of the same sound scene, producing a complete audio track of up to nearly two minutes in a single run and supporting more than 20 languages.

Doubao said individual users can access the model directly through the Volcano Ark experience center, while international developers can connect through BytePlus. Seed Audio 1.0 has also been integrated into the Doubao app itself.

Built for full-scene audio generation rather than standard TTS

The report says that traditional audio production usually breaks the process into separate steps. Voice-over, music, and effects are often created independently, then aligned on a timeline and manually assembled in editing software. Seed Audio 1.0 is aimed at removing that handoff process by generating the entire sound scene in one inference pass.

That puts it in a different category from conventional text-to-speech, or TTS, which focuses on turning text into spoken output. Seed Audio 1.0 instead maps speech, background music, effects, and ambient sound into one unified acoustic representation. Rather than generating a speaker’s voice and the surrounding sound bed separately, it builds the scene as one object.

Zero-shot reference support and multi-character prompting

According to the article, Seed Audio 1.0 supports zero-shot multimodal reference. Users do not need to retrain the model. By providing either a text sample or an audio sample, they can have the system imitate the timbre and style of the referenced voice.

A single prompt can set up multi-character dialogue, background music, and Foley-style effects at the same time. Even when the generated clip approaches two minutes in length, the voices of multiple characters remain stable instead of drifting during playback.

The model’s timeline control reaches 100-millisecond precision, allowing users to specify when dialogue begins and when effects enter a scene with tight timing. It supports more than 20 languages, including Chinese, English, and Japanese. The report also says timbre and style can be adjusted separately, so users do not need to swap to a different voice just to change the delivery style.

Targeted at dubbing, scoring, Foley, and mixing workflows

The capability set is positioned for film-style audio creation, covering dubbing, music scoring, Foley, and mixing. The report also extends those use cases to audiobooks, radio dramas, podcasts, short-form video, games, and interactive media. In practical terms, it pulls together work that has often been split across separate post-production roles and tools.

One representation space instead of separate pipelines

The article argues that older speech-generation workflows usually split tasks into separate lines: one for speech synthesis, one for music, and one for sound effects. Each model is trained and run independently, and the user is left to align everything afterward on the same timeline.

Seed Audio 1.0 flips that logic. It places all audio elements into the same representation space first and then generates them together. The advantage described in the report is that character voices, music, and effects are synchronized from the start rather than matched later through post-adjustment.

Comparison with ElevenLabs

BlockTempo compares the product with ElevenLabs Studio. Under that setup, speech, music, and sound effects are handled through three separate APIs, and users need to stitch and align the results themselves on a timeline. Doubao is trying to replace those steps with a single inference process.

On pricing, the article says ElevenLabs Pro costs $99 a month, while Doubao uses Volcano Engine’s token-based pricing model, which charges based on usage. The report says that in Chinese-language scenarios, the cost is clearly lower.

The article closes by framing voice-over, music, and sound effects as tasks that have typically required three roles, three software stacks, and three export steps. Seed Audio 1.0 is an attempt to compress those procedures into one model run. Whether it can truly replace professional specialization in areas such as dubbing, audiobooks, and podcasts will depend on how creators judge it after using the product.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
600

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.