TTS

ByteDance
2026-07-22 07:07:25

ByteDance launches Doubao Seed Audio 1.0 to generate full audio scenes in a single pass

ByteDance’s Doubao team this week introduced Seed Audio 1.0, a new audio generation model that is now available for creator testing through the Volcano Ark experience center, with enterprise API access also opened for invitation-based trials. The model treats vocals, background music, sound effects, and ambient noise as parts of one unified audio scene rather than separate tracks, allowing it to generate a complete soundtrack of up to nearly two minutes in a single inference run. According to the report, Seed Audio 1.0 supports more than 20 languages, including Chinese, English, and Japanese. It also supports zero-shot multimodal reference, meaning users can provide text or audio samples and have the system imitate the referenced voice timbre and style without retraining. A single prompt can arrange multi-character dialogue, music, and Foley-style effects, while timing can be controlled at 100-millisecond precision. The model has been integrated into the Doubao app. Individual users can try it through Volcano Ark, while international developers can access it via BytePlus. The report contrasts Doubao’s one-step generation approach with ElevenLabs Studio, which separates speech, music, and sound effects into three APIs. It also notes that ElevenLabs Pro costs $99 per month, while Doubao uses Volcano Engine’s token-based pricing model.

1060
ByteDance launches Doubao Seed Audio 1.0 to generate full audio scenes in a single pass
xAI
2026-07-10 11:52:13

xAI API Launches Secure Voice Cloning With Multilingual Support

xAI has added secure voice cloning to its API, enabling users to build production-ready voice models from one minute of speech. The feature supports multilingual output, streaming via REST and WebSocket, and a two-step voice ownership verification process.

220
xAI API Launches Secure Voice Cloning With Multilingual Support