Tencent open-sources AuK, a 1.5B speech model for voice editing and cloning
Tencent Hunyuan has open-sourced AuK, a 1.5 billion-parameter speech generation and editing model that the company describes as an "audio version of Nano Banana." The model combines voice cloning, word-level correction, emotion transfer, accent removal, speaking rate and pitch adjustment, denoising, and multi-speaker and music separation in a single system, with natural-language controls. According to the announcement cited by BlockBeats, AuK can edit existing recordings directly. Users can change only the incorrect words in a clip instead of re-recording the full segment. It can also change emotion, timbre, or accent without altering the spoken content, and supports tasks such as lyric editing, removing laughter or coughs, and turning normal speech into a whisper. Tencent said AuK led several speech generation and general speech editing benchmarks. On SpeechEditBench, it scored 49.73, compared with 28.70 for the runner-up, Ming-UniAudio. Tencent also released AuK-Flash, a distilled version that needs only four inference steps and runs about 4.5 times faster. The paper notes that AuK still cannot reliably understand all free-form natural-language instructions and currently depends on a Prompt Enhancer to classify tasks, organize parameters, and rewrite prompts. The code and model weights have been released under the MIT license.

