Tencent open-sources AuK, a 1.5B speech model for voice editing and cloning

Tencent open-sources AuK, a 1.5B speech model for voice editing and cloning

N
News Editor
2026-09-10 12:14:24
Tencent Hunyuan has open-sourced AuK, a 1.5 billion-parameter speech generation and editing model that the company describes as an "audio version of Nano Banana." The model combines voice cloning, word-level correction, emotion transfer, accent removal, speaking rate and pitch adjustment, denoising, and multi-speaker and music separation in a single system, with natural-language controls. According to the announcement cited by BlockBeats, AuK can edit existing recordings directly. Users can change only the incorrect words in a clip instead of re-recording the full segment. It can also change emotion, timbre, or accent without altering the spoken content, and supports tasks such as lyric editing, removing laughter or coughs, and turning normal speech into a whisper. Tencent said AuK led several speech generation and general speech editing benchmarks. On SpeechEditBench, it scored 49.73, compared with 28.70 for the runner-up, Ming-UniAudio. Tencent also released AuK-Flash, a distilled version that needs only four inference steps and runs about 4.5 times faster. The paper notes that AuK still cannot reliably understand all free-form natural-language instructions and currently depends on a Prompt Enhancer to classify tasks, organize parameters, and rewrite prompts. The code and model weights have been released under the MIT license.

Tencent Hunyuan has open-sourced AuK, a 1.5 billion-parameter model for speech generation and editing, which it describes as an "audio version of Nano Banana." The system combines voice cloning, text replacement, emotion control, accent removal, speaking rate and pitch adjustment, denoising, and multi-speaker and music separation in one model, with natural-language control built in.

According to the BlockBeats flash update citing Beating AI, AuK can edit existing recordings directly. If only a few words are wrong, users can revise just those words instead of re-recording the full passage. It can also switch emotion, timbre, or accent while keeping the content unchanged. The model also supports lyric changes, removing laughter and coughing sounds, and converting normal speech into a whisper.

For voice cloning, AuK does not require a transcript for the reference audio in advance. Tencent said a new voice output can be generated with only a sample of speech and new text.

In Tencent's official tests, AuK ranked ahead on multiple speech generation and general speech editing benchmarks. On SpeechEditBench, it posted a score of 49.73, while Ming-UniAudio, listed as the second-place model, scored 28.70.

Tencent also released AuK-Flash, a faster version produced through distillation. The company said it requires only four inference steps and runs about 4.5x faster.

The paper also states that AuK still cannot reliably understand all arbitrary natural-language instructions. At this stage, it still relies on a Prompt Enhancer to identify the task, organize parameters, and rewrite the instruction.

The code and model weights have already been published under the MIT license.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
300

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.