Uses natural-language instructions to unify speech generation, editing, enhancement, and source separation
Download with PC ClientMinimum 32GB RAM. 36GB+ storage recommended.
macOS 15+: M-series chips required.
Windows 10/11: NVIDIA GPU with 12GB+ VRAM required.
Note: For NVIDIA GPUs, install a newer driver.AuK is an open-source speech generation and editing foundation model developed by Tencent. In simple terms, it can do much more than turn text into speech: it can also modify existing audio through natural-language instructions.
For example, you can provide a reference voice and ask AuK to speak new text in a similar voice. You can also ask it to replace a word in an existing recording, change the speaking speed or pitch, modify the speaker's emotion, remove noise, or separate different speakers from an audio recording.
One of AuK's key characteristics is that these different speech tasks are handled through a unified natural-language instruction interface, rather than requiring a completely separate model for each function.
1. AI Speech Generation
AuK supports zero-shot text-to-speech (Zero-shot TTS). By providing a reference voice, users can generate new speech in a similar voice without specifically training a model for that speaker.
It also supports instruction-based TTS, where users can describe the desired voice characteristics with text without providing a reference recording.
2. Direct Editing of Existing Speech
AuK is not limited to text-to-speech generation. It can also edit existing recordings.
For example, it can:
This makes speech editing more like editing a text document, allowing certain parts of a recording to be changed without having to recreate the entire recording.
3. Control Over Speaking Style and Voice Characteristics
AuK can modify how speech sounds while keeping the spoken content intact. It supports tasks such as:
In other words, AuK can control not only what is being said, but also how it is said.
4. Speech Enhancement and Audio Separation
For low-quality recordings, AuK can perform denoising, dereverberation, and speech restoration.
It can also separate voices from mixed audio—for example, keeping a particular speaker, separating speakers in a conversation, extracting vocals from music, or isolating a target speaker based on what they say.
AuK can be viewed as an AI audio-editing engine, rather than simply another TTS system.
Potential applications include:
A particularly important characteristic is that AuK brings speech generation, speech editing, enhancement, and source separation into a unified model and instruction framework.
AuK was developed and open-sourced by Tencent's Hunyuan team.
Both the source code and model weights have been publicly released. The project is licensed under the MIT License, allowing developers to use, modify, and build upon the released software and model components.
AuK is a 1.5-billion-parameter speech foundation model, trained on millions of hours of diverse audio data.
Its implementation uses a modern generative AI architecture built around components including a text encoder, an audio VAE, and Transformer/flow-matching-based generation and editing modules. The source code specifically shows the use of Qwen2.5-Omni components, BigVGAN Flow VAE, and Flux2Edit/CFMEdit in the model pipeline.
The project provides two main model variants: