Clones voice and emotion from a short sample for accent-free dubbing across 14 languages

16GB RAM recommended. 21GB+ storage recommended.
macOS 15+: M-series chips required.
Windows 10/11 64-bit: NVIDIA GPU with 8GB+ VRAM required.
Note: For NVIDIA GPUs, install a newer driver.Confucius4-TTS is a next-generation multilingual text-to-speech (TTS) and voice cloning engine developed and open-sourced by the NetEase Youdao Team as part of Youdao's "Confucius" Large Model 4.0 ecosystem. It fully supports high-quality synthesis and accent-free cross-lingual cloning across 14 languages: Chinese, English, Japanese, Korean, German, French, Spanish, Portuguese, Russian, Italian, Vietnamese, Thai, Indonesian, and Arabic, providing creators, developers, and enterprises with accessible AI voice solutions.
Confucius4-TTS is built on modern generative AI architecture, leveraging a Speech Encoder + LLM (Large Language Model) hybrid generative pipeline paired with advanced neural vocoders, achieving high-fidelity, low-latency audio generation.