English
Discover amazing AI tools and applications
An upgraded TTS system featuring multilingual support, real-time style switching and efficient inference
a node-based user interface for Stable Diffusion.
Turn Text into Realistic Podcasts with Multi-Speaker, Multilingual & Emotional Speech
Zero-shot TTS system by Bilibili featuring fast voice cloning, cross-lingual synthesis, precise emotion/speed control
Generate 1-minute videos quickly with only 6GB of low VRAM.
Automatically generates HD short videos with script, footage, voiceover, subtitles and music
Clone voice in 5 seconds — GPT-SoVITS enables multilingual AI speech.
Zero-shot voice cloning and natural-language-driven voice design across multiple languages and dialects
Supporting Mandarin, English and Cantonese, with natural speech synthesis and zero-shot voice cloning
Animate Digital Humans' Lip Movements
Automatically handles video translation, subtitle generation and dubbing
Voice cloning with text-tag control over 11 emotions, 4 paralinguistic sounds (like laughter/breathing), and 14 Chinese dialects
Zero-shot voice cloning, supports multiple languages, allows voice parameter control
Zero-shot cloning and emotional control across 23 languages
An ultra-fast, lightweight open-source music model, delivering commercial-grade audio on local hardware with less than 4GB VRAM
Supporting 600+ languages, voice design, voice cloning, natural speech, and ultra-fast inference
Highly expressive, long-form, multi-speaker conversational audio generation
Upgraded model architecture with greatly improved sound quality and song integrity, supports ultra-long duration and multilingual creation
Creates high-quality realistic sound effects from text or video with only English prompts supported
Featuring ultra-realistic 48kHz zero-shot voice cloning with rich emotional and lifelike expressiveness
Turns lyrics into songs in seconds, generates music by style keywords.
A speech synthesis system that generates multi-speaker conversations with voice cloning and multilingual support.
Multilingual support for 52 languages/dialects and exceptional robustness in song and contextual transcription
An open-source fast-thinking multilingual translation model supporting 33 languages
Multi-language emotional expression, and real-time streaming generation to enable human-like natural speech with low-resource cross-scenario deployment.
Clones voices from short audio and generates natural speech.
Multilingual speech recognition, emotion & audio event detection—efficient and accurate
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
multilingual, real-time/offline recognition, easy to use and efficient
Delivers single-pass transcription, speaker diarization, and timestamping for long audio across 50+ languages
Real-time talking-head framework, high-fidelity, long-duration stable audio-visual synchronization
An open-source translation model, 33 languages + 5 dialects, accurate & flexible
A joint zero-shot SVS project supporting multilingual synthesis and dual-mode control
An expressive open-source text-to-speech model supporting 31 languages, featuring stable zero-shot voice cloning and precise inline pause control
An open-source LLM-DiT speech model by Xiaohongshu featuring zero-shot cloning for 24 languages and 21 dialects
Turns static portraits into video/audio-driven 3D models in real time
Generate high-quality classical music, supports generation by period, composer and instrumentation
An all-in-one local AI assistant supporting cross-platform operation and mobile remote control
A desktop graphical tool developed by ValueCell-ai based on OpenClaw, featuring one-click installation
Transforms simple text prompts into up to 6 minutes of studio-quality music and lifelike sound effects
A large model runner with a visual interface.
Supporting 17 languages, with accurate dialect and low-volume recognition
Generation and understanding, featuring high-fidelity song synthesis and controllable structural creation
An ultra-lightweight TTS tool, supporting multilingual speech synthesis and zero-shot voice cloning on ordinary CPU
Taming Bad Noise for Effective Video Object Removal
A desktop client for ChatGPT, Claude and other LLMs
PartPacker enables part-level 3D object generation from single-view images
Turns a simple description into a high-fidelity sound effect up to 30 seconds long, perfect for video voiceovers and game asset creation.
Removing hard-coded subtitles from videos and text watermarks from images with lossless resolution
Generates depth-aware 3D panoramas and scene models from single images
Supporting high-quality TTS and zero-shot voice cloning with extremely high timbre similarity
Add perfectly fitting foley sounds to silent videos.
Get up and running with large language models
A workflow automation platform supporting no-code/code dual-mode building
A security-first workflow automation tool with enhanced reliability and efficiency, supporting visual drag-and-drop operation
Clones voice and emotion from a short sample for accent-free dubbing across 14 languages
An easy-to-deploy, extensible open-source AI chatbot
A smart-searchable library of 2000+ ready-to-use n8n automation workflows
A lightweight, efficient open-source document parser that accurately converts PDFs, images, and e-books into Markdown/JSON
Zero-shot voice cloning, emotion expression capabilities
Zero-shot voice cloning and expressive editing of emotion, style, and paralinguistic cues
Making local speech-to-text and translation easy.
Zero-shot voice cloning, emotion expression and streaming inference
Generates pixel-aligned high-fidelity 3D models with PBR textures from a single image
An open-source enterprise AI assistant integrating RAG pipelines, multi-modal interaction, and workflow orchestration.
A free video editor featuring drag-and-drop simplicity, rich effects, and 4K video export
Easily displays technical and metadata information of audio and video files
A self-hosted machine translation tool supporting multi-language translation, usable offline with controllable data privacy
Generates lightweight interactive 3D Gaussian scenes compatible with mainstream 3D software from one single image rapidly
A cross-platform video editor offering broad format support, powerful filters, and easy timeline editing
Seamless integration with chat apps like WeChat, DingTalk, and Feishu for multi-agent collaborative workflows
A lightweight and easy-to-use multi-terminal personal AI assistant, supporting multi-channel connection and custom skills
A cross-platform software for high-performance live streaming and screen recording