AI 语音合成 · GitHub star 排行 Top 28
文本转语音、声音克隆与语音识别的开源实现。下表按 GitHub star 数从高到低排列,共收录 28 个相关开源项目, 使用上方的筛选器可以按活跃度、编程语言进一步过滤。
| 排名 | 项目 | Star | 语言 | 最后更新 |
|---|---|---|---|---|
| 1 | harry0703/MoneyPrinterTurbo 利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow. | 102k | Python | 2026-08 |
| 2 | unslothai/unsloth Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, GLM and other models. | 70k | Python | 2026-08 |
| 3 | RVC-Boss/GPT-SoVITS 1 min voice data can also be used to train a good TTS model! (few shot voice cloning) | 60k | Python | 2026-07 |
| 4 | coqui-ai/TTS 🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production | 46k | Python | 2024-08 |
| 5 | calesthio/OpenMontage World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Tur… | 45k | Python | 2026-08 |
| 6 | 2noise/ChatTTS A generative speech model for daily dialogue. | 40k | Python | 2026-04 |
| 7 | myshell-ai/OpenVoice Instant voice cloning by MIT and MyShell. Audio foundation model. | 37k | Python | 2025-04 |
| 8 | babysor/MockingBird 🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time | 37k | Python | 2026-03 |
| 9 | OpenBMB/VoxCPM VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning | 35k | Python | 2026-07 |
| 10 | QwenAudio/CosyVoice Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability. | 23k | Python | 2026-05 |
| 11 | index-tts/index-tts An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System | 22k | Python | 2026-08 |
| 12 | nari-labs/dia A TTS model capable of generating ultra-realistic dialogue in one pass. | 19k | Python | 2025-11 |
| 13 | jianchang512/pyvideotrans Translate the video from one language to another and embed dubbing & subtitles. | 19k | Python | 2026-08 |
| 14 | leon-ai/leon 🧠 Leon is your open-source personal assistant. | 17k | TypeScript | 2026-08 |
| 15 | k2-fsa/sherpa-onnx Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Int… | 14k | C++ | 2026-08 |
| 16 | supertone-inc/supertonic Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX. | 14k | Swift | 2026-07 |
| 17 | abus-aikorea/voice-pro Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper aud… | 12k | Python | 2026-07 |
| 18 | rany2/edge-tts Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key | 12k | Python | 2026-03 |
| 19 | mozilla/TTS :robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts) | 10k | Jupyter Notebook | 2023-11 |
| 20 | open-mmlab/Amphion Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researcher… | 10k | Python | 2026-03 |
| 21 | espnet/espnet End-to-End Speech Processing Toolkit | 9.9k | Python | 2026-08 |
| 22 | netease-youdao/EmotiVoice EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine | 8.5k | Python | 2024-08 |
| 23 | jaywalnut310/vits VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech | 7.9k | Python | 2023-12 |
| 24 | Blaizzy/mlx-audio A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis o… | 7.7k | Python | 2026-08 |
| 25 | myshell-ai/MeloTTS High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean. | 7.6k | Python | 2024-12 |
| 26 | espeak-ng/espeak-ng eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents. | 6.7k | C | 2026-08 |
| 27 | yl4579/StyleTTS2 StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models | 6.3k | Python | 2024-08 |
| 28 | argmaxinc/argmax-oss-swift On-device Speech AI for Apple Silicon | 6.3k | Swift | 2026-08 |
怎么看这份榜单
star 数反映的是知名度而不是代码质量——教程、清单类仓库天然更容易积累 star, 而一些工程质量极高的底层库反而不显眼。选型时建议把「最后更新」这一列一起看: 超过一年没有提交的项目,往往意味着维护已经停滞,接入前需要评估风险。
想看其他方向的排行,可以返回GitHub 高星项目榜切换维度。