跳到主要内容

AI 语音合成开源项目排行

按 star 排行的开源项目速查

匹配 28 个项目数据采集于 2026-08-04
MoneyPrinterTurbo★ 102kfork 15kPythonharry0703

利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.

ai-video-generatorcontent-creationffmpeginstagram-reels
unsloth★ 70kfork 6.3kPythonunslothai官网

Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, GLM and other models.

agentdeepseekfine-tuninggemma
GPT-SoVITS★ 60kfork 6.6kPythonRVC-Boss

1 min voice data can also be used to train a good TTS model! (few shot voice cloning)

text-to-speechttsvitsvoice-clone
TTS★ 46kfork 6.2kPythoncoqui-ai2024-08 后未更新官网

🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production

deep-learningglow-ttshifiganmelgan
OpenMontage★ 45kfork 5.5kPythoncalesthio官网

World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Tur…

agentagentic-aiaiclaude
ChatTTS★ 40kfork 4.3kPython2noise官网

A generative speech model for daily dialogue.

agentchatchatgptchattts
OpenVoice★ 37kfork 4.1kPythonmyshell-ai2025-04 后未更新官网

Instant voice cloning by MIT and MyShell. Audio foundation model.

text-to-speechttsvoice-clonezero-shot-tts
MockingBird★ 37kfork 5.2kPythonbabysor

🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time

aideep-learningpytorchspeech
VoxCPM★ 35kfork 4.0kPythonOpenBMB官网

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

audiodeeplearningminicpmmultilingual
CosyVoice★ 23kfork 2.6kPythonQwenAudio官网

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

audio-generationcantonesechatbotchatgpt
index-tts★ 22kfork 2.7kPythonindex-tts

An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

bigvgancross-lingualindexttstext-to-speech
dia★ 19kfork 1.7kPythonnari-labs

A TTS model capable of generating ultra-realistic dialogue in one pass.

aiopen-weighttext-to-speech
pyvideotrans★ 19kfork 2.3kPythonjianchang512官网

Translate the video from one language to another and embed dubbing & subtitles.

speech-to-texttext-to-speechvideo-transition
leon★ 17kfork 1.5kTypeScriptleon-ai官网

🧠 Leon is your open-source personal assistant.

aiai-agentai-assistantartificial-intelligence
sherpa-onnx★ 14kfork 1.6kC++k2-fsa官网

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Int…

aarch64androidarm32asr
supertonic★ 14kfork 1.5kSwiftsupertone-inc官网

Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.

cppcsharpfluttergo
voice-pro★ 12kfork 1.7kPythonabus-aikorea官网

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper aud…

audiobookfaster-whispergradiokaraoke
edge-tts★ 12kfork 1.1kPythonrany2官网

Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key

speech-synthesistext-to-speechtts
TTS★ 10kfork 1.3kJupyter Notebookmozilla2023-11 后未更新

:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)

dataset-analysisdeep-learningganttsglow-tts
Amphion★ 10kfork 847Pythonopen-mmlab官网

Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researcher…

audio-generationaudio-synthesisaudioldmaudit
espnet★ 9.9kfork 2.4kPythonespnet官网

End-to-End Speech Processing Toolkit

chainerdeep-learningend-to-endkaldi
EmotiVoice★ 8.5kfork 756Pythonnetease-youdao2024-08 后未更新

EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine

aideep-learningemotionemotivoice
vits★ 7.9kfork 1.4kPythonjaywalnut3102023-12 后未更新官网

VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

deep-learningpytorchspeech-synthesistext-to-speech
mlx-audio★ 7.7kfork 682PythonBlaizzy官网

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis o…

apple-siliconaudio-processingmlxmultimodal
MeloTTS★ 7.6kfork 1.1kPythonmyshell-ai2024-12 后未更新

High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.

chineseenglishfrenchjapanese
espeak-ng★ 6.7kfork 1.3kCespeak-ng

eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.

androidespeakespeak-ngspeech-synthesis
StyleTTS2★ 6.3kfork 694Pythonyl45792024-08 后未更新

StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

adversarial-trainingdeep-learningdiffusion-modelsgan
argmax-oss-swift★ 6.3kfork 593Swiftargmaxinc

On-device Speech AI for Apple Silicon

inferenceiosmacospyannote

AI 语音合成 · GitHub star 排行 Top 28

文本转语音、声音克隆与语音识别的开源实现。下表按 GitHub star 数从高到低排列,共收录 28 个相关开源项目, 使用上方的筛选器可以按活跃度、编程语言进一步过滤。

排名项目Star语言最后更新
1harry0703/MoneyPrinterTurbo
利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.
102kPython2026-08
2unslothai/unsloth
Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, GLM and other models.
70kPython2026-08
3RVC-Boss/GPT-SoVITS
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
60kPython2026-07
4coqui-ai/TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
46kPython2024-08
5calesthio/OpenMontage
World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Tur…
45kPython2026-08
62noise/ChatTTS
A generative speech model for daily dialogue.
40kPython2026-04
7myshell-ai/OpenVoice
Instant voice cloning by MIT and MyShell. Audio foundation model.
37kPython2025-04
8babysor/MockingBird
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
37kPython2026-03
9OpenBMB/VoxCPM
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
35kPython2026-07
10QwenAudio/CosyVoice
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
23kPython2026-05
11index-tts/index-tts
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
22kPython2026-08
12nari-labs/dia
A TTS model capable of generating ultra-realistic dialogue in one pass.
19kPython2025-11
13jianchang512/pyvideotrans
Translate the video from one language to another and embed dubbing & subtitles.
19kPython2026-08
14leon-ai/leon
🧠 Leon is your open-source personal assistant.
17kTypeScript2026-08
15k2-fsa/sherpa-onnx
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Int…
14kC++2026-08
16supertone-inc/supertonic
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
14kSwift2026-07
17abus-aikorea/voice-pro
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper aud…
12kPython2026-07
18rany2/edge-tts
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
12kPython2026-03
19mozilla/TTS
:robot: :speech_balloon: Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)
10kJupyter Notebook2023-11
20open-mmlab/Amphion
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researcher…
10kPython2026-03
21espnet/espnet
End-to-End Speech Processing Toolkit
9.9kPython2026-08
22netease-youdao/EmotiVoice
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
8.5kPython2024-08
23jaywalnut310/vits
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
7.9kPython2023-12
24Blaizzy/mlx-audio
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis o…
7.7kPython2026-08
25myshell-ai/MeloTTS
High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.
7.6kPython2024-12
26espeak-ng/espeak-ng
eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.
6.7kC2026-08
27yl4579/StyleTTS2
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
6.3kPython2024-08
28argmaxinc/argmax-oss-swift
On-device Speech AI for Apple Silicon
6.3kSwift2026-08

怎么看这份榜单

star 数反映的是知名度而不是代码质量——教程、清单类仓库天然更容易积累 star, 而一些工程质量极高的底层库反而不显眼。选型时建议把「最后更新」这一列一起看: 超过一年没有提交的项目,往往意味着维护已经停滞,接入前需要评估风险。

想看其他方向的排行,可以返回GitHub 高星项目榜切换维度。

常见问题

AI 语音合成方向最值得关注的项目是哪个?
按 star 数排序,当前榜首是 MoneyPrinterTurbo。不过 star 高低只代表知名度,选型时建议结合项目的最后更新时间、issue 响应速度和自己的技术栈综合判断。
这个榜单多久更新一次?
榜单为人工触发的快照式采集,当前数据采集于 2026-08-04。开源项目的 star 数每日变动,排名可能与你访问 GitHub 时略有出入。