ishandutta2007
Awesome-Text-to-Speech

๐ŸŽค A curated list of the latest and most influential tools, models, and resources in the Text-to-Speech sector. ๐ŸŒŸ Star if you like it! ๐ŸŒŸ

Last updated Aug 8, 2026
170
Stars
38
Forks
2
Issues
+3
Stars/day
Attention Score
78
Language breakdown
No language data available.
โ–ธ Files click to expand
README

Awesome Text-to-Speech Banner

Awesome Text-to-Speech (TTS) ๐Ÿ—ฃ๏ธ: Best AI Voice Generation Models & Tools 2025

AwesomeDiscord Awesome GitHub stars GitHub forks GitHub license Python Version PRs Welcome GitHub Sponsors Last Commit Contributors Follow on Twitter GitHub followers


๐Ÿš€ The Ultimate Guide to AI Voice Generation, Speech Synthesis, and Neural Voice Cloning

Welcome to the most comprehensive, meticulously curated, and continuously updated list of Text-to-Speech (TTS) resources. Whether you are looking for the best open-source TTS models of 2025, searching for low-latency TTS APIs for AI agents, or exploring high-fidelity voice cloning for content creation, you've found the right place.

[!TIP]
Looking for the best ElevenLabs alternatives? This repository tracks the rapidly evolving landscape of both commercial SaaS and local-first neural speech synthesis.

๐Ÿ—บ๏ธ Quick Navigation


๐Ÿ”Š Why Explore Text-to-Speech?

Text-to-Speech technology has moved beyond robotic voices. Today, it powers:

  • Accessibility First: High-quality screen readers for the visually impaired.
  • Automated Content Creation: Realistic voiceovers for YouTube, podcasts, and e-learning.
  • Next-Gen AI Agents: Real-time conversational AI with human-like prosody.
  • Multilingual Support: Instant translation and dubbing for global reach.
  • Personalization: Custom voice clones for gaming and virtual assistants.

๐Ÿ“ˆ Current State of Text-to-Speech (2025 Update)

The landscape of AI voice synthesis has shifted from basic concatenation to advanced Generative Speech Models. Key highlights:

  • Hyper-realistic and Natural Speech Synthesis: Innovations in deep learning and neural network architectures have led to highly natural, expressive, and emotionally nuanced synthetic voices. ๐ŸŽค
  • Next-Generation Architectures: The adoption of State Space Models (SSMs), Diffusion Models, and advanced transformer-based architectures is offering superior performance, efficiency, and voice quality in speech generation. ๐Ÿง 
  • Real-time Conversational AI: Significant advancements in reducing latency now enable real-time TTS, making conversational AI, virtual assistants, and live dubbing more natural and responsive. โšก
  • Advanced Voice Cloning and Style Transfer: Cutting-edge techniques allow for high-fidelity voice cloning from minimal audio samples and the transfer of speaking style and emotion across different voices. ๐ŸŽญ
  • Multilingual and Cross-Lingual TTS: Models are increasingly capable of generating speech in numerous languages with accurate pronunciation and intonation, breaking down language barriers.

๐Ÿ“š Comprehensive List of Text-to-Speech (TTS) Resources ๐ŸŒ

โ˜๏ธ Cloud-based & Commercial AI Voice Generation Platforms

Leading platforms offering robust, scalable, and high-quality Text-to-Speech APIs and services for various applications.

| ๐Ÿ› ๏ธ Service/Model | ๐Ÿข Organization | ๐ŸŒŸ Key Features | ๐Ÿ’ฐ Min. Monthly Subscription | ๐Ÿ‘ฅ Company Size | ๐Ÿ”— Link | | --- | --- | --- | --- | --- | --- | | NVIDIA NeMo | NVIDIA | Platform for building, training, and deploying generative AI models, including TTS and ASR. | Free (API Credits) / Enterprise | $3.3T+ (Market Cap) | NVIDIA NeMo | | Azure AI Speech | Microsoft | High-quality neural voices with advanced fine-tuning, emotion, and enterprise scalability. | Free (0.5M chars/mo) / PAYG | $3.2T+ (Market Cap) | Azure AI Speech | | Google Cloud TTS | Google | Powerful TTS API with a large variety of natural-sounding voices and extensive customization. | Free (1M+ chars/mo) / PAYG | $2.2T+ (Market Cap) | Google Cloud TTS | | AWS Polly | Amazon | Generative, Neural and Standard TTS voices with deep AWS ecosystem integration. | Free (1M+ chars/mo) / PAYG | $1.9T+ (Market Cap) | AWS Polly | | OpenAI TTS | OpenAI | High-quality, real-time streaming TTS models for applications requiring natural AI voices. | Pay-as-you-go ($5 free credit) | $852B (Valuation) | OpenAI TTS | | ElevenLabs | ElevenLabs | State-of-the-art AI voice generator offering realistic voices, voice cloning, and AI dubbing. | Free (10k chars/mo) / $5 | $11B (Valuation) | ElevenLabs | | Speechify | Speechify | Highly popular consumer text-to-speech with premium voices (including celebrity voices). | Free (Basic Voices) / $139/yr | $1.5B (Valuation) | Speechify | | Deepgram Aura | Deepgram | Specializing in low-latency TTS designed for real-time conversational AI and virtual interactions. | Free ($200 credit) / PAYG | $1.2B (Valuation) | Deepgram Aura | | Cartesia Sonic | Cartesia | Sub-100ms ultra-low latency TTS designed for real-time AI agents. | Free (API Credits) / PAYG | $200M (Valuation) | Cartesia | | Murf.ai | Murf.ai | AI voiceovers with a built-in video editor, ideal for creators and presentations. | Free (10 mins total) / $19 | $46M (Valuation) | Murf.ai | | LMNT | LMNT | Lightning-fast TTS API with exceptional naturalness, great for interactive voice apps. | Free (API Credits) / PAYG | $21M (Estimated) | LMNT | | Play.ht | Play.ht | Professional AI voices and "Ultra-Realistic" studio editor for long-form content. | Free (12.5k chars) / $39 | $15M (Valuation) | Play.ht | | Soniox TTS | Soniox | Real-time streaming TTS API for conversational AI voice agents in 60+ languages with multilingual voices. | Pay-as-you-go (~$0.70/hr) | $10M (Estimated) | Soniox | | Neets.ai | Neets | Extremely fast and affordable TTS APIs starting at $0.0004 per 1k characters. | Free (API Credits) / PAYG | $5M (Estimated) | Neets.ai | | Gandr | Gandr | TTS API for voice agents. Word error rate 1.98 percent against a 2.17 percent human reference on the same scorer, one voice in 23 languages, every render watermarked. Python/JS SDKs, LiveKit plugin, MCP server. | From $10/mo. Flat-rate unmetered streams from $150/mo | <$1M (Indie) | Gandr | | Spokio | Spokio | Offline macOS text-to-speech app with local voice cloning, batch export, and no cloud uploads. | Free (API Credits) / Enterprise | <$1M (Indie) | Spokio | | PHANTOM VOICES | PHANTOM VOICES | 10 free professional AI voice clones via public REST API. Zero cost, commercial rights cleared. 29 platform configs (Vapi, Retell AI, etc). Multilingual (9+ languages). AI-powered recommendation. | Free (API Credits) / Enterprise | <$1M (Indie) | PHANTOM VOICES | | RunAPI ElevenLabs SDK | RunAPI | Multi-language SDKs for ElevenLabs text-to-speech, dialogue generation, sound effects, transcription, and audio isolation workflows. | Pay-as-you-go | <$1M (Indie) | RunAPI ElevenLabs SDK |

๐Ÿ—๏ธ Open-Source Text-to-Speech Libraries & Local-First Projects

If you are looking for free text-to-speech models for commercial use or want to run TTS locally on a CPU, these open-source projects provide the best balance of quality and privacy.

Sound Wave Animation

| ๐Ÿ› ๏ธ Service/Model | ๐Ÿข Organization | ๐ŸŒŸ Key Features | ๐Ÿ—ฃ๏ธ Primary Language | ๐Ÿ“ Repository | | --- | --- | --- | --- | --- | | ๐Ÿธ Coqui TTS | Coqui | Best overall open-source TTS. Supports 1100+ languages, zero-shot voice cloning, and fine-tuning. | Python / Multilingual | GitHub stars | | ChatTTS | 2noise | Conversational text-to-speech model specially optimized for dialogue and natural conversational flow. | Python | GitHub stars | | OpenVoice | MyShell | Highly versatile and instant voice cloning that requires only a short audio clip. | Python | GitHub stars | | Fish Speech | Fish Audio | SOTA multilingual, multi-speaker model with superior naturalness. | Python | GitHub stars | | Chatterbox | Resemble AI | Advanced neural voice synthesis with emotion control and high-fidelity cloning. | Python | GitHub stars | | CosyVoice | Alibaba | Excellent multilingual and zero-shot voice cloning model capable of high fidelity. | Python | GitHub stars | | KittenTTS | KittenML | ONNX-based library for low-latency TTS without requiring a GPU. | Python / ONNX | GitHub stars | | F5-TTS | SWivid | Flow Matching TTS. Incredible naturalness and prosody using DiT architectures. | Python | GitHub stars | | Tortoise-TTS | James Betker | Powerful multi-voice TTS system known for its exceptional voice cloning capabilities. | Python | GitHub stars | | Piper | Rhasspy | Fastest local TTS. Optimized for low-end hardware and offline use. | C++ / Python | GitHub stars | | Amphion | Amphion | Open-source audio, music and speech generation toolkit containing multiple SOTA TTS models. | Python | GitHub stars | | Kokoro-82M | Hexgrad | Best SOTA CPU TTS. Ultra-fast, studio quality, only 82M parameters. | ONNX / Python | GitHub stars | | Parler-TTS | Hugging Face | Lightweight, controllable speech generation with high naturalness. | Python | HF | | Matcha-TTS | Shivam Mehta | Fast TTS architecture employing conditional flow matching, producing highly natural output. | Python | GitHub stars | | LocalMode | LocalMode | In-browser TTS. Runs Kokoro (29 voices) and other AI models 100% in the browser via WebGPU/WASM. No server, no API keys, offline after first load. | JavaScript / TypeScript | GitHub stars | | Vocello | PowerBeef | Native Mac & iPhone app. Qwen3-TTS with preset speakers, natural-language voice design, and voice cloning. Runs entirely on Apple Silicon with no Python runtime, faster than realtime on an 8 GB M2. | Swift / MLX | GitHub stars |

Advanced Voice Cloning & Neural Voice Synthesis ๐Ÿงฌ

Dedicated resources and examples focusing on the latest in voice replication and advanced synthetic voice generation.

  • XTTS-v2 by Coqui: A breakthrough in voice cloning, capable of replicating a voice from just a 6-second audio clip, preserving emotion and speaking style.
  • Resemble AI's Chatterbox: Offers advanced zero-shot voice cloning capabilities, enabling instant voice replication without extensive training data.
  • ElevenLabs Voice Cloning: Provides robust tools for creating highly realistic voice clones, suitable for personalized audio content.
  • Suno Bark: A transformer-based text-to-audio model that generates highly naturalistic, multilingual speech, music, and sound effects. It excels at expressive speech with nuances like laughter, sighs, and crying.
* Bark on GitHub
  • MeloTTS: A multi-language, multi-speaker Text-to-Speech model capable of generating high-quality audio.
* MeloTTS on GitHub

Hugging Face ๐Ÿค— - The Hub for TTS Models

Hugging Face has emerged as a central ecosystem for sharing, discovering, and experimenting with a vast array of pretrained Text-to-Speech models. Explore their extensive collection for diverse applications and research.

Notable Research Papers & Community Discussions ๐Ÿ“

Stay updated with the latest breakthroughs and discussions in the TTS community.

  • [[N] Baidu AI Can Clone Your Voice in Seconds](https://www.reddit.com/r/MachineLearning/comments/7zb2jm/nbaiduaicancloneyourvoiceinseconds/) (Reddit discussion on voice cloning technology)
  • [[R] Expressive Speech Synthesis with Tacotron](https://www.reddit.com/r/MachineLearning/comments/87klvo/rexpressivespeechsynthesisswith_tacotron/) (Reddit discussion on making TTS more human-like)
  • [[D] Realtime Neural Voice Style Transfer Feasibility and Implications](https://www.reddit.com/r/MachineLearning/comments/8opn4c/drealtimeneuralvoicetransfer/) (Discussion on the challenges and potential of real-time voice style transfer)
  • [[D] Is there an implementation of Neural Voice Cloning?](https://www.reddit.com/r/MachineLearning/comments/8o7mkt/disthereanimplementationofneural_voice/) (Community quest for neural voice cloning implementations)
  • [[D] Are the hyper-realistic results of Tacotron-2 and Wavenet not reproducible?](https://www.reddit.com/r/MachineLearning/comments/845uji/darethehyperrealisticresultsoftacotron2_and/) (Discussion on reproducibility in advanced TTS models)
  • [[P] Voice Style Transfer: Speaking like Kate Winslet](https://www.reddit.com/r/MachineLearning/comments/7a0wcv/pvoicetransferspeakinglikekatewinslet/) (Showcase of voice style transfer examples)
  • F5-TTS: A Strong Baseline for Zero-Shot Text-to-Speech with Diffusion Transformer: A highly influential paper on Flow Matching and DiT architectures for seamless voice cloning.
  • MambaVoiceCloning (MVC): Research on using State Space Models (SSMs) to achieve human-level speech synthesis with linear-time complexity.
  • VALL-E R: Robust and Efficient Zero-Shot TTS: Advancements in monotonic alignment for more stable autoregressive speech generation.
  • Reddit: The ElevenLabs Killer Quest (r/TTS): Community-driven search for high-fidelity, local open-source alternatives to proprietary SaaS.
  • TTS Arena by Artificial Analysis: The definitive community leaderboard for blind-testing the naturalness of modern TTS models.
  • [[D] Why Flow Matching is replacing traditional Diffusion in TTS](https://www.reddit.com/r/MachineLearning/): Technical deep-dive into the "straightening" of ODE paths for faster, higher-quality audio generation.
  • An Automated Failure-Mode QA Framework for Neural Text-to-Speech Systems: A production case study introducing automated failure-mode detection for TTS output, with ttsproof (https://github.com/Mormolykos/ttsproof) as its open-source implementation.

Exemplary Code Samples & Project Demos ๐Ÿ’ป

A collection of influential code repositories and product demonstrations showcasing various Text-to-Speech implementations and their output quality.

| Project/Samples | Pretrained Models | Code Link | Paper/Arxiv ID | Output Quality | Year of Launch | Description | | --- | :---: | :---: | :---: | :---: | :---: | --- | | [Fish Speech v1.5 | -- | Code | Codebase | A+ | 2026 | SOTA multilingual, multi-speaker model with superior naturalness. | | Kokoro-82M Samples | -- | Code | -- | A | 2025 | Ultra-efficient CPU-based model with studio-quality output. | | F5-TTS Samples | -- | Code | 2410.06885 | A | 2024 | Diffusion-based zero-shot cloning with impressive prosody. | | MeloTTS Samples | -- | Code | Codebase | B | 2024 | Multilingual, multi-speaker TTS model for high-quality audio generation. | | Parler-TTS Samples | -- | Code | 2402.01912 | B | 2024 | Samples from a lightweight model producing natural-sounding speech. | | XTTS-v2 Samples | -- | Code | 2309.02055 | A | 2023 | Demonstrations of Coqui's advanced voice cloning with emotion transfer. | | Bark Samples (Suno.ai) | -- | Code | -- | A | 2023 | Samples from Suno's expressive text-to-audio model, including non-speech sounds. | | rayhane's Tacotron2 Samples | -- | -- | -- | D | 2019 | Audio samples from an early Tacotron 2 implementation. | | Google Tacotron + Style Transfer Sample (Official) | -- | -- | 1803.09047 | A | 2018 | Official samples showcasing prosody and style transfer with Tacotron. | | NVIDIA's WaveGlow Samples | Download Model | Code | 1811.00002 | A | 2018 | High-fidelity audio generated by NVIDIA's WaveGlow vocoder. | | NVIDIA's Tacotron2 + WaveGlow Samples | Download Model | Code | -- | A | 2018 | Combined high-quality speech synthesis from Tacotron 2 and WaveGlow. | | mazzzystar's Tacotron-WaveRNN Samples | Get Model | Code | -- | A | 2018 | Demonstrations from a Tacotron and WaveRNN hybrid model. | | syang1993's Tacotron + Style Transfer Samples | Model ErnstTmp (232k iter) | -- | 1803.09047 and 1803.09017 | C | 2018 | Samples demonstrating Tacotron with global style tokens for voice style transfer. | | Kyubyong's Tacotron on LJ Dataset Samples | Download model | -- | -- | D | 2018 | Audio generated from Tacotron trained on the LJSpeech dataset. | | Kyubyong's Tacotron on Nick Dataset Samples | -- | -- | -- | D | 2018 | Tacotron samples from the Nick dataset. | | Kyubyong's Tacotron on Web Dataset Samples | Download model | -- | -- | D | 2018 | Tacotron speech output from the Web dataset. | | Kyubyong's Expressive Tacotron Samples | -- | Code | 1803.09047 | D | 2018 | Samples demonstrating expressive speech synthesis with Tacotron. | | Kyubyong's DC-TTS on Nick Dataset Samples | -- | -- | -- | D | 2018 | DC-TTS samples generated from the Nick dataset. | | Baidu's Deep Voice Samples (Official) | -- | -- | -- | D | 2017 | Official audio demonstrations from Baidu's Deep Voice project. | | Baidu's Deep Voice 3 Samples (Official) | -- | -- | 1710.07654 | B | 2017 | Official samples from Deep Voice 3, showcasing advanced speech synthesis. | | Google Tacotron2 Samples (Official) | -- | -- | 1712.05884 | A | 2017 | Official, high-quality audio samples from the groundbreaking Tacotron 2 model. | | DeepMind Neural Discrete Representation Learning Samples (Official) | -- | -- | 1711.00937 | B | 2017 | Samples demonstrating speech generated using VQ-VAE for neural discrete representation learning. | | r9y9's Wavenet Vocoder Tacotron2 Samples | Download Tacotron2 model - Download Wavenet model - Get models | -- | 1712.05884 and 1611.09482 | B | 2017 | Samples from a Tacotron 2 and WaveNet vocoder combination. | | dhgrs's Implementation of Neural Discrete Representation Learning Samples | Download Model | Code | 1711.00937 | D | 2017 | Audio generated using a Chainer implementation of VQ-VAE for speech. | | keithito's Tacotron Samples | Get model | -- | -- | D | 2017 | Audio samples from keithito's Tacotron implementation. | | Kyubyong's DC-TTS on LJ Dataset Samples | Get model | -- | -- | D | 2017 | DC-TTS generated speech from the LJSpeech dataset. | | Kyubyong's DC-TTS Kate Samples | -- | -- | -- | D | 2017 | DC-TTS samples featuring the "Kate" voice. | | andabi's Deep Voice Conversion | -- | -- | -- | D | 2017 | Demonstrations of deep voice conversion techniques. | | Facebook Loop Samples (Official) | Get model | -- | -- | D | 2017 | Official audio samples from Facebook's Loop project. | | mazzzystar's RandomCNN Voice Transfer | -- | -- | 1712.08363 | D | 2017 | Speech conversion samples using Random CNNs. | | Griffin-Lim Samples | -- | -- | -- | A | 1984 | Classic samples from the Griffin-Lim algorithm for spectrogram inversion. |

Work in Progress & Future of Text-to-Speech ๐Ÿšง

Ongoing projects and cutting-edge research shaping the next generation of AI voice synthesis.

  • https://github.com/ErnstTmp is implementing https://arxiv.org/abs/1807.06736
  • https://github.com/nii-yamagishilab/self-attention-tacotron
  • https://github.com/nii-yamagishilab/tacotron2
If I missed your output sample/demo in this consolidation, just add and send a pull request. I will be more than happy to add it. Thanks!

Codelabs & Interactive Tutorials ๐Ÿงช

Practical guides and interactive notebooks for experimenting with Text-to-Speech models.

Product Demos & Showcase Videos ๐ŸŽฅ

Visual demonstrations of advanced Text-to-Speech and voice cloning in action.

Related Works & Foundational Research ๐Ÿ“š

Broader projects and research efforts that contribute to the Text-to-Speech ecosystem.

  • https://github.com/tensorflow/magenta

Arxiv Sanity Preserver - Key Papers in Speech Synthesis ๐Ÿ“„

Explore influential academic papers and preprints in the field of Text-to-Speech and voice AI.

Star History

Star History Chart

๐Ÿ’ฌ Community & Support for Text-to-Speech Enthusiasts

Connect with the community, get support, and stay informed about the latest in TTS.

  • ๐Ÿ“š Documentation: Check out our official documentation for detailed guides and tutorials on utilizing TTS technologies.
  • ๐Ÿ—ฃ๏ธ Forum: Join our community forum to ask questions, share your Text-to-Speech projects, and connect with other users and developers.
  • ๐Ÿ’ฌ Discord: Chat with us on Discord for real-time support and discussions on AI voice generation.
  • ๐Ÿฆ Twitter: Follow us on Twitter for the latest news, updates, and insights into the world of synthetic speech.
  • ๐Ÿฆ Github: Follow me on Github for the latest commits and updates on this and other AI projects.

๐ŸŽฏ Key Use Cases for AI Voice Generation

Explore how Text-to-Speech and Voice Cloning are being used across industries:

  • ๐ŸŽ™๏ธ Podcast Automation: Convert written articles into high-quality audio episodes instantly.
  • ๐ŸŽฎ Video Game Development: Dynamic NPC dialogue using local-first TTS like Piper or Kokoro.
  • ๐Ÿ› ๏ธ Customer Support: Low-latency conversational AI for 24/7 automated support.
  • ๐Ÿ“– Accessible E-Learning: Making educational content accessible with natural-sounding voices.
  • ๐ŸŽฌ Content Localization: Dubbing videos into multiple languages while preserving the original speaker's emotion.

โ“ Frequently Asked Questions (FAQ) & SEO Insights

What is the best open-source Text-to-Speech model in 2025?

As of 2025, Kokoro-82M is widely considered the best for CPU-based local inference due to its studio quality and small footprint. For high-fidelity and expressive speech, F5-TTS and Fish Speech are leading the way in naturalness.

Are there free ElevenLabs alternatives for voice cloning?

Yes! Projects like Coqui XTTS-v2 and OpenVoice offer high-quality voice cloning for free. If you are looking for local-first alternatives, check out F5-TTS.

How do I achieve low-latency TTS for AI agents?

To achieve sub-200ms latency, it is recommended to use Deepgram Aura, Cartesia Sonic, or optimized local models like Piper (C++ implementation) and Kokoro-82M with ONNX runtime.

Can I use these TTS models for commercial projects?

Many models listed here (like OpenAI TTS, ElevenLabs, and Azure AI Speech) have clear commercial tiers. For open-source models, look for those with MIT or Apache 2.0 licenses, such as Piper and Kokoro.

๐Ÿ’– Support & Sponsorship

If you find this collection of Text-to-Speech resources helpful, or if it has saved you time and effort in your AI voice generation endeavors, please consider sponsoring the development. Your support helps maintain the project, add new cutting-edge models and tools, and keep this initiative open-source and accessible to everyone.

Sponsor @ishandutta2007 on GitHub

Every contribution, no matter how small, makes a huge difference in advancing the Text-to-Speech landscape! ๐Ÿ™

๐Ÿ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.

ยฉ 2026 GitRepoTrend ยท ishandutta2007/Awesome-Text-to-Speech ยท Updated daily from GitHub