1
ElevenLabs' expressive text-to-speech model with audio tags and 70+ languages
8.8/10
What it is ElevenLabs' speech models power its voiceover, dubbing and audiobook tools. The flagship, Eleven v3, reached general availability in February 2026 after an alpha in mid-2025. The lineup around it: v3 Conversational for real-time agents (about 280 ms), Multilingual v2 for steady long-form narration, and Flash v2.5 for low-latency bulk work (about 75 ms). ElevenLabs' CLI made v3… Read more →
Pros
- Very expressive delivery with audio tags
- 70+ languages on v3
- Voice cloning from $6/month
- Mature API and editor
Cons
- Free plan has no commercial licence
- Credits burn quickly on long projects
- 5,000-character cap per v3 request
Pricing: Free plan, paid from $6/mo · text-to-speech voice-cloning multilingual audio-tags elevenlabs
2
OpenAI's open-source speech recognition for transcripts and subtitles in 99 languages
★ 110k · MIT · updated 2026-08-31
8.2/10
What it is Whisper is OpenAI's open-source speech recognition model, released under the MIT licence. It transcribes and translates speech in 99 languages and detects which language is spoken. The newest open checkpoint is turbo (large-v3-turbo). It is an 809M-parameter pruned version of large-v3, about 8x faster with little loss in accuracy. OpenAI's paid API still offers whisper-1, but it… Read more →
Pros
- MIT licence, free for commercial use
- 99 languages
- Runs offline for privacy
- Huge ecosystem of ports and apps
Cons
- Can make up text during silence
- No built-in speaker separation
- Open line not updated since turbo
Pricing: Open source · speech-to-text open-source subtitles transcription openai
3
Resemble AI's MIT-licensed TTS with zero-shot voice cloning and built-in watermarking
★ 27k · MIT · updated 2026-07-21
7.6/10
What it is Chatterbox is Resemble AI's open-source text-to-speech family, released under the MIT licence. The current general-purpose model, Chatterbox Multilingual V3 (June 2026), is a 0.5B-parameter model that clones a voice from a short sample in 23+ languages. Next to it sit Chatterbox-Turbo (350M, English, low-latency, with paralinguistic tags) and Chatterbox-Nano (110M, English, runs on a CPU). Key features… Read more →
Pros
- MIT licence
- Zero-shot cloning in 23+ languages
- PerTh watermark on every output
- Turbo and Nano variants for speed and CPU
Cons
- Needs a GPU for comfortable speed (except Nano)
- Quality varies by language
- Easy cloning needs responsible use
Pricing: Open source · open-source voice-cloning text-to-speech watermarking multilingual
4
Cartesia's low-latency TTS for voice agents and narration, 44 languages
7.6/10
What it is Sonic is Cartesia's text-to-speech family. It is built on state space models (SSMs) rather than transformers, which makes it very fast. Sonic-3.6, generally available since 27 August 2026, is the current version. Cartesia also makes Ink-2, a streaming speech-to-text model. Cartesia pitches Sonic mainly at real-time voice agents, but it also works for narration and voiceovers. Key… Read more →
Pros
- Sub-90 ms latency
- Commercial licence from $5/month
- 44 languages
- Instant voice cloning on Pro
Cons
- Aimed at voice agents more than creators
- No long-form editing studio
- Few documented creator-app integrations
Pricing: Free plan, paid from $5/mo · text-to-speech low-latency voice-agents voice-cloning api
5
Tiny Apache-licensed TTS model with 54 voices in 8 languages that runs anywhere
★ 9.0k · Apache-2.0 · updated 2025-08-06
7.4/10
What it is Kokoro is an open-weight text-to-speech model with only 82 million parameters, published by the developer hexgrad. It is built on the StyleTTS 2 architecture. Version 1.0 (January 2025) is still the current release. It became popular because it sounds far better than its size suggests and runs quickly on a CPU, a laptop or even in a… Read more →
Pros
- Apache-2.0, free commercial use
- Runs fast on CPU and in the browser
- 54 voices
- Very cheap when hosted
Cons
- Only 8 languages
- No voice cloning
- Limited emotional range
- No release since January 2025
Pricing: Open source · open-weights text-to-speech lightweight offline apache-2
6
OpenAI's steerable text-to-speech API with 13 voices and prompt-based style control
7.4/10
What it is OpenAI's text-to-speech family is API-first. The current recommended model is gpt-4o-mini-tts. Its latest snapshot is dated December 2025. You tell it what to say and also how to say it: tone, emotion, pace and accent come from a plain-language instruction. The older tts-1 and tts-1-hd are still available. For live voice agents, OpenAI points to its separate… Read more →
Pros
- Plain-language style control
- Low per-token pricing
- Streaming output
- Clear consent and disclosure rules
Cons
- Only 13 voices, optimised for English
- API only, no editor
- Custom voices limited to eligible customers
Pricing: Paid · text-to-speech api openai steerable developers