1
ElevenLabs' expressive text-to-speech model with audio tags and 70+ languages
8.8/10
What it is ElevenLabs' speech models power its voiceover, dubbing and audiobook tools. The flagship, Eleven v3, reached general availability in February 2026 after an alpha in mid-2025. The lineup around it: v3 Conversational for real-time agents (about 280 ms), Multilingual v2 for steady long-form narration, and Flash v2.5 for low-latency bulk work (about 75 ms). ElevenLabs' CLI made v3… Read more →
Pros
- Very expressive delivery with audio tags
- 70+ languages on v3
- Voice cloning from $6/month
- Mature API and editor
Cons
- Free plan has no commercial licence
- Credits burn quickly on long projects
- 5,000-character cap per v3 request
Pricing: Free plan, paid from $6/mo · text-to-speech voice-cloning multilingual audio-tags elevenlabs
2
OpenAI's open-source speech recognition for transcripts and subtitles in 99 languages
★ 110k · MIT · updated 2026-08-31
8.2/10
What it is Whisper is OpenAI's open-source speech recognition model, released under the MIT licence. It transcribes and translates speech in 99 languages and detects which language is spoken. The newest open checkpoint is turbo (large-v3-turbo). It is an 809M-parameter pruned version of large-v3, about 8x faster with little loss in accuracy. OpenAI's paid API still offers whisper-1, but it… Read more →
Pros
- MIT licence, free for commercial use
- 99 languages
- Runs offline for privacy
- Huge ecosystem of ports and apps
Cons
- Can make up text during silence
- No built-in speaker separation
- Open line not updated since turbo
Pricing: Open source · speech-to-text open-source subtitles transcription openai
3
Google's prompt-directed TTS with 2,000+ voices, 100+ languages and consented cloning
8.0/10
What it is Gemini TTS is Google's speech generation line in the Gemini API. On 23 September 2026 Google launched two new models. Gemini 3.8 Flash TTS is built for top quality and directed performances. Gemini 3.8 Flash-Lite TTS is built for high-volume dubbing and agents. Both take natural-language direction for delivery and replace the earlier 3.1 Flash and 2.5… Read more →
Pros
- 2,000+ voices and 100+ languages
- Natural-language delivery direction
- SynthID watermark on all audio
- Consent-verified voice cloning
Cons
- Launched 23 Sept 2026, still very new
- Prices double from January 2027
- Cloning unavailable in the EU, UK, India and some US states
Pricing: Paid · text-to-speech multilingual voice-cloning synthid google
4
Cartesia's low-latency TTS for voice agents and narration, 44 languages
7.6/10
What it is Sonic is Cartesia's text-to-speech family. It is built on state space models (SSMs) rather than transformers, which makes it very fast. Sonic-3.6, generally available since 27 August 2026, is the current version. Cartesia also makes Ink-2, a streaming speech-to-text model. Cartesia pitches Sonic mainly at real-time voice agents, but it also works for narration and voiceovers. Key… Read more →
Pros
- Sub-90 ms latency
- Commercial licence from $5/month
- 44 languages
- Instant voice cloning on Pro
Cons
- Aimed at voice agents more than creators
- No long-form editing studio
- Few documented creator-app integrations
Pricing: Free plan, paid from $5/mo · text-to-speech low-latency voice-agents voice-cloning api
5
Tiny Apache-licensed TTS model with 54 voices in 8 languages that runs anywhere
★ 9.0k · Apache-2.0 · updated 2025-08-06
7.4/10
What it is Kokoro is an open-weight text-to-speech model with only 82 million parameters, published by the developer hexgrad. It is built on the StyleTTS 2 architecture. Version 1.0 (January 2025) is still the current release. It became popular because it sounds far better than its size suggests and runs quickly on a CPU, a laptop or even in a… Read more →
Pros
- Apache-2.0, free commercial use
- Runs fast on CPU and in the browser
- 54 voices
- Very cheap when hosted
Cons
- Only 8 languages
- No voice cloning
- Limited emotional range
- No release since January 2025
Pricing: Open source · open-weights text-to-speech lightweight offline apache-2
6
OpenAI's steerable text-to-speech API with 13 voices and prompt-based style control
7.4/10
What it is OpenAI's text-to-speech family is API-first. The current recommended model is gpt-4o-mini-tts. Its latest snapshot is dated December 2025. You tell it what to say and also how to say it: tone, emotion, pace and accent come from a plain-language instruction. The older tts-1 and tts-1-hd are still available. For live voice agents, OpenAI points to its separate… Read more →
Pros
- Plain-language style control
- Low per-token pricing
- Streaming output
- Clear consent and disclosure rules
Cons
- Only 13 voices, optimised for English
- API only, no editor
- Custom voices limited to eligible customers
Pricing: Paid · text-to-speech api openai steerable developers