Chatterbox Multilingual V3 🧠 Model Open source
Resemble AI's MIT-licensed TTS with zero-shot voice cloning and built-in watermarking
- GitHub stars
- 27k
- Stars this week
- –
- Forks
- 3.6k
- Licence
- MIT
- Last push
- 2026-07-21
- Maintainer
- Resemble AI
Where you can use Chatterbox Multilingual V3
Apps in this directory that let you generate with Chatterbox Multilingual V3, from their own documentation.
Good for
About Chatterbox Multilingual V3
What it is
Chatterbox is Resemble AI's open-source text-to-speech family, released under the MIT licence. The current general-purpose model, Chatterbox Multilingual V3 (June 2026), is a 0.5B-parameter model that clones a voice from a short sample in 23+ languages. Next to it sit Chatterbox-Turbo (350M, English, low-latency, with paralinguistic tags) and Chatterbox-Nano (110M, English, runs on a CPU).
Key features
- Zero-shot voice cloning from a short reference clip
- 23+ languages, with dedicated packs for Chinese, Spanish, Portuguese and Hindi
- Turbo supports tags like [laugh] and [cough] and targets about 75 ms latency
- Emotion exaggeration and guidance controls
- Every output carries Resemble's PerTh watermark for detecting synthetic audio
Where you can use it
Install it with pip and run it locally, or use the Hugging Face demo. Resemble AI also offers Chatterbox through its own hosted platform, and NVIDIA offers it as a NIM microservice.
Pricing and rights
Free under MIT for commercial and personal use. You pay only for your own compute, or Resemble AI's hosted rates if you use their service. The built-in PerTh watermark makes outputs traceable. Because cloning is zero-shot, only clone voices you own or have written consent to use.
Who it is for
Developers and creators who want free, self-hosted voice cloning for narration, dubbing or characters, with provenance built in.
Verdict
Chatterbox is one of the strongest permissively licensed cloning models, and the watermark is a responsible default. Running it well locally still takes a GPU and some setup. Multilingual quality varies by language.
Pros
- MIT licence
- Zero-shot cloning in 23+ languages
- PerTh watermark on every output
- Turbo and Nano variants for speed and CPU
Cons
- Needs a GPU for comfortable speed (except Nano)
- Quality varies by language
- Easy cloning needs responsible use
Similar AI models
All speech & voice models →Eleven v3 🧠 ModelFreemium
ElevenLabs' expressive text-to-speech model with audio tags and 70+ languages
Whisper large-v3-turbo 🧠 ModelOpen source
OpenAI's open-source speech recognition for transcripts and subtitles in 99 languages
Gemini 3.8 Flash TTS 🧠 ModelPaid
Google's prompt-directed TTS with 2,000+ voices, 100+ languages and consented cloning
Sonic-3.6 🧠 ModelFreemium
Cartesia's low-latency TTS for voice agents and narration, 44 languages
Kokoro-82M v1.0 🧠 ModelOpen source
Tiny Apache-licensed TTS model with 54 voices in 8 languages that runs anywhere
GPT-4o mini TTS 🧠 ModelPaid
OpenAI's steerable text-to-speech API with 13 voices and prompt-based style control
More from Resemble AI
Resemble AI Paid
Consent-first voice cloning and open Chatterbox TTS, now led by deepfake detection