Try “Veo”, “voiceover”, “thumbnail” or “Descript” · Esc to close

Chatterbox Multilingual V3 🧠 Model Open source

Resemble AI's MIT-licensed TTS with zero-shot voice cloning and built-in watermarking

Speech & voice models · Open source ★ 27k · MIT · updated 2026-07-21

7.6editor score
Visit Chatterbox Multilingual V3 ↗
GitHub stars
27k
Stars this week
–
Forks
3.6k
Licence
MIT
Last push
2026-07-21
Maintainer
Resemble AI

Where you can use Chatterbox Multilingual V3

Apps in this directory that let you generate with Chatterbox Multilingual V3, from their own documentation.

See every app × model on the model map →

Good for

About Chatterbox Multilingual V3

What it is

Chatterbox is Resemble AI's open-source text-to-speech family, released under the MIT licence. The current general-purpose model, Chatterbox Multilingual V3 (June 2026), is a 0.5B-parameter model that clones a voice from a short sample in 23+ languages. Next to it sit Chatterbox-Turbo (350M, English, low-latency, with paralinguistic tags) and Chatterbox-Nano (110M, English, runs on a CPU).

Key features

  • Zero-shot voice cloning from a short reference clip
  • 23+ languages, with dedicated packs for Chinese, Spanish, Portuguese and Hindi
  • Turbo supports tags like [laugh] and [cough] and targets about 75 ms latency
  • Emotion exaggeration and guidance controls
  • Every output carries Resemble's PerTh watermark for detecting synthetic audio

Where you can use it

Install it with pip and run it locally, or use the Hugging Face demo. Resemble AI also offers Chatterbox through its own hosted platform, and NVIDIA offers it as a NIM microservice.

Pricing and rights

Free under MIT for commercial and personal use. You pay only for your own compute, or Resemble AI's hosted rates if you use their service. The built-in PerTh watermark makes outputs traceable. Because cloning is zero-shot, only clone voices you own or have written consent to use.

Who it is for

Developers and creators who want free, self-hosted voice cloning for narration, dubbing or characters, with provenance built in.

Verdict

Chatterbox is one of the strongest permissively licensed cloning models, and the watermark is a responsible default. Running it well locally still takes a GPU and some setup. Multilingual quality varies by language.

open-source voice-cloning text-to-speech watermarking multilingual

Pros

  • MIT licence
  • Zero-shot cloning in 23+ languages
  • PerTh watermark on every output
  • Turbo and Nano variants for speed and CPU

Cons

  • Needs a GPU for comfortable speed (except Nano)
  • Quality varies by language
  • Easy cloning needs responsible use

Eleven v3 🧠 ModelFreemium

ElevenLabs' expressive text-to-speech model with audio tags and 70+ languages

ElevenLabs · 2026-02

8.8 Visit ↗

Whisper large-v3-turbo 🧠 ModelOpen source

OpenAI's open-source speech recognition for transcripts and subtitles in 99 languages

OpenAI · 2024-10 · open weights

Gemini 3.8 Flash TTS 🧠 ModelPaid

Google's prompt-directed TTS with 2,000+ voices, 100+ languages and consented cloning

Google DeepMind · 2026-09

8.0 Visit ↗

Sonic-3.6 🧠 ModelFreemium

Cartesia's low-latency TTS for voice agents and narration, 44 languages

Cartesia · 2026-08

7.6 Visit ↗

Kokoro-82M v1.0 🧠 ModelOpen source

Tiny Apache-licensed TTS model with 54 voices in 8 languages that runs anywhere

hexgrad · 2025-01 · open weights

7.4 Visit ↗

GPT-4o mini TTS 🧠 ModelPaid

OpenAI's steerable text-to-speech API with 13 voices and prompt-based style control

OpenAI · 2025-12

7.4 Visit ↗

More from Resemble AI

Resemble AI Paid

Consent-first voice cloning and open Chatterbox TTS, now led by deepfake detection

★ 27k

6.8 Visit ↗

Popular searches