Try “Veo”, “voiceover”, “thumbnail” or “Descript” · Esc to close

Kokoro-82M v1.0 🧠 Model Open source

Tiny Apache-licensed TTS model with 54 voices in 8 languages that runs anywhere

Speech & voice models · Open source ★ 9.0k · Apache-2.0 · updated 2025-08-06

7.4editor score
Visit Kokoro-82M v1.0 ↗
GitHub stars
9.0k
Stars this week
–
Forks
981
Licence
Apache-2.0
Last push
2025-08-06
Maintainer
hexgrad

Where you can use Kokoro-82M v1.0

Apps in this directory that let you generate with Kokoro-82M v1.0, from their own documentation.

No app mapped yet.

See every app × model on the model map →

Good for

About Kokoro-82M v1.0

What it is

Kokoro is an open-weight text-to-speech model with only 82 million parameters, published by the developer hexgrad. It is built on the StyleTTS 2 architecture. Version 1.0 (January 2025) is still the current release. It became popular because it sounds far better than its size suggests and runs quickly on a CPU, a laptop or even in a browser.

Key features

  • 82M parameters, 24 kHz output, fast on CPU
  • 54 preset voices across 8 languages
  • Apache-2.0 weights, so commercial use is allowed
  • Python package plus community ports (ONNX, JavaScript, browser)
  • Hosted by several inference providers at under $1 per million characters

Where you can use it

Download from Hugging Face (hexgrad/Kokoro-82M) or GitHub and run it with the kokoro Python package. You can also use it through hosted inference providers such as fal, Replicate and DeepInfra. We found no creator app in our directory that documents Kokoro by name.

Pricing and rights

The weights are free under Apache-2.0, including commercial use. You pay only for your own compute or a host's per-character fee. Kokoro does no voice cloning, so you cannot copy a real person's voice with it. The trade-off is that you are limited to the preset voices.

Who it is for

Developers, hobbyists and creators on a budget who want free, private, offline narration for videos, audiobooks or apps.

Verdict

Kokoro is a strong open TTS pick for size and licence. It runs almost anywhere and costs nothing. Its expressiveness is well behind ElevenLabs v3, it covers only 8 languages, it has no voice cloning, and it has not been updated since early 2025.

open-weights text-to-speech lightweight offline apache-2

Pros

  • Apache-2.0, free commercial use
  • Runs fast on CPU and in the browser
  • 54 voices
  • Very cheap when hosted

Cons

  • Only 8 languages
  • No voice cloning
  • Limited emotional range
  • No release since January 2025

Eleven v3 🧠 ModelFreemium

ElevenLabs' expressive text-to-speech model with audio tags and 70+ languages

ElevenLabs · 2026-02

8.8 Visit ↗

Whisper large-v3-turbo 🧠 ModelOpen source

OpenAI's open-source speech recognition for transcripts and subtitles in 99 languages

OpenAI · 2024-10 · open weights

Gemini 3.8 Flash TTS 🧠 ModelPaid

Google's prompt-directed TTS with 2,000+ voices, 100+ languages and consented cloning

Google DeepMind · 2026-09

8.0 Visit ↗

Chatterbox Multilingual V3 🧠 ModelOpen source

Resemble AI's MIT-licensed TTS with zero-shot voice cloning and built-in watermarking

Resemble AI · 2026-06 · open weights

7.6 Visit ↗

Sonic-3.6 🧠 ModelFreemium

Cartesia's low-latency TTS for voice agents and narration, 44 languages

Cartesia · 2026-08

7.6 Visit ↗

GPT-4o mini TTS 🧠 ModelPaid

OpenAI's steerable text-to-speech API with 13 voices and prompt-based style control

OpenAI · 2025-12

7.4 Visit ↗

Popular searches