Try “Veo”, “voiceover”, “thumbnail” or “Descript” · Esc to close

Sonic-3.6 🧠 Model Freemium

Cartesia's low-latency TTS for voice agents and narration, 44 languages

Speech & voice models · Free plan, paid from $5/mo

7.6editor score
Visit Sonic-3.6 ↗
Maker
Cartesia
Latest release
2026-08
Weights
Closed
Licence
–
Apps offering it
0

Where you can use Sonic-3.6

Apps in this directory that let you generate with Sonic-3.6, from their own documentation.

No app mapped yet.

See every app × model on the model map →

Good for

About Sonic-3.6

What it is

Sonic is Cartesia's text-to-speech family. It is built on state space models (SSMs) rather than transformers, which makes it very fast. Sonic-3.6, generally available since 27 August 2026, is the current version. Cartesia also makes Ink-2, a streaming speech-to-text model. Cartesia pitches Sonic mainly at real-time voice agents, but it also works for narration and voiceovers.

Key features

  • Sub-90 ms latency, built for streaming
  • 44 languages, newly including Odia and Urdu, plus expanded Hinglish
  • Natural pacing and disfluencies (such as 'uhm') for conversational delivery
  • Instant voice cloning (Pro) and professional cloning (Startup and up)
  • Volume, speed and emotion controls through API parameters and SSML

Where you can use it

The Cartesia playground and API, plus cloud marketplaces such as AWS SageMaker JumpStart for Sonic 3. We found no creator app in our directory that documents Sonic by name. Most use comes through developer and voice-agent platforms.

Pricing and rights

Cartesia's Free plan gives 20,000 credits a month. Pro costs $5 a month for 100,000 credits and adds a commercial licence and instant voice cloning. Startup costs $49 and adds professional cloning. Scale costs $299. Commercial use starts on Pro.

Who it is for

Developers building voice agents, apps and games, and budget-minded creators who want a cheap commercial licence for narration.

Verdict

Sonic-3.6 is one of the fastest natural-sounding TTS models, and $5 a month is a low price for commercial use. It is built for agents first. There is no long-form editor like ElevenLabs Studio, and there are fewer ready-made creator integrations.

text-to-speech low-latency voice-agents voice-cloning api

Pros

  • Sub-90 ms latency
  • Commercial licence from $5/month
  • 44 languages
  • Instant voice cloning on Pro

Cons

  • Aimed at voice agents more than creators
  • No long-form editing studio
  • Few documented creator-app integrations

Eleven v3 🧠 ModelFreemium

ElevenLabs' expressive text-to-speech model with audio tags and 70+ languages

ElevenLabs · 2026-02

8.8 Visit ↗

Whisper large-v3-turbo 🧠 ModelOpen source

OpenAI's open-source speech recognition for transcripts and subtitles in 99 languages

OpenAI · 2024-10 · open weights

Gemini 3.8 Flash TTS 🧠 ModelPaid

Google's prompt-directed TTS with 2,000+ voices, 100+ languages and consented cloning

Google DeepMind · 2026-09

8.0 Visit ↗

Chatterbox Multilingual V3 🧠 ModelOpen source

Resemble AI's MIT-licensed TTS with zero-shot voice cloning and built-in watermarking

Resemble AI · 2026-06 · open weights

7.6 Visit ↗

Kokoro-82M v1.0 🧠 ModelOpen source

Tiny Apache-licensed TTS model with 54 voices in 8 languages that runs anywhere

hexgrad · 2025-01 · open weights

7.4 Visit ↗

GPT-4o mini TTS 🧠 ModelPaid

OpenAI's steerable text-to-speech API with 13 voices and prompt-based style control

OpenAI · 2025-12

7.4 Visit ↗

Popular searches