Sonic-3.6 🧠 Model Freemium
Cartesia's low-latency TTS for voice agents and narration, 44 languages
- Maker
- Cartesia
- Latest release
- 2026-08
- Weights
- Closed
- Licence
- –
- Apps offering it
- 0
Where you can use Sonic-3.6
Apps in this directory that let you generate with Sonic-3.6, from their own documentation.
Good for
About Sonic-3.6
What it is
Sonic is Cartesia's text-to-speech family. It is built on state space models (SSMs) rather than transformers, which makes it very fast. Sonic-3.6, generally available since 27 August 2026, is the current version. Cartesia also makes Ink-2, a streaming speech-to-text model. Cartesia pitches Sonic mainly at real-time voice agents, but it also works for narration and voiceovers.
Key features
- Sub-90 ms latency, built for streaming
- 44 languages, newly including Odia and Urdu, plus expanded Hinglish
- Natural pacing and disfluencies (such as 'uhm') for conversational delivery
- Instant voice cloning (Pro) and professional cloning (Startup and up)
- Volume, speed and emotion controls through API parameters and SSML
Where you can use it
The Cartesia playground and API, plus cloud marketplaces such as AWS SageMaker JumpStart for Sonic 3. We found no creator app in our directory that documents Sonic by name. Most use comes through developer and voice-agent platforms.
Pricing and rights
Cartesia's Free plan gives 20,000 credits a month. Pro costs $5 a month for 100,000 credits and adds a commercial licence and instant voice cloning. Startup costs $49 and adds professional cloning. Scale costs $299. Commercial use starts on Pro.
Who it is for
Developers building voice agents, apps and games, and budget-minded creators who want a cheap commercial licence for narration.
Verdict
Sonic-3.6 is one of the fastest natural-sounding TTS models, and $5 a month is a low price for commercial use. It is built for agents first. There is no long-form editor like ElevenLabs Studio, and there are fewer ready-made creator integrations.
Pros
- Sub-90 ms latency
- Commercial licence from $5/month
- 44 languages
- Instant voice cloning on Pro
Cons
- Aimed at voice agents more than creators
- No long-form editing studio
- Few documented creator-app integrations
Similar AI models
All speech & voice models →Eleven v3 🧠 ModelFreemium
ElevenLabs' expressive text-to-speech model with audio tags and 70+ languages
Whisper large-v3-turbo 🧠 ModelOpen source
OpenAI's open-source speech recognition for transcripts and subtitles in 99 languages
Gemini 3.8 Flash TTS 🧠 ModelPaid
Google's prompt-directed TTS with 2,000+ voices, 100+ languages and consented cloning
Chatterbox Multilingual V3 🧠 ModelOpen source
Resemble AI's MIT-licensed TTS with zero-shot voice cloning and built-in watermarking
Kokoro-82M v1.0 🧠 ModelOpen source
Tiny Apache-licensed TTS model with 54 voices in 8 languages that runs anywhere
GPT-4o mini TTS 🧠 ModelPaid
OpenAI's steerable text-to-speech API with 13 voices and prompt-based style control