GPT-4o mini TTS 🧠 Model Paid
OpenAI's steerable text-to-speech API with 13 voices and prompt-based style control
- Maker
- OpenAI
- Latest release
- 2025-12
- Weights
- Closed
- Licence
- –
- Apps offering it
- 1
Where you can use GPT-4o mini TTS
Apps in this directory that let you generate with GPT-4o mini TTS, from their own documentation. You can also call it directly through the maker's API.
Good for
About GPT-4o mini TTS
What it is
OpenAI's text-to-speech family is API-first. The current recommended model is gpt-4o-mini-tts. Its latest snapshot is dated December 2025. You tell it what to say and also how to say it: tone, emotion, pace and accent come from a plain-language instruction. The older tts-1 and tts-1-hd are still available. For live voice agents, OpenAI points to its separate Realtime models (gpt-realtime).
Key features
- 13 built-in voices. OpenAI recommends marin and cedar for best quality
- Style instructions in plain language, with no SSML needed
- Follows Whisper's language coverage, but voices are optimised for English
- Streaming output for low-latency playback
- Custom voices for eligible customers, with a recorded consent statement from the speaker
Where you can use it
Mostly developers and automation tools calling the OpenAI API directly. We found no creator app in our directory that documents gpt-4o-mini-tts by name. ChatGPT's own voice features run on OpenAI's realtime voice stack instead.
Pricing and rights
Usage-based pricing only: $0.60 per million input text tokens and $12 per million audio output tokens. There is no subscription. OpenAI's usage policies require you to tell listeners clearly that the voice is AI-generated. Custom voices need the speaker's consent recording.
Who it is for
Developers building narration, accessibility or app voice features who already use the OpenAI API and want cheap, steerable voices.
Verdict
GPT-4o mini TTS is inexpensive and easy to steer, and it is a sensible default inside OpenAI-based pipelines. The voice list is small, non-English delivery trails ElevenLabs and Gemini, and there is no consumer editor. It is a building block, not a creator app.
Pros
- Plain-language style control
- Low per-token pricing
- Streaming output
- Clear consent and disclosure rules
Cons
- Only 13 voices, optimised for English
- API only, no editor
- Custom voices limited to eligible customers
Similar AI models
All speech & voice models →Eleven v3 🧠 ModelFreemium
ElevenLabs' expressive text-to-speech model with audio tags and 70+ languages
Whisper large-v3-turbo 🧠 ModelOpen source
OpenAI's open-source speech recognition for transcripts and subtitles in 99 languages
Gemini 3.8 Flash TTS 🧠 ModelPaid
Google's prompt-directed TTS with 2,000+ voices, 100+ languages and consented cloning
Chatterbox Multilingual V3 🧠 ModelOpen source
Resemble AI's MIT-licensed TTS with zero-shot voice cloning and built-in watermarking
Sonic-3.6 🧠 ModelFreemium
Cartesia's low-latency TTS for voice agents and narration, 44 languages
Kokoro-82M v1.0 🧠 ModelOpen source
Tiny Apache-licensed TTS model with 54 voices in 8 languages that runs anywhere
More from OpenAI
ChatGPT Freemium
OpenAI's general assistant for writing, research, images and voice chat
GPT Image 2.5 🧠 ModelFreemium
OpenAI's image model behind ChatGPT Images, with 4K output and up to 16 reference images
GPT-6 Astra 🧠 ModelFreemium
OpenAI's GPT-6 family (Astra, Sol, Luna) behind ChatGPT and the OpenAI API
Whisper large-v3-turbo 🧠 ModelOpen source
OpenAI's open-source speech recognition for transcripts and subtitles in 99 languages