Whisper large-v3-turbo 🧠 Model Open source
OpenAI's open-source speech recognition for transcripts and subtitles in 99 languages
- GitHub stars
- 110k
- Stars this week
- –
- Forks
- 13k
- Licence
- MIT
- Last push
- 2026-08-31
- Maintainer
- OpenAI
Where you can use Whisper large-v3-turbo
Apps in this directory that let you generate with Whisper large-v3-turbo, from their own documentation.
Good for
About Whisper large-v3-turbo
What it is
Whisper is OpenAI's open-source speech recognition model, released under the MIT licence. It transcribes and translates speech in 99 languages and detects which language is spoken. The newest open checkpoint is turbo (large-v3-turbo). It is an 809M-parameter pruned version of large-v3, about 8x faster with little loss in accuracy. OpenAI's paid API still offers whisper-1, but it now points users to newer models (gpt-4o-transcribe, gpt-realtime-whisper).
Key features
- 99-language transcription, translation to English and language detection
- Six sizes from tiny (39M) to large (1.55B). Turbo needs about 6 GB of VRAM
- Timestamps for subtitle (SRT/VTT) generation
- Runs fully offline on your own machine
- The base for many community ports (whisper.cpp, faster-whisper) and desktop apps
Where you can use it
Install it with pip and run it locally, or call whisper-1 through the OpenAI API. Many transcription and captioning tools build on Whisper, but few name it on their pricing or product pages, so we do not link specific apps here.
Pricing and rights
The open-source model is free under the MIT licence, including commercial use. You pay only for your own hardware. The hosted whisper-1 API costs $0.006 per minute. Accuracy drops on noisy audio and rare languages, and Whisper can produce made-up text during silence. Proofread before you publish subtitles.
Who it is for
Podcasters, video editors and developers who want free, private, offline transcripts and subtitles, or a base for their own captioning tool.
Verdict
Whisper is still the default open speech-to-text model: free, multilingual and easy to run. It has had no major new open checkpoint since turbo. It does not separate speakers on its own, and hallucinated text on silent passages means you still need a human pass.
Pros
- MIT licence, free for commercial use
- 99 languages
- Runs offline for privacy
- Huge ecosystem of ports and apps
Cons
- Can make up text during silence
- No built-in speaker separation
- Open line not updated since turbo
Similar AI models
All speech & voice models →Eleven v3 🧠 ModelFreemium
ElevenLabs' expressive text-to-speech model with audio tags and 70+ languages
Gemini 3.8 Flash TTS 🧠 ModelPaid
Google's prompt-directed TTS with 2,000+ voices, 100+ languages and consented cloning
Chatterbox Multilingual V3 🧠 ModelOpen source
Resemble AI's MIT-licensed TTS with zero-shot voice cloning and built-in watermarking
Sonic-3.6 🧠 ModelFreemium
Cartesia's low-latency TTS for voice agents and narration, 44 languages
Kokoro-82M v1.0 🧠 ModelOpen source
Tiny Apache-licensed TTS model with 54 voices in 8 languages that runs anywhere
GPT-4o mini TTS 🧠 ModelPaid
OpenAI's steerable text-to-speech API with 13 voices and prompt-based style control
More from OpenAI
ChatGPT Freemium
OpenAI's general assistant for writing, research, images and voice chat
GPT Image 2.5 🧠 ModelFreemium
OpenAI's image model behind ChatGPT Images, with 4K output and up to 16 reference images
GPT-6 Astra 🧠 ModelFreemium
OpenAI's GPT-6 family (Astra, Sol, Luna) behind ChatGPT and the OpenAI API
GPT-4o mini TTS 🧠 ModelPaid
OpenAI's steerable text-to-speech API with 13 voices and prompt-based style control