AI Models for Content (2026): 41 Video, Image, Voice & Text Models and Where to Use Them
The generative models behind the apps: language, image, video, voice and music models, and the apps where you can use each one.
The same model is often sold by several apps at different prices and limits. Each model page lists where you can use it; the model map compares them all.
Claude Fable 5.1 🧠 ModelFreemium
Anthropic's Claude models (Fable 5.1, Opus 5.5, Sonnet 5, Haiku 4.5) with 1M-token context
GPT Image 2.5 🧠 ModelFreemium
OpenAI's image model behind ChatGPT Images, with 4K output and up to 16 reference images
GPT-6 Astra 🧠 ModelFreemium
OpenAI's GPT-6 family (Astra, Sol, Luna) behind ChatGPT and the OpenAI API
Nano Banana 2 🧠 ModelFreemium
Google's Gemini image models (Nano Banana 2, Pro and 2 Lite) for generation and editing up to 4K
Eleven v3 🧠 ModelFreemium
ElevenLabs' expressive text-to-speech model with audio tags and 70+ languages
Veo 3.1 🧠 ModelPaid
Google DeepMind's text- and image-to-video model with native sound, up to 4K
Gemini 3.1 Pro 🧠 ModelFreemium
Google's Gemini models (3.1 Pro, 3.8 Flash) in the Gemini app, NotebookLM and the Gemini API
FLUX.2 🧠 ModelOpen source
Black Forest Labs' FLUX.2 family: hosted [max]/[pro]/[flex] plus open-weight [dev] and [klein]
Kling 3.0 🧠 ModelFreemium
Kuaishou's video model with native multilingual audio, lip sync and 15-second clips
Seedance 2.5 🧠 ModelPaid
ByteDance's video model: 30-second clips with audio, many references and local edits
Suno v6 🧠 ModelFreemium
Suno's song model: full tracks with vocals and lyrics from a prompt
Midjourney V8.2 🧠 ModelPaid
Midjourney's own image model, known for aesthetics and personalization, now in V8.2
Seedream 5.0 Pro 🧠 ModelFreemium
ByteDance's Seedream image models, used in CapCut, Dreamina and many creative apps
Whisper large-v3-turbo 🧠 ModelOpen source
OpenAI's open-source speech recognition for transcripts and subtitles in 99 languages
MiniMax H3 (Hailuo 3.0) 🧠 ModelFreemium
MiniMax's open-weight video model: native 2K, stereo sound, 15-second clips
Runway Gen-4.5 🧠 ModelFreemium
Runway's own video model, now with native audio and one-minute multi-shot scenes
Wan 3.0 🧠 ModelOpen source
Alibaba's video family: open-weight Wan 2.x and the 30-second Wan 3.0 API
DeepSeek V4 Pro 🧠 ModelOpen source
DeepSeek's MIT-licensed V4 models with 1M-token context and very low API prices
Gemini 3.8 Flash TTS 🧠 ModelPaid
Google's prompt-directed TTS with 2,000+ voices, 100+ languages and consented cloning
Ideogram 4.0 🧠 ModelOpen source
Ideogram's text-rendering image model with JSON prompting, now also as non-commercial open weights
Lyria 3.5 🧠 ModelPaid
Google DeepMind's music model for full songs with vocals, lyrics and SynthID
Eleven Music v2.5 🧠 ModelFreemium
ElevenLabs' music model: songs up to 5 minutes with vocals, cleared for most uses
LTX-2.5 🧠 ModelOpen source
Lightricks' open-weight video model with synchronised audio and 4K output
Qwen-Image 3.0 🧠 ModelOpen source
Alibaba's Qwen image models, known for dense multilingual text rendering, with open-weight versions
Qwen3.8 🧠 ModelOpen source
Alibaba's Qwen3.8 family, from Apache-2.0 open weights (27B to 2.4T) to the hosted Qwen3.8-Max
Firefly Image 5 🧠 ModelFreemium
Adobe's own image model, trained on licensed content, with native 4MP output and Content Credentials
Chatterbox Multilingual V3 🧠 ModelOpen source
Resemble AI's MIT-licensed TTS with zero-shot voice cloning and built-in watermarking
Luma Ray3.2 🧠 ModelPaid
Luma's HDR video model with up to 16 keyframes in a 20-second 1080p clip
Sonic-3.6 🧠 ModelFreemium
Cartesia's low-latency TTS for voice agents and narration, 44 languages
Grok 4.7 🧠 ModelFreemium
xAI's Grok models, built into Grok and X, with a 500K-token context on Grok 4.7
Stable Audio 3.0 🧠 ModelOpen source
Stability AI's licensed-data audio models up to six minutes, with open-weight tiers
ACE-Step 1.5 🧠 ModelOpen source
MIT-licensed open music model that makes full songs with vocals locally in seconds
Kokoro-82M v1.0 🧠 ModelOpen source
Tiny Apache-licensed TTS model with 54 voices in 8 languages that runs anywhere
GPT-4o mini TTS 🧠 ModelPaid
OpenAI's steerable text-to-speech API with 13 voices and prompt-based style control
Mistral Medium 3.5 🧠 ModelOpen source
Mistral AI's European models: open-weight Medium 3.5, Large 3 and Small 4, used in Mistral Vibe
Vidu Q3 🧠 ModelFreemium
ShengShu's video model with native audio and 16-second multi-shot scenes
PixVerse V6 🧠 ModelFreemium
PixVerse's video model: 1-15 s clips at 1080p with native audio and multi-shot
HunyuanVideo 1.5 🧠 ModelOpen source
Tencent's 8.3B open-weight video model that runs on a single consumer GPU
Meta Muse Spark 1.3 🧠 ModelFreemium
Meta's model line: closed Muse Spark 1.3, open Muse Glimmer 30B and the older open Llama 4
Stable Diffusion 3.5 🧠 ModelOpen source
Stability AI's open-weight image models (SD 3.5 Large, Turbo, Medium) for local and custom workflows
MusicGen (AudioCraft) 🧠 ModelOpen source
Meta's open research model for short instrumental music from text or a melody