Veo 3.1 🧠 Model Paid
Google DeepMind's text- and image-to-video model with native sound, up to 4K
- Maker
- Google DeepMind
- Latest release
- 2025-10
- Weights
- Closed
- Licence
- –
- Apps offering it
- 30
Where you can use Veo 3.1
Apps in this directory that let you generate with Veo 3.1, from their own documentation. You can also call it directly through the maker's API.
Good for
About Veo 3.1
What it is
Veo is Google DeepMind's video generation family. The current version, Veo 3.1, turns a text prompt or reference images into short clips with sound generated in the same pass: dialogue, sound effects and ambience. A cheaper Lite tier reached the API in March 2026. Gemini Omni Flash, a separate conversational video model launched in May 2026, now sits next to Veo in the Gemini app and Flow.
Key features
- Native audio: speech, sound effects and ambience come out with the picture
- 1080p and 4K output, in 16:9 landscape or native 9:16 portrait
- 8-second base clips. Scene extension builds longer sequences
- Several reference images to keep a character or product consistent
- Every clip carries an invisible SynthID watermark
Where you can use it
Gemini app, Google Flow, Google Vids, AI Studio and the Gemini API. Outside Google, Veo 3.1 is a documented option in Adobe Firefly (as a partner model in Firefly Boards), Higgsfield, Krea, Freepik, Leonardo.Ai, Luma's app and invideo AI. OpenArt lists Veo 3, and Canva's 'Create a Video Clip' runs on Veo 3 (paid plans).
Pricing and rights
Gemini API prices per generated second, audio included (September 2026): Veo 3.1 Standard $0.40 at 720p/1080p and $0.60 at 4K. Fast is $0.10-0.30. Lite is $0.05 at 720p and $0.08 at 1080p. In the Gemini app and Flow it comes with Google AI subscriptions (prices not verified here). Partner apps charge their own credits.
Who it is for
Filmmakers, ad makers and social teams who want realistic short shots with sound, and developers who want a per-second API.
Verdict
Veo 3.1 gives strong realism and some of the best built-in audio, and it is available in more apps than most video models. Base clips are short. Google's own docs say short lines of dialogue can still sound off. Google's newer effort is going into Gemini Omni, so expect the line to change.
Pros
- Sound (dialogue, SFX, ambience) generated in the same pass
- 4K and native vertical output
- Three API tiers from $0.05 to $0.60 per second
- Available in many creator apps
- SynthID watermark on every clip
Cons
- 8-second base clips
- Standard tier is expensive per second
- Google says short dialogue lines are still hit-and-miss
- Google's newer video work is going into Gemini Omni
Similar AI models
All video models →Kling 3.0 🧠 ModelFreemium
Kuaishou's video model with native multilingual audio, lip sync and 15-second clips
Seedance 2.5 🧠 ModelPaid
ByteDance's video model: 30-second clips with audio, many references and local edits
MiniMax H3 (Hailuo 3.0) 🧠 ModelFreemium
MiniMax's open-weight video model: native 2K, stereo sound, 15-second clips
Runway Gen-4.5 🧠 ModelFreemium
Runway's own video model, now with native audio and one-minute multi-shot scenes
Wan 3.0 🧠 ModelOpen source
Alibaba's video family: open-weight Wan 2.x and the 30-second Wan 3.0 API
LTX-2.5 🧠 ModelOpen source
Lightricks' open-weight video model with synchronised audio and 4K output
More from Google DeepMind
Nano Banana 2 🧠 ModelFreemium
Google's Gemini image models (Nano Banana 2, Pro and 2 Lite) for generation and editing up to 4K
Gemini 3.1 Pro 🧠 ModelFreemium
Google's Gemini models (3.1 Pro, 3.8 Flash) in the Gemini app, NotebookLM and the Gemini API
Gemini 3.8 Flash TTS 🧠 ModelPaid
Google's prompt-directed TTS with 2,000+ voices, 100+ languages and consented cloning
Lyria 3.5 🧠 ModelPaid
Google DeepMind's music model for full songs with vocals, lyrics and SynthID