Veo 3 is the video generation model of Google DeepMind, presented in May 2025: you describe the scene and it generates an eight-second clip with synchronized audio inside, dialogues included. You use it in the Gemini app, where it is also free with daily limits, in Flow (Google's tool dedicated to AI video) and through the APIs. It is the engine that follows the prompt with the most precision among the five big ones: you ask for a lateral tracking shot with rain against the light and you get that scene, not a cousin of it. Here you find how to start, what it really costs, the differences between Veo 2, 3 and 3.1 and the limits nobody writes.
In my weekly work, Veo 3 clips are the shots that come out right on the first take most often. Not because it is magic: because it is the engine most faithful to the sentence you write. The others gift you beautiful variants of the scene you asked for; Veo tends to make you the scene you asked for. If you make videos for clients, where the brief is not negotiated, this difference is everything.
This guide is the zoom on Veo 3: the complete map of the five engines, with who wins where, is in the main guide create videos with AI.
The other models give you a scene similar to the one you asked for. Veo gives you that scene: it is the reason it ends up in my projects when the brief is written.
What Veo 3 is (and where it comes from)
Veo 3 is a text-to-video model by Google DeepMind: you give it a description, or a starting image, and it returns an eight-second clip with camera, movement and audio generated together. The audio is the turning point that set it apart from the previous generation: rain that has its noise, footsteps that follow the step, characters speaking with lips in sync.
The story in three moves helps you get oriented, because on Veo you still find everything and it gets confusing:
- Veo 2 is the previous generation: clips without audio, lower resolution, less camera control. You can still find it as a "fast" option: fine for trying, a different craft for producing.
- Veo 3, presented in May 2025, brought synchronized audio and the quality that stands comparison with cinema photography. It is the moment AI video stopped being mute.
- Veo 3.1, the current version, added continuity: the possibility of starting again from the last frame to extend the scene, and references to keep the same character and the same objects across different clips.
When someone searches "Veo 2" today, they almost always want to know whether the jump to 3 is worth it: it is, and audio is the number one reason.
Veo 3 for free: where you really try it
The question I get most often is whether Veo 3 can be used for free. Yes, with precise limits:
- Inside Gemini, free plan: you can generate a few videos per day with the Fast version of Veo, at reduced resolution and with watermark. The exact number is not officially declared and varies: in my weeks it sits around three per day. Enough to understand whether the tool is for you.
- Inside Flow: Google Labs' video tool starts with trial credits, which run out fast. Flow is beautiful for building scenes, not for producing for free.
- The serious free routes, all together, are in the dedicated guide: free AI video generators, where Veo is one entry in the complete picture.
My production honesty: Veo's free tier is for learning to write prompts, not for doing work. If you have a project with a deadline, the daily limits stop you halfway, and stopping halfway is the worst way to get to know a tool.
How to use it: the first clip in five steps
The path I would give to anyone starting today, tried on Gemini's free plan:
- Open Gemini (app or web) and ask for a video in natural language: "generate a video of...". No commands needed: Veo reads the sentence.
- Write the scene with the five elements: subject, action, camera, light, audio, in this order and in one sentence. Below I leave you an example I really use.
- Watch what comes out and iterate on the wrong part. The camera wanders? Re-declare it. The light is flat? Name the light source. Do not regenerate blind: every attempt has a cost, even when it is free, because time is the first budget.
- When the clip holds, move it into Flow if the scene needs continuing: there you can extend from the last frame and keep the ingredients, meaning the character and object references that make subsequent clips consistent.
- Edit outside. Eight seconds are a shot, not a video: cuts, rhythm and final audio are decided in editing. No model gives you that.
On Veo writing pays more than elsewhere, because the model executes. The structure I use, with an example that works for a first try:
A woman in a beige trench coat walks alone through a night market, camera tracks laterally at walking pace, warm string lights and steam from food stalls, ambient chatter and her footsteps.
Subject (who is in it), action (one only), camera (declared), light (the source named), audio (the sound asked for). Veo takes every piece seriously: if you declare the camera, the camera does that; if you name the light, the light arrives from there. The complete rules, including the traps like "no people in the background" that makes people appear in the background, are in the prompt section of the main guide.
Veo 2, Veo 3, Veo 3.1: what changes in practice
| Veo 2 | Veo 3 | Veo 3.1 | |
|---|---|---|---|
| Audio | no, mute | yes, synchronized, dialogues included | yes, improved |
| Clip duration | short clips | 8 seconds | 8 seconds, extendable from the last frame |
| Consistency | limited | good on the single clip | character and object references across clips |
| Where you find it | as a fast option in some tools | Gemini, Flow, API | Gemini, Flow, API: it is the current version |
| When to choose it | never, if you have 3 within reach | written scene that must come out like that | projects with a recurring character |
How much Veo 3 costs
Three roads, three pricing logics. Figures verified in August 2026: in this sector they move every month, so take them as a snapshot and recheck the official pages before deciding.
| Road | What you pay | When it makes sense |
|---|---|---|
| Gemini free | 0: a few generations per day with Veo Fast, reduced resolution, watermark | Learning to write, trying scenes |
| Google AI plans | the monthly subscription unlocks more generations, higher quality and Flow with more credits | Producing every week |
| Gemini API | per clip: about $0.40 in 720-1080p, $0.60 in 4K | Building products and automatic flows |
Where Veo stumbles (and how to get around it)
I use Veo every week and I recommend it, so I owe you the real cons:
- Eight seconds per clip. It is the format of the shot, not of the video: you extend it in Flow or you edit. Whoever arrives from "make me a one minute video" has to recalibrate the expectation: one minute is eight clips, meaning eight written scenes.
- The safety guards. Famous faces, real people, sensitive content: Veo filters hard. In work for brands it is almost always an advantage, but know it before writing the prompt with your ambassador inside.
- Text inside the video. Like everyone: writing comes out crooked. Texts get composed in editing, always.
- Lip sync in Italian. English is followed well, Italian less: for Italian dialogues the voice goes in post.
- The invisible watermark. Veo videos carry SynthID inside, Google's invisible signature declaring the AI origin. You cannot see it, you cannot remove it, and social platforms recognize it: the provenance has to be declared, always.
When I choose Veo, and when not
In my way of splitting scenes between engines, Veo takes the shots where precision pays: the scene is already written, the client has approved the reference, it must come out like that. When the scene asks for something else, I change engine:
- people walking, gestures, body physics: Kling moves better;
- scenes with several actions in a row, long clips: Seedance remembers them all;
- whole projects with camera, consistency and extensions to manage together: I go through the directing platform, Higgsfield, where the engines live with the tools around them.
And if you need the video for your brand and prefer to have someone who does it every day: that is my craft. Write to me and we start from an idea.
The guides of the video path
- Create videos with AI, the map of all the tools (the pillar of this guide);
- Higgsfield, the directing platform where the engines work;
- Seedance, the engine of long scenes;
- Kling AI, the realist, and the best free school;
- Sora 2, the storyboard pioneer, today with competition on its heels;
- Free AI video generators, all the ways at zero cost.
Sources
- Google DeepMind, official model page: deepmind.google/models/veo
- Gemini, where Veo lives for free too: gemini.google.com
- Flow, the Google Labs video tool: labs.google/flow
- Gemini API, video generation pricing: ai.google.dev
- Free plan limits and model behavior: verified in the app in August 2026, same check as the main guide
For updates and news, there is my Telegram channel
I publish the latest news on AI video tools there first.
Leave a comment
Comments are reviewed and approved before they appear.