Skip to content
erika.taranto
ITENDE中文RU
Blog Guides

Veo 3: what it is, how to use it and what it costs

Google's Veo 3 explained by someone who produces with it: what it is, where to try it free with Gemini and Flow, the differences between Veo 2, 3 and 3.1, the real costs and the honest limits.

Erika Taranto Erika Taranto
Official selection · AI for Good 2026 11 min read
Veo 3: what it is, how to use it and what it costs
In short

Veo 3 is the video generation model of Google DeepMind, presented in May 2025: you describe the scene and it generates an eight-second clip with synchronized audio inside, dialogues included. You use it in the Gemini app, where it is also free with daily limits, in Flow (Google's tool dedicated to AI video) and through the APIs. It is the engine that follows the prompt with the most precision among the five big ones: you ask for a lateral tracking shot with rain against the light and you get that scene, not a cousin of it. Here you find how to start, what it really costs, the differences between Veo 2, 3 and 3.1 and the limits nobody writes.

In my weekly work, Veo 3 clips are the shots that come out right on the first take most often. Not because it is magic: because it is the engine most faithful to the sentence you write. The others gift you beautiful variants of the scene you asked for; Veo tends to make you the scene you asked for. If you make videos for clients, where the brief is not negotiated, this difference is everything.

This guide is the zoom on Veo 3: the complete map of the five engines, with who wins where, is in the main guide create videos with AI.

The other models give you a scene similar to the one you asked for. Veo gives you that scene: it is the reason it ends up in my projects when the brief is written.

What Veo 3 is (and where it comes from)

Veo 3 is a text-to-video model by Google DeepMind: you give it a description, or a starting image, and it returns an eight-second clip with camera, movement and audio generated together. The audio is the turning point that set it apart from the previous generation: rain that has its noise, footsteps that follow the step, characters speaking with lips in sync.

The story in three moves helps you get oriented, because on Veo you still find everything and it gets confusing:

  • Veo 2 is the previous generation: clips without audio, lower resolution, less camera control. You can still find it as a "fast" option: fine for trying, a different craft for producing.
  • Veo 3, presented in May 2025, brought synchronized audio and the quality that stands comparison with cinema photography. It is the moment AI video stopped being mute.
  • Veo 3.1, the current version, added continuity: the possibility of starting again from the last frame to extend the scene, and references to keep the same character and the same objects across different clips.

When someone searches "Veo 2" today, they almost always want to know whether the jump to 3 is worth it: it is, and audio is the number one reason.

The Veo timeline in three steps: Veo 2 without audio, Veo 3 with native synchronized audio, Veo 3.1 with references and scene continuity
The three moves of Veo: from the mute clip to native audio, up to continuity between scenes.

Veo 3 for free: where you really try it

The question I get most often is whether Veo 3 can be used for free. Yes, with precise limits:

  • Inside Gemini, free plan: you can generate a few videos per day with the Fast version of Veo, at reduced resolution and with watermark. The exact number is not officially declared and varies: in my weeks it sits around three per day. Enough to understand whether the tool is for you.
  • Inside Flow: Google Labs' video tool starts with trial credits, which run out fast. Flow is beautiful for building scenes, not for producing for free.
  • The serious free routes, all together, are in the dedicated guide: free AI video generators, where Veo is one entry in the complete picture.

My production honesty: Veo's free tier is for learning to write prompts, not for doing work. If you have a project with a deadline, the daily limits stop you halfway, and stopping halfway is the worst way to get to know a tool.

The three houses of Veo: Gemini to generate by talking to the model, Flow to build extended scenes, APIs to develop products
The three houses of Veo: one for talking, one for building, one for developing.

How to use it: the first clip in five steps

The path I would give to anyone starting today, tried on Gemini's free plan:

  1. Open Gemini (app or web) and ask for a video in natural language: "generate a video of...". No commands needed: Veo reads the sentence.
  2. Write the scene with the five elements: subject, action, camera, light, audio, in this order and in one sentence. Below I leave you an example I really use.
  3. Watch what comes out and iterate on the wrong part. The camera wanders? Re-declare it. The light is flat? Name the light source. Do not regenerate blind: every attempt has a cost, even when it is free, because time is the first budget.
  4. When the clip holds, move it into Flow if the scene needs continuing: there you can extend from the last frame and keep the ingredients, meaning the character and object references that make subsequent clips consistent.
  5. Edit outside. Eight seconds are a shot, not a video: cuts, rhythm and final audio are decided in editing. No model gives you that.
The Google logo with the words Veo 3
Veo 3 is Google's: you generate inside Gemini and Flow.
## The prompt on Veo: five pieces, one sentence

On Veo writing pays more than elsewhere, because the model executes. The structure I use, with an example that works for a first try:

A woman in a beige trench coat walks alone through a night market, camera tracks laterally at walking pace, warm string lights and steam from food stalls, ambient chatter and her footsteps.

Subject (who is in it), action (one only), camera (declared), light (the source named), audio (the sound asked for). Veo takes every piece seriously: if you declare the camera, the camera does that; if you name the light, the light arrives from there. The complete rules, including the traps like "no people in the background" that makes people appear in the background, are in the prompt section of the main guide.

Write in English even if you publish in Italian. Veo trains mostly in English and follows it with more precision: my prompts always start in English, the video ends up dubbed or subtitled in Italian. Lip sync follows English well; for Italian dialogues the voice goes in post.

Veo 2, Veo 3, Veo 3.1: what changes in practice

 Veo 2Veo 3Veo 3.1
Audiono, muteyes, synchronized, dialogues includedyes, improved
Clip durationshort clips8 seconds8 seconds, extendable from the last frame
Consistencylimitedgood on the single clipcharacter and object references across clips
Where you find itas a fast option in some toolsGemini, Flow, APIGemini, Flow, API: it is the current version
When to choose itnever, if you have 3 within reachwritten scene that must come out like thatprojects with a recurring character
Practical comparison between Veo 2, Veo 3 and Veo 3.1: audio, clip extension and consistency references
What really changes from Veo 2 to Veo 3 to Veo 3.1.

How much Veo 3 costs

Three roads, three pricing logics. Figures verified in August 2026: in this sector they move every month, so take them as a snapshot and recheck the official pages before deciding.

RoadWhat you payWhen it makes sense
Gemini free0: a few generations per day with Veo Fast, reduced resolution, watermarkLearning to write, trying scenes
Google AI plansthe monthly subscription unlocks more generations, higher quality and Flow with more creditsProducing every week
Gemini APIper clip: about $0.40 in 720-1080p, $0.60 in 4KBuilding products and automatic flows
The real bill is the attempts. Even on Veo, which gets it wrong less than the others, the first version of a clip is almost never the good one: every regeneration costs like a successful one. If you need ten clips, budget sixteen. The complete rule, with all the competitors' numbers, is in the costs section of the main guide.
The three pricing roads of Veo 3: the free tier with daily limits, the Google AI plans by subscription, the paid APIs per clip
The three pricing roads. Figures verified in August 2026.

Where Veo stumbles (and how to get around it)

I use Veo every week and I recommend it, so I owe you the real cons:

  • Eight seconds per clip. It is the format of the shot, not of the video: you extend it in Flow or you edit. Whoever arrives from "make me a one minute video" has to recalibrate the expectation: one minute is eight clips, meaning eight written scenes.
  • The safety guards. Famous faces, real people, sensitive content: Veo filters hard. In work for brands it is almost always an advantage, but know it before writing the prompt with your ambassador inside.
  • Text inside the video. Like everyone: writing comes out crooked. Texts get composed in editing, always.
  • Lip sync in Italian. English is followed well, Italian less: for Italian dialogues the voice goes in post.
  • The invisible watermark. Veo videos carry SynthID inside, Google's invisible signature declaring the AI origin. You cannot see it, you cannot remove it, and social platforms recognize it: the provenance has to be declared, always.
One minute of video equals eight eight-second clips: the tape of shots to write and edit
Eight seconds are a shot, not a video: one minute is eight written and edited scenes.

When I choose Veo, and when not

In my way of splitting scenes between engines, Veo takes the shots where precision pays: the scene is already written, the client has approved the reference, it must come out like that. When the scene asks for something else, I change engine:

  • people walking, gestures, body physics: Kling moves better;
  • scenes with several actions in a row, long clips: Seedance remembers them all;
  • whole projects with camera, consistency and extensions to manage together: I go through the directing platform, Higgsfield, where the engines live with the tools around them.

And if you need the video for your brand and prefer to have someone who does it every day: that is my craft. Write to me and we start from an idea.

The guides of the video path

Sources

Telegram

For updates and news, there is my Telegram channel

I publish the latest news on AI video tools there first.

Join the channel →
Did you like it? Share it.
Comments

Leave a comment

Comments are reviewed and approved before they appear.

Your email will not be published.

FAQ

Frequently asked questions

Is Veo 3 free? +
Partly yes. Veo 3 is included in Gemini's free plan: you can generate a few videos per day with the Fast version, at reduced resolution and with watermark. The exact number is not officially declared and varies. To produce seriously you need the paid Google AI plans or the APIs.
What is Google's Veo 3? +
It is the video generation model of Google DeepMind, presented in May 2025. It generates eight-second clips with synchronized audio, dialogues included, and lives inside Gemini, inside Flow (Google's tool dedicated to AI video) and in the APIs for developers. The current version is Veo 3.1, which added references to keep characters and objects consistent from one clip to the next.
How much does Veo 3 cost? +
Three roads: included in Gemini's free tier with daily limits; the paid Google AI plans, which raise the limits and the quality; the APIs for developers, where in August 2026 a clip cost 0.40 dollars in 720-1080p and 0.60 in 4K. Prices move: always check the official pages before deciding.
What is the difference between Veo 2 and Veo 3? +
Veo 2 is the previous generation: clips without audio, lower resolution, less control. Veo 3 brought synchronized audio, dialogues included, and a quality that stands comparison with cinema photography. Veo 3.1, the current version, added continuity between clips and references on characters and objects.
Where do you use Veo 3? +
In three ways: in the Gemini app and on the web, talking to the model in natural language; in Flow, the Google Labs tool built to construct scenes and keep characters consistent; through the Gemini APIs for those who develop. Gemini is the entry door I recommend to those starting out.
How long does a video generated with Veo 3 last? +
A Veo 3 clip lasts eight seconds. It can be extended in Flow by continuing the scene, and in practice clips get selected and edited, as in cinema: no model gives you the edit, and that is where the video becomes yours.
Does Veo 3 work in Italian? +
Yes, it accepts prompts in Italian, but video models train mostly in English and follow it better: my prompts always start in English, even for videos I then publish in Italian. Lip sync follows English well; for Italian the voice is better added in post, as in dubbing.
Keep reading
Free AI video generators: the real ways in 2026 Guides
September 2026

Free AI video generators: the real ways in 2026

Read
Create videos with AI: the 2026 guide Guides
September 2026

Create videos with AI: the 2026 guide

Read
Higgsfield: what it is, how to use it and what it costs Guides
September 2026

Higgsfield: what it is, how to use it and what it costs

Read