Veo 3.1
by Google DeepMind
Video with sound generated in the same pass, in 1080p or 4K.
What it is
Veo 3.1 is Google DeepMind's video model, described by its makers as "our leading video generation model, designed to empower filmmakers and storytellers". It generates eight-second sequences in 1080p and 4K, and it is one of the few models that produces audio natively rather than leaving you to lay sound over a silent clip.
That audio is the reason people pick it: sound effects, ambient noise and dialogue arrive with the picture, already matched to what is on screen. For a social cut or an ad, that removes an entire step.
It also does the things a short piece actually needs , camera moves you specify, an image turned into a moving shot, a transition built from a first and a last frame, and scene extension. In this studio all of that runs from the same screen.
What it is good at
Google's own framing is "video, meet audio": effects, ambience and dialogue generated natively inside the video. No separate pass, no library search.
Framing and shot movement , zoom, pan, dolly , are things you ask for. The camera presets in the Studio write those moves into the prompt for you, word for word.
Google describes "smooth, artful, and epic transitions between images provided for the first and last frame". Approve two stills, get the move between them.
Style reference matching, character consistency, scene extension, object insertion and removal, and outpainting , the toolkit for a second shot that belongs with the first.
Where it struggles
- Google is explicit about one thing: natural, consistent spoken audio, especially for short speech segments, "remains an area of active development". Long, clean dialogue is not what this is for yet.
- Sequences are eight seconds. Anything longer is several shots that have to be made to belong together , which is what a script does in this studio.
- Audio raises the price. The Studio shows the difference on the button before you run.
How to prompt it
Put the motion first. What moves, and what the camera does while it moves , then the setting, then the light. Because the audio is generated with the picture, naming what should be heard is worth doing: footsteps on gravel, a room tone, a single line of dialogue. Google keeps a dedicated prompt guide for the model.
An example that shows what that means
A chef slides a cast-iron pan onto a flame and the oil catches, camera pushing in slowly from waist height to a tight shot on the pan, warm kitchen light from a window on the left, the rest of the room falling dark. Audio: the hiss of oil, a low extractor hum, no music.
What we offer here
Straight from the catalogue, so this list never goes stale. The price is the provider’s published rate with our markup on top, and it is on the button before you run anything.
Questions
Does it really make sound, or do I add it after?
It makes it. Google generates sound effects, ambient noise and dialogue natively inside the video. You can turn audio off, and the Studio shows what that saves before you run.
How long can a clip be?
Eight seconds per generation. For thirty seconds you build a script: several shots that share one style block, so they read as one piece instead of separate attempts.
Can I control the camera move?
Yes. Zoom, pan and dolly are things you specify, and the camera presets in the Studio put the exact phrasing into the prompt so you are not guessing at the wording.
Can I start from an image I already have?
Yes , image to video, and first frame to last frame for a controlled transition. Both run from the same screen here.