Why creators pick Veo 3.1
Veo 3.1 is Google DeepMind's video model and the strongest option on WandVideo for photorealism. It generates synchronized audio, including speech, ambience and sound effects, and can output 720p, 1080p or 4K. Clips are 4, 6 or 8 seconds long.
Modes available on WandVideo
- Text to video in 16:9 or 9:16.
- Image to video from a single start image, with the aspect ratio inferred automatically or forced to 16:9 / 9:16.
Both modes accept a negative prompt and an optional seed for reproducible results.
Prompting tips
- Write dialogue in quotes and state who says it: "A barista looks up and says, 'Your usual?'"
- Describe sound explicitly when it matters: rain on a tin roof, distant traffic, a vinyl crackle.
- Use 720p while iterating; switch to 1080p or 4K for the final render to save credits.
- Veo follows lens and lighting language well: anamorphic, golden hour, overcast softbox.
Credits
Veo 3.1 is the premium model on WandVideo. Cost scales with duration, resolution tier (720p/1080p versus 4K) and whether audio is enabled. Silent 720p clips are the most economical; 4K with audio is the most expensive combination.
Safety and watermarking
Veo applies its own safety filters to prompts, input images and output. It blocks explicit content and likenesses of real public figures. All Veo output carries an invisible SynthID watermark. Failed generations are refunded.


