Veo 3.1 is the opposite of Seedance: it wants full sentences, not comma-separated blocks. And it's the only major model where what you write becomes what you hear — audio is part of the prompt, not an afterthought.
Veo was trained on descriptive prose. Sentences with grammar outperform keyword soup, and the model uses sentence order to infer what happens first.
A weathered fisherman mends his net on a wooden dock at dawn. The camera slowly pushes in as mist rolls off the water behind him. Soft golden light catches the frayed rope in his hands. Shot on 35mm film with natural grain.
fisherman, dock, dawn, mist, push in, golden light, 35mm, cinematic, 8K, masterpiece
Veo generates synchronized sound in the same pass. You cue it three ways, and they behave very differently:
She looks directly at the camera and says: "We were never supposed to find this."
Ambient audio: distant traffic, light rain on the window, a kettle beginning to whistle.
No music, no narration — only natural room tone.
If nobody speaks and you're generating at volume, you're paying for audio you won't use — Kling costs less than half as much for silent clips. If the shot has to look expensive above all else, Seedance leads on pure visual quality. Compare directly in Veo vs Sora and Seedance vs Veo, or price your job in the cost calculator.
Ready-made prompts in this style: 12 Veo examples. Writing for the other models instead? See the Seedance guide · Kling guide · Sora guide.
Write it as a sentence with the line in quotation marks: She looks at the camera and says: "...". Veo lip-syncs to the quoted text. Keep it to one sentence per clip and avoid fast camera movement during speech.
Veo generates audio by default and fills silence with a score. State your intent explicitly — "no music, only natural room tone" — whenever you plan to add your own soundtrack.
Full sentences. Veo was trained on descriptive prose and uses sentence order to infer sequence. Keyword lists work better on Seedance; on Veo they waste the model's strongest capability.