Generation basics
t2v — text-to-video
You type a description, the model invents the whole scene. Most flexible, least predictable.
v2v — video-to-video
Restyle existing footage while keeping its motion. Runway's specialty.
Seed
The random starting number. Same seed + same prompt = same output. Lock it when you want to tweak a prompt without changing everything else.
CFG / guidance scale
How strictly the model obeys your prompt. Low = creative drift, high = literal but often stiff. 7–12 is the usual sweet spot.
Steps / sampling steps
How many refinement passes per frame. More steps ≠ better past a point; it mostly costs time.
Quality & artifacts
Temporal coherence
Whether things stay themselves across frames. Poor coherence is why faces morph and objects melt — the single biggest quality differentiator between models.
Morphing / melting
Geometry losing its shape mid-clip. Worse on longer durations and complex scenes — a reason to keep clips at 5 seconds.
Upscaling
Generating small, then enlarging with a second model. Cheaper than native high-res and often sharper.
Interpolation / frame blending
Inventing in-between frames to raise frame rate or slow footage down. Great for smoothing, bad at hiding real coherence problems.
Control & consistency
Keyframe / first & last frame
Supplying both ends of a clip so the model fills the middle. The most reliable way to control where a shot lands.
Character consistency
Keeping the same person across shots. Reference images beat prompt descriptions every time.
LoRA
A small add-on trained on a specific style, character or object. Common in open models (Wan), rare in closed APIs.
Negative prompt
Terms telling the model what to avoid. Supported by Kling, Runway and Wan; Sora and Veo want the intent folded into the positive prompt instead.
Money & workflow
Credits
The in-app currency. Conversion rates differ wildly per tool — always translate to dollars-per-second before comparing.
Faceless channel
A channel with no on-camera presenter — AI visuals plus TTS narration. The dominant format for AI-video monetization.
TTS — text-to-speech
Synthetic narration. Free options exist (edge-tts); paid ones sound noticeably more human.
Formats
Native audio
Sound generated with the video in one pass. Veo 3.1, Sora 2 and Seedance 2.5 have it; Kling, Runway and Wan output silent clips.
New to all of this? Start with the prompt generator and the tool picks by use case.
FAQ
What does i2v mean in AI video?
Image-to-video: you supply a starting image and the model animates it, rather than inventing the entire scene from text. It's far more controllable than text-to-video because composition, characters and style are already fixed by your image.
What is temporal coherence?
Whether objects, faces and backgrounds stay consistent from frame to frame. Weak temporal coherence produces morphing faces and melting geometry, and it's the clearest quality gap between budget and flagship models.
What is the retry tax?
The gap between list price and real cost. Since roughly 3 generations are needed per usable clip, your true spend is about triple the advertised per-second rate.