AI Video Glossary

Every term you'll hit in an AI video tool's UI, explained in plain English — plus whether it actually matters for what you're making.

Generation basics

t2v — text-to-video

You type a description, the model invents the whole scene. Most flexible, least predictable.

i2v — image-to-video

You supply the first frame, the model only adds motion. Far more controllable — see the i2v prompt guide.

v2v — video-to-video

Restyle existing footage while keeping its motion. Runway's specialty.

Seed

The random starting number. Same seed + same prompt = same output. Lock it when you want to tweak a prompt without changing everything else.

CFG / guidance scale

How strictly the model obeys your prompt. Low = creative drift, high = literal but often stiff. 7–12 is the usual sweet spot.

Steps / sampling steps

How many refinement passes per frame. More steps ≠ better past a point; it mostly costs time.

Quality & artifacts

Temporal coherence

Whether things stay themselves across frames. Poor coherence is why faces morph and objects melt — the single biggest quality differentiator between models.

Flicker

Frame-to-frame brightness or texture jitter. Fix with negative prompts and slower described motion — see the negative prompt list.

Morphing / melting

Geometry losing its shape mid-clip. Worse on longer durations and complex scenes — a reason to keep clips at 5 seconds.

Upscaling

Generating small, then enlarging with a second model. Cheaper than native high-res and often sharper.

Interpolation / frame blending

Inventing in-between frames to raise frame rate or slow footage down. Great for smoothing, bad at hiding real coherence problems.

Control & consistency

Keyframe / first & last frame

Supplying both ends of a clip so the model fills the middle. The most reliable way to control where a shot lands.

Character consistency

Keeping the same person across shots. Reference images beat prompt descriptions every time.

LoRA

A small add-on trained on a specific style, character or object. Common in open models (Wan), rare in closed APIs.

Negative prompt

Terms telling the model what to avoid. Supported by Kling, Runway and Wan; Sora and Veo want the intent folded into the positive prompt instead.

Camera phrase

Film vocabulary the model understands: dolly, orbit, crane, FPV. One per shot — see the 9 camera moves.

Money & workflow

Credits

The in-app currency. Conversion rates differ wildly per tool — always translate to dollars-per-second before comparing.

Retry tax

The hidden multiplier: nobody nails a shot on take one, so real cost ≈ list price × 3. Model it in the cost calculator.

RPM

Revenue per 1,000 views. The number that decides whether your content is a business — Shorts sit near $0.20, long-form tech/finance reaches $8–20. See the RPM calculator.

Faceless channel

A channel with no on-camera presenter — AI visuals plus TTS narration. The dominant format for AI-video monetization.

TTS — text-to-speech

Synthetic narration. Free options exist (edge-tts); paid ones sound noticeably more human.

Formats

Aspect ratio

9:16 for Shorts/TikTok/Reels, 16:9 for standard YouTube, 21:9 for cinematic. Convert in the ratio calculator.

Bitrate

Data per second of video — the main driver of file size and export quality. Check yours in the file size calculator.

Native audio

Sound generated with the video in one pass. Veo 3.1, Sora 2 and Seedance 2.5 have it; Kling, Runway and Wan output silent clips.

New to all of this? Start with the prompt generator and the tool picks by use case.

FAQ

What does i2v mean in AI video?

Image-to-video: you supply a starting image and the model animates it, rather than inventing the entire scene from text. It's far more controllable than text-to-video because composition, characters and style are already fixed by your image.

What is temporal coherence?

Whether objects, faces and backgrounds stay consistent from frame to frame. Weak temporal coherence produces morphing faces and melting geometry, and it's the clearest quality gap between budget and flagship models.

What is the retry tax?

The gap between list price and real cost. Since roughly 3 generations are needed per usable clip, your true spend is about triple the advertised per-second rate.