AI video glossary
Generative video terms, explained in plain English.
Clear definitions for AI video generation, text to video, image to video, talking avatars, lip sync, neural voices, dubbing, and more.
Term index
Talking avatar
A digital character — usually built from a single photo — whose lips, jaw, and expressions are animated by AI to match a chosen voice or script.
Read more →Lip sync
The frame-by-frame alignment of a face's mouth movements to a target audio track so that the speaker visibly forms the right sounds.
Read more →Text-to-video
A workflow that turns a written script into a finished video — usually generating a voiceover, animating a face or scene and assembling the result.
Read more →Photo-to-video
Generating a moving video from a single still photograph — most commonly by animating the subject's face to speak.
Read more →AI presenter
A virtual on-camera spokesperson generated by AI — used for explainer videos, product demos, courses, and internal communications.
Read more →Voice cloning
Synthesizing a new voice that sounds like a specific real person, typically from a short audio sample.
Read more →Neural voice
A text-to-speech voice generated by a deep neural network, producing more natural intonation and emotion than older concatenative or formant TTS.
Read more →AI dubbing
Automatically re-voicing a video into a new language, ideally with matched lip-sync and the original speaker's vocal identity.
Read more →Deepfake
Synthetic media where a person's face or voice is swapped or animated by AI. Consent-based talking avatars are a legitimate use of the same tech.
Read more →Text-to-speech (TTS)
Technology that converts written text into spoken audio. Modern neural TTS voices sound nearly human and power most talking-avatar tools today.
Read more →AI video generator
Software that creates video from prompts, scripts, images, audio, or existing footage. It may generate one shot or plan and assemble a complete multi-scene video.
Read more →Talking-head video
A video format where one person (or avatar) talks directly to camera — the dominant format for tutorials, sales, courses and social explainers.
Read more →
