
If you've just begun to involveAI video productionYou have to understand these concepts first
I think I understand what these terms mean
DoAI VideoIt starts with four large blocks:
AUDIO, VIDEO BOTTOM, AI GENERATION RELATED, CURRENT-LINE ENGINEERING CONCEPTS
Each concept states only:What is it
I. Audio portion (coding, subtitles)
1. TTS TEXT-TO-SPEECH Generates the text directly into a human voice。
- Common people focus: There are clouds (Edge-TTS, Eleven Labs) and local (Kokoro); there is no local connection and the cloud is better sounded。
2. ASR VOICEOVER TEXT Turn the sound back into text for subtitles。Whisper is an ASR tool.
- FOCUS ON NORMAL PEOPLE: SUBTITLES ARE NOT A DIRECT PASTING OF THE SCRIPT; THEY'RE AI LISTENING TO AUDIO, ALIGNING THE SOUND AXIS。
Timetamp / Font Timetamp Every word, every sentence corresponds to the beginning, the end。
- The common man's focus is on time-stamping to synchronize with people's voices。
4. SRT SUBTITLE FILES postfixes.srt, plain text subtitle files, ffmpeg, clippings。
5. SEPARATING HUMAN VOICES (UVR5) It's not like you're going to be able to use the sound of people and background music。
- The general human focus: secondary clips, negatives will be used。
SOUND CLONING RVC Take a couple of real-life audios and reset the person's voice。
- The focus of ordinary people is to limit themselves to their own voices and not to steal others ' voices。


II. Bottom video tools (synthetic, graphic, compulsory)
1. ffmpeg Command line audio and video toolboxes, clipping, adjoining, crushing the video, burning subtitles, reformatting。
- ORDINARY PEOPLE FOCUS: ALMOST ALL OF THE AI VIDEO STREAMING LINES RELY ON IT, THERE IS NO GRAPHICAL INTERFACE, SCRIPTS CALL IT, AND ORDINARY PEOPLE DO NOT WRITE COMPLEX COMMANDS BY HAND。
2. Frame / Frame series The nature of the video is a quick-playing picture, a picture called a frame; a bunch of pictures called frame sequences。
- Focus on normal people: HyperFrames and ComfyUI often output a bunch of pictures and hand them over to ffmpeg for video。
3. FPS FRAME RATE How many pictures are played per second, short video is usually 24/30 frames。
RIFE AI CREATES AN INTERMEDIATE IMAGE THAT FLOWS DOWN THE LOW FRAME RATE, SOLVES THE AI VIDEO CARDON, A ONE-ON-ONE。
5. Overscoring AI zooms in to make it clear. Representative: Real-ESRgan。
6. Face repair CodeFormer / GFPgan AI PRODUCES VIDEO, OFTEN DISTORTED AND BROKEN FACES, A TOOL DEDICATED TO REPAIRING HUMAN FACES。
7. code (libx264) VIDEO COMPRESSION ALGORITHMS, WE NORMALLY USE MP4 SHORT VIDEOS, MOSTLY H264 CODED。


III. AI GENERATING IMAGE CORE CONCEPT
Prompt Phrasing TELL AI WHAT IMAGE TO GENERATE FOR THE TEXT DESCRIPTION。
Seeds A STRING OF NUMBERS, THE SAME FEED + THE SAME HINT, AI GENERATES ALMOST IDENTICAL IMAGES; A SEED IMAGE CHANGES。
- The focus of ordinary people is: if the picture is to remain stable, the seed is to be fixed。
3. LoRA Small model files to fix the person's appearance and style, and to avoid differences in the size of each frame。
4. Time-series consistency AI VIDEOS WITH MAXIMUM PAIN: FRONT-AND-BACK FRAME OBJECTS, HUMAN FACES, FLASH. TIME-SERIES CONSISTENCY = IMAGE OBJECTS DO NOT DRIFT。
5. TUSHENG VIDEO I2V Enter a picture that AI turns into a dynamic video clip (Wan2.2, AnimateDiff)。
6. VINCENT VIDEO T2V ENTER TEXT ONLY, AI DIRECTLY GENERATE VIDEO。
7. Video redraw V2V Denoise Go Noise TAKE A READY-TO-USE VIDEO TO ALLOW AI TO RE-ENGINEER THE IMAGE; THE VALUE 0:1, THE MORE THE VALUE CHANGES。
Lip-Sync Audio-drive pictures / Virtual human lips follow voice。
ControlNet Control Network COMPRESSING AI IMAGES, SUCH AS FIXED PERSON MOVEMENTS, CONTOURS, NOT ALLOWING AI TO CHANGE POSITIONS。


IV. Flow & Tool Concept (automated video focus)
- Inference THE AI MODEL RUNS UP AND GENERATES PICTURES AND VIDEOS. THE LOCAL COMPUTER RAN THE THEORY TO EAT THE GRAPHICS。
- Quantitative DECOMPRESS THE AI MODEL, LOWER THE GRAPHIC CARD REQUIREMENTS, SLIGHTLY LOSE THE PAINT QUALITY, AND LOW-END COMPUTERS ARE COMMONLY USED。
- ComfyUI NODE-BASED VISUALIZATION AI TOOL, DRAG AND DRAG NODE SERIALS: BIOGRAPHS, RAW VIDEOS, FRAMES, HYPERPOINTS, LOCAL AI VIDEO MAINSTREAM TOOL。
- HyperFrames Animation with HTML/CSS, rendering frame and ffmpeg synthetic video; suitable for animation in knowledge。
- MoviePy Python tool, seal ffmpeg, write scripts for automatic clips。
- Smart Body (WorkBuddy / Codex) You can't write a code, you can describe a video process in a natural language, and it helps you generate Python scripts, and links all the tools above。
- common human focus: smarts only output code, environment, software to install themselves; codes have bugs, thrown back and changed。
7. Waterlines A set of fixed sequences: Synthetic subtitling images are created to enhance the quality and synthesis of the product. Previous output file as next input。


• New hands to avoid pit awareness (very important, not noun, to understand)
- AI VIDEO IS NOT A BUTTON-POINTING, AND IT IS A MULTI-TOOL RELAY TO PROCESS FILES。
- Priority is given to running through the minimum flow line closed loops, with additional frames, overscores, living video modules and no one-time stacking。
- LOCALLY RUN AI, GRAPHIC CARD DISPLAY DETERMINES THE SPEED; SMALLER MODELS ARE USED TO QUANTIFY IF THERE IS INSUFFICIENT VISIBILITY。
- all tools are read and write documents: txt, mp3, srt, png, mp4, understand the flow of documents and understand the whole logic。
Here's the most complete list of proprietary terms
🎬 AI VIDEO CREATION. FULL PROPRIETARY TERM SCHEMATIC





