{"id":56606,"date":"2026-10-01T09:22:52","date_gmt":"2026-10-01T01:22:52","guid":{"rendered":"https:\/\/www.1ai.net\/?p=56606"},"modified":"2026-09-02T14:27:52","modified_gmt":"2026-09-02T06:27:52","slug":"ai%e7%94%9f%e6%88%90%e8%a7%86%e9%a2%91%e5%85%a5%e9%97%a8%ef%bc%9a%e6%99%ae%e9%80%9a%e4%ba%ba%e7%ac%ac%e4%b8%80%e6%ac%a1%e5%81%9aai%e8%a7%86%e9%a2%91%ef%bc%8c%e5%ba%94%e8%af%a5%e5%85%88%e6%90%9e","status":"publish","type":"post","link":"https:\/\/www.1ai.net\/en\/56606.html","title":{"rendered":"AI GENERATED AN INTRODUCTORY VIDEO: FOR THE FIRST TIME, A COMMON PERSON DOES AN AI VIDEO, AND SHOULD FIRST UNDERSTAND THE UNDERLYING CONCEPTS"},"content":{"rendered":"<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56609\" title=\"e91eka98jkkq4fj00vkd000v90kcip\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/e91eca98j00tkq4fj00vkd000v900cip.jpg\" alt=\"e91eka98jkkq4fj00vkd000v90kcip\" width=\"1125\" height=\"450\" \/><\/p>\n<p>If you've just begun to involve<a href=\"https:\/\/www.1ai.net\/en\/tag\/ai%e8%a7%86%e9%a2%91%e5%88%b6%e4%bd%9c\" title=\"_OTHER ORGANISER\" target=\"_blank\" >AI video production<\/a>You have to understand these concepts first<\/p>\n<p>I think I understand what these terms mean<\/p>\n<p>Do<a href=\"https:\/\/www.1ai.net\/en\/tag\/ai%e8%a7%86%e9%a2%91\" title=\"[View articles tagged with [AI Video]]\" target=\"_blank\" >AI Video<\/a>It starts with four large blocks:<\/p>\n<p><em><strong>AUDIO, VIDEO BOTTOM, AI GENERATION RELATED, CURRENT-LINE ENGINEERING CONCEPTS<\/strong><\/em><\/p>\n<p>Each concept states only:<strong>What is it<\/strong><\/p>\n<p>I. Audio portion (coding, subtitles)<\/p>\n<p><strong>1. TTS TEXT-TO-SPEECH<\/strong>\u00a0Generates the text directly into a human voice\u3002<\/p>\n<blockquote>\n<ul>\n<li>Common people focus: There are clouds (Edge-TTS, Eleven Labs) and local (Kokoro); there is no local connection and the cloud is better sounded\u3002<\/li>\n<\/ul>\n<\/blockquote>\n<p><strong>2. ASR VOICEOVER TEXT<\/strong>\u00a0Turn the sound back into text for subtitles\u3002<strong>Whisper is an ASR tool<\/strong>.<\/p>\n<blockquote>\n<ul>\n<li>FOCUS ON NORMAL PEOPLE: SUBTITLES ARE NOT A DIRECT PASTING OF THE SCRIPT; THEY'RE AI LISTENING TO AUDIO, ALIGNING THE SOUND AXIS\u3002<\/li>\n<\/ul>\n<\/blockquote>\n<p><strong>Timetamp \/ Font Timetamp<\/strong>\u00a0Every word, every sentence corresponds to the beginning, the end\u3002<\/p>\n<blockquote>\n<ul>\n<li>The common man's focus is on time-stamping to synchronize with people's voices\u3002<\/li>\n<\/ul>\n<\/blockquote>\n<p><strong>4. SRT SUBTITLE FILES<\/strong>\u00a0postfixes.srt, plain text subtitle files, ffmpeg, clippings\u3002<\/p>\n<p><strong>5. SEPARATING HUMAN VOICES (UVR5)<\/strong>\u00a0It's not like you're going to be able to use the sound of people and background music\u3002<\/p>\n<blockquote>\n<ul>\n<li>The general human focus: secondary clips, negatives will be used\u3002<\/li>\n<\/ul>\n<\/blockquote>\n<p><strong>SOUND CLONING RVC<\/strong>\u00a0Take a couple of real-life audios and reset the person's voice\u3002<\/p>\n<blockquote>\n<ul>\n<li>The focus of ordinary people is to limit themselves to their own voices and not to steal others ' voices\u3002<\/li>\n<\/ul>\n<\/blockquote>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56607\" title=\"9b6cc28dj00tkq4fx00mgd000v90jp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/9b6cc28dj00tkq4fx00mgd000v900jjp.jpg\" alt=\"9b6cc28dj00tkq4fx00mgd000v90jp\" width=\"1125\" height=\"703\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56608\" title=\"4508832aj00tkq4g300iod000v90jp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/4508832aj00tkq4g300iod000v900jjp.jpg\" alt=\"4508832aj00tkq4g300iod000v90jp\" width=\"1125\" height=\"703\" \/><\/p>\n<p>II. Bottom video tools (synthetic, graphic, compulsory)<\/p>\n<p><strong>1. ffmpeg<\/strong>\u00a0Command line audio and video toolboxes, clipping, adjoining, crushing the video, burning subtitles, reformatting\u3002<\/p>\n<blockquote>\n<ul>\n<li>ORDINARY PEOPLE FOCUS: ALMOST ALL OF THE AI VIDEO STREAMING LINES RELY ON IT, THERE IS NO GRAPHICAL INTERFACE, SCRIPTS CALL IT, AND ORDINARY PEOPLE DO NOT WRITE COMPLEX COMMANDS BY HAND\u3002<\/li>\n<\/ul>\n<\/blockquote>\n<p><strong>2. Frame \/ Frame series<\/strong>\u00a0The nature of the video is a quick-playing picture, a picture called a frame; a bunch of pictures called frame sequences\u3002<\/p>\n<blockquote>\n<ul>\n<li>Focus on normal people: HyperFrames and ComfyUI often output a bunch of pictures and hand them over to ffmpeg for video\u3002<\/li>\n<\/ul>\n<\/blockquote>\n<p><strong>3. FPS FRAME RATE<\/strong>\u00a0How many pictures are played per second, short video is usually 24\/30 frames\u3002<\/p>\n<p><strong>RIFE<\/strong>\u00a0AI CREATES AN INTERMEDIATE IMAGE THAT FLOWS DOWN THE LOW FRAME RATE, SOLVES THE AI VIDEO CARDON, A ONE-ON-ONE\u3002<\/p>\n<p><strong>5. Overscoring<\/strong>\u00a0AI zooms in to make it clear. Representative: Real-ESRgan\u3002<\/p>\n<p><strong>6. Face repair CodeFormer \/ GFPgan<\/strong>\u00a0AI PRODUCES VIDEO, OFTEN DISTORTED AND BROKEN FACES, A TOOL DEDICATED TO REPAIRING HUMAN FACES\u3002<\/p>\n<p><strong>7. code (libx264)<\/strong>\u00a0VIDEO COMPRESSION ALGORITHMS, WE NORMALLY USE MP4 SHORT VIDEOS, MOSTLY H264 CODED\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56610\" title=\"4d9d88a5j00tkq4gc00jd000v90jp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/4d9d88a5j00tkq4gc00ijd000v900jjp.jpg\" alt=\"4d9d88a5j00tkq4gc00jd000v90jp\" width=\"1125\" height=\"703\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56611\" title=\"5552315 dj00tkq4gf00jad000v90jp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/5552315dj00tkq4gf00jad000v900jjp.jpg\" alt=\"5552315 dj00tkq4gf00jad000v90jp\" width=\"1125\" height=\"703\" \/><\/p>\n<p>III. AI GENERATING IMAGE CORE CONCEPT<\/p>\n<p><strong>Prompt Phrasing<\/strong>\u00a0TELL AI WHAT IMAGE TO GENERATE FOR THE TEXT DESCRIPTION\u3002<\/p>\n<p><strong>Seeds<\/strong>\u00a0A STRING OF NUMBERS, THE SAME FEED + THE SAME HINT, AI GENERATES ALMOST IDENTICAL IMAGES; A SEED IMAGE CHANGES\u3002<\/p>\n<blockquote>\n<ul>\n<li>The focus of ordinary people is: if the picture is to remain stable, the seed is to be fixed\u3002<\/li>\n<\/ul>\n<\/blockquote>\n<p><strong>3. LoRA<\/strong>\u00a0Small model files to fix the person's appearance and style, and to avoid differences in the size of each frame\u3002<\/p>\n<p><strong>4. Time-series consistency<\/strong>\u00a0AI VIDEOS WITH MAXIMUM PAIN: FRONT-AND-BACK FRAME OBJECTS, HUMAN FACES, FLASH. TIME-SERIES CONSISTENCY = IMAGE OBJECTS DO NOT DRIFT\u3002<\/p>\n<p><strong>5. TUSHENG VIDEO I2V<\/strong>\u00a0Enter a picture that AI turns into a dynamic video clip (Wan2.2, AnimateDiff)\u3002<\/p>\n<p><strong>6. VINCENT VIDEO T2V<\/strong>\u00a0ENTER TEXT ONLY, AI DIRECTLY GENERATE VIDEO\u3002<\/p>\n<p><strong>7. Video redraw V2V Denoise Go Noise<\/strong>\u00a0TAKE A READY-TO-USE VIDEO TO ALLOW AI TO RE-ENGINEER THE IMAGE; THE VALUE 0:1, THE MORE THE VALUE CHANGES\u3002<\/p>\n<p><strong>Lip-Sync<\/strong>\u00a0Audio-drive pictures \/ Virtual human lips follow voice\u3002<\/p>\n<p><strong>ControlNet Control Network<\/strong>\u00a0COMPRESSING AI IMAGES, SUCH AS FIXED PERSON MOVEMENTS, CONTOURS, NOT ALLOWING AI TO CHANGE POSITIONS\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56612\" title=\"628db51aj00tkq4gm00jod000v90jp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/628db51aj00tkq4gm00jod000v900jjp.jpg\" alt=\"628db51aj00tkq4gm00jod000v90jp\" width=\"1125\" height=\"703\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56616\" title=\"ecbd437ej00tkq4gt00jwd000v90jp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/ecbd437ej00tkq4gt00jwd000v900jjp.jpg\" alt=\"ecbd437ej00tkq4gt00jwd000v90jp\" width=\"1125\" height=\"703\" \/><\/p>\n<p>IV. Flow &amp; Tool Concept (automated video focus)<\/p>\n<ol>\n<li><strong>Inference<\/strong>\u00a0THE AI MODEL RUNS UP AND GENERATES PICTURES AND VIDEOS. THE LOCAL COMPUTER RAN THE THEORY TO EAT THE GRAPHICS\u3002<\/li>\n<li><strong>Quantitative<\/strong>\u00a0DECOMPRESS THE AI MODEL, LOWER THE GRAPHIC CARD REQUIREMENTS, SLIGHTLY LOSE THE PAINT QUALITY, AND LOW-END COMPUTERS ARE COMMONLY USED\u3002<\/li>\n<li><strong>ComfyUI<\/strong>\u00a0NODE-BASED VISUALIZATION AI TOOL, DRAG AND DRAG NODE SERIALS: BIOGRAPHS, RAW VIDEOS, FRAMES, HYPERPOINTS, LOCAL AI VIDEO MAINSTREAM TOOL\u3002<\/li>\n<li><strong>HyperFrames<\/strong>\u00a0Animation with HTML\/CSS, rendering frame and ffmpeg synthetic video; suitable for animation in knowledge\u3002<\/li>\n<li><strong>MoviePy<\/strong>\u00a0Python tool, seal ffmpeg, write scripts for automatic clips\u3002<\/li>\n<li><strong>Smart Body (WorkBuddy \/ Codex)<\/strong>\u00a0You can't write a code, you can describe a video process in a natural language, and it helps you generate Python scripts, and links all the tools above\u3002<\/li>\n<\/ol>\n<blockquote>\n<ul>\n<li>common human focus: smarts only output code, environment, software to install themselves; codes have bugs, thrown back and changed\u3002<\/li>\n<\/ul>\n<\/blockquote>\n<p><strong>7. Waterlines<\/strong>\u00a0A set of fixed sequences: Synthetic subtitling images are created to enhance the quality and synthesis of the product. Previous output file as next input\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56614\" title=\"43350b0bj00tkq4h000gwd000v90jp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/43350b0bj00tkq4h000gwd000v900jjp.jpg\" alt=\"43350b0bj00tkq4h000gwd000v90jp\" width=\"1125\" height=\"703\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56615\" title=\"ba943a6dj00tkq4h400i3d000v90jp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/ba943a6dj00tkq4h400i3d000v900jjp.jpg\" alt=\"ba943a6dj00tkq4h400i3d000v90jp\" width=\"1125\" height=\"703\" \/><\/p>\n<p>\u2022 New hands to avoid pit awareness (very important, not noun, to understand)<\/p>\n<ol>\n<li>AI VIDEO IS NOT A BUTTON-POINTING, AND IT IS A MULTI-TOOL RELAY TO PROCESS FILES\u3002<\/li>\n<li>Priority is given to running through the minimum flow line closed loops, with additional frames, overscores, living video modules and no one-time stacking\u3002<\/li>\n<li>LOCALLY RUN AI, GRAPHIC CARD DISPLAY DETERMINES THE SPEED; SMALLER MODELS ARE USED TO QUANTIFY IF THERE IS INSUFFICIENT VISIBILITY\u3002<\/li>\n<li>all tools are read and write documents: txt, mp3, srt, png, mp4, understand the flow of documents and understand the whole logic\u3002<\/li>\n<\/ol>\n<p>Here's the most complete list of proprietary terms<\/p>\n<p>\ud83c\udfac AI VIDEO CREATION. FULL PROPRIETARY TERM SCHEMATIC<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56613\" title=\"2aeda4bj00tkq4hd006d000v90ep\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/2aedaa4bj00tkq4hd006dd000v900iep.jpg\" alt=\"2aeda4bj00tkq4hd006d000v90ep\" width=\"1125\" height=\"662\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56618\" title=\"0990ea04j00tkq4hi007xd000v90hrp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/0990ea04j00tkq4hi007xd000v900hrp.jpg\" alt=\"0990ea04j00tkq4hi007xd000v90hrp\" width=\"1125\" height=\"639\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56617\" title=\"55e68451j00tkq4m0075d000v9000l6p\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/55e68451j00tkq4hm0075d000v900l6p.jpg\" alt=\"55e68451j00tkq4m0075d000v9000l6p\" width=\"1125\" height=\"762\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56619\" title=\"0d3ad21bj00tkq4ht006bd000v90emp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/0d3ad21bj00tkq4ht006bd000v900emp.jpg\" alt=\"0d3ad21bj00tkq4ht006bd000v90emp\" width=\"1125\" height=\"526\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56620\" title=\"4cdebd4dj00tkq4hx0090d000v90l4p\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/4cdebd4dj00tkq4hx0090d000v900l4p.jpg\" alt=\"4cdebd4dj00tkq4hx0090d000v90l4p\" width=\"1125\" height=\"760\" \/><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56621\" title=\"7c9a8f47jkkk4i40004ed000v9000cmp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/7c9a8f47j00tkq4i4004ed000v900cmp.jpg\" alt=\"7c9a8f47jkkk4i40004ed000v9000cmp\" width=\"1125\" height=\"454\" \/><\/p>","protected":false},"excerpt":{"rendered":"<p>If you start with AI video production, you have to understand these concepts, and perhaps understand the meaning of these terms, and make it into four blocks: audio, video bottom, AI generation related, streaming-wire engineering, each of which says only what: + the focus that ordinary people need to know is \ud83d\udd0a1 +, the audio part (coding, subtitles) 1.TTS text-to-speeching directly to human voice. Common people focus: There are clouds (Edge-TTS, Eleven Labs) and local (Kokoro); there is no local connection and the cloud is better sounded. 2. ASR Voice-to-Speech Text Turns Sound Back to Text for Subtitles. Whisper is an ASR tool. == sync, corrected by elderman ==<\/p>","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[149,144],"tags":[3765,956,5640,5321,2894],"collection":[],"class_list":{"0":"post-56606","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"hentry","6":"category-jiaocheng","7":"category-baike","8":"tag-ai","12":"tag-2894"},"acf":[],"_links":{"self":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts\/56606","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/comments?post=56606"}],"version-history":[{"count":0,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts\/56606\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/media?parent=56606"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/categories?post=56606"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/tags?post=56606"},{"taxonomy":"collection","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/collection?post=56606"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}