{"id":55989,"date":"2026-08-19T10:09:40","date_gmt":"2026-08-19T02:09:40","guid":{"rendered":"https:\/\/www.1ai.net\/?p=55989"},"modified":"2026-08-19T10:10:25","modified_gmt":"2026-08-19T02:10:25","slug":"ai%e6%95%b0%e5%ad%97%e4%ba%ba%e6%9c%ac%e5%9c%b0%e9%83%a8%e7%bd%b2%e6%95%99%e7%a8%8b%ef%bc%8c0%e6%88%90%e6%9c%ac%e7%9a%84%e6%95%b0%e5%ad%97%e4%ba%ba%e5%8f%a3%e6%92%ad%e6%95%99%e7%a8%8b","status":"publish","type":"post","link":"https:\/\/www.1ai.net\/en\/55989.html","title":{"rendered":"LOCAL DEPLOYMENT OF THE AID DIGITAL POPULATION AT ZERO COST"},"content":{"rendered":"<p>Local deployment: zero cost<a href=\"https:\/\/www.1ai.net\/en\/tag\/%e6%95%b0%e5%ad%97%e4%ba%ba\" title=\"[View articles tagged with [digital people]]\" target=\"_blank\" >Digital Human<\/a><a href=\"https:\/\/www.1ai.net\/en\/tag\/%e5%8f%a3%e6%92%ad\" title=\"[Sees articles with [oral] labels]\" target=\"_blank\" >Oral<\/a>Tutorial<\/p>\n<p>all the digital population programmes seen before were based on heygen, 1 minute, about $10, expensive\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-55990\" title=\"f892b98j00tjzv1b00n8d000v90idp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/08\/f892bb98j00tjzv1b00n8d000v900idp.jpg\" alt=\"f892b98j00tjzv1b00n8d000v90idp\" width=\"1125\" height=\"661\" \/><\/p>\n<p>The programme is a local telecast digital programme, which requires only one Nvidia graphic card\u3002<\/p>\n<p><strong>Chapter 1<\/strong><\/p>\n<p><strong>1.1 Three common configurations<\/strong><\/p>\n<p>Let's see if you have a computer configuration before you decide whether to use local services or cloud services\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-55991\" title=\"49d2d36bj00tjzv26005rd000v9000bkp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/08\/49d2d36bj00tjzv26005rd000v900bkp.jpg\" alt=\"49d2d36bj00tjzv26005rd000v9000bkp\" width=\"1125\" height=\"416\" \/><\/p>\n<p><strong>WHAT ABOUT THE NVIDIA<\/strong><\/p>\n<p>Skip local digital synthesis. Articles can continue to be published orally, audio, subtitles, illustrations, animations and video renderings\u3002<\/p>\n<p>The digital person simply folds a good mouth into a basic video without affecting the production of the previous content. You can also choose the cloud-based digital programme\u3002<\/p>\n<p><strong>Chapter 2<\/strong><\/p>\n<p><strong>2.1 Start-up hints that can be copied directly<\/strong><\/p>\n<p>copys the text dropped to codex, replacing only the contents in square brackets\u3002<\/p>\n<p>please install and use angent-video-pipeline<br \/>\nhttps:\/\/github.com\/JayceHuang\/agent-video-pipeline.git<\/p>\n<p>Make a video of [article path, attachment or lower body]\u3002<\/p>\n<p>Workspace is [absolute path]\u3002<br \/>\nThe project directory is [absolute path]\u3002<br \/>\nThe release platform is [insert]\u3002<br \/>\nThe painting is [16:9 or 9:16]\u3002<br \/>\nThis time [no digitals or local digitals]\u3002<\/p>\n<p>When installed, read README.md and SKILL.md in full, check the hardware and Git, Python 3.10 or higher, ffmpeg, ffprobe, Node.js. Priority is given to installation or repair in the work area where there is a lack of dependency and no contamination of the system environment\u3002<\/p>\n<p>check the complete .agent-video\/ area. starts the neutral configuration for running the command below when the missing or the structure is incomplete\u3002<\/p>\n<p>python scripts\/init_config_root.py \u2013workspace<\/p>\n<p>automatically fills out workspace.yaml, runtime.local.yaml and material configurations that can be extrapolated, and lists the final configurations. runs later the program_profile.py to verify and freeze the project configuration\u3002<\/p>\n<p>when the input is a long article, you first call the adapt-longform-for-speech to generate the script and wait for my confirmation\u3002<\/p>\n<p>If any configuration or quality check fails, only the downstream phase is suspended. Please diagnose and repair the current failure layer and re-run the check until it passes. I am asked only if there is a lack of evidence, of the mandate, of what must be decided by me, or if there is no automatic repair\u3002<\/p>\n<p>after the oral presentation, the gentleman is a 60 to 90 seconds, audio, subtitle, visual, and live video test. run all quality checks, provide test video, cover and inspection reports, pending my confirmation\u3002<\/p>\n<p>Do not generate complete video without confirmation. The complete package and the final delivery package will continue after confirmation\u3002<\/p>\n<p><strong>2.2 Leave sufficient time for the first operation<\/strong><\/p>\n<p>Video production is much more important than normal documentation tasks. The first run is also to download dependencies, models and browser environments at a speed that does not depend only on the final rendering time\u3002<\/p>\n<p>My own two tests are very different\u3002<\/p>\n<ul>\n<li>Windows handles a long article that takes more than two hours from start-up to completion<\/li>\n<li>Linux generated a non-digital test for about 80 seconds, the first time taking about 21 minutes<\/li>\n<\/ul>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-55992\" title=\"bf4ffe13j00tjzv2z007d000i000rgp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/08\/bf4ffe13j00tjzv2z007dd000i000rgp.jpg\" alt=\"bf4ffe13j00tjzv2z007d000i000rgp\" width=\"648\" height=\"988\" \/><\/p>\n<p>These two figures represent only the machines and networks at the time, and the second will usually be much faster, as dependence, models and some intermediates already exist\u3002<\/p>\n<p><strong>2.3 What do you want to see after running<\/strong><\/p>\n<p>When Codex first stops, check in the following order\u3002<\/p>\n<ol>\n<li>Did you add or delete the key facts<\/li>\n<li>Was there any misreading, typographical, sudden velocity or apparent mechanical sense in the sound<\/li>\n<li>Does the subtitle follow the sound<\/li>\n<li>Did the screen cover the subtitles<\/li>\n<li>Does animated capture content<\/li>\n<li>Could the test footage represent the true style of the whole video<\/li>\n<\/ol>\n<p>Somewhere wrong, write the time and target state to Codex. Tell it 18 seconds too close, 32 seconds too fast. It would be more useful to point out clear issues than a vague concept\u3002<\/p>\n<p><strong>Chapter 3<\/strong><\/p>\n<p><strong>3.1 Digital human synthesis<\/strong><\/p>\n<p>The basic video is screened for content, sound, subtitles and images, and then the digital person is generated\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-55994\" title=\"69410052j00tjzv3o00hkd000v9000g3p\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/08\/69410052j00tjzv3o00hkd000v900g3p.jpg\" alt=\"69410052j00tjzv3o00hkd000v9000g3p\" width=\"1125\" height=\"579\" \/><\/p>\n<p>Duix. Avatar is an open-source digital person that can be deployed locally and can be called through API. In the official note, a reference video of about 10 seconds can be used to create a digital image and sound, followed by text or audio-driven mouth. The work stream of this paper takes only this part of the oral synthesis, and ultimately the sound continues to use the previously identified chords\u3002<\/p>\n<p>Note: Duix. Avatar's current README states that commercial licences are required for locally deployed open-source support, free of charge, for businesses with more than 100,000 users or more than $10 million per year\u3002<\/p>\n<p>The authorization is confirmed before any personal image or voice is used\u3002<\/p>\n<p><strong>3.2 Preparation of reference videos<\/strong><\/p>\n<p>To the extent possible, reference videos satisfy the following\u3002<\/p>\n<ul>\n<li>Man on camera<\/li>\n<li>Full face, clear and clear<\/li>\n<li>Light and background stability<\/li>\n<li>Not too much head and body<\/li>\n<li>There's only one key person in the picture<\/li>\n<li>The material does belong to you or has been authorized<\/li>\n<\/ul>\n<p>You can take a section from a public video that you've taken before, or a special one. Once the image is selected, the follow-up video recycles the character's movement, so that it is not too large to be seen\u3002<\/p>\n<p><strong>3.3 Generate silent oral video<\/strong><\/p>\n<p>Hand over the reference video and .wav` to Codex for local digital services\u3002<\/p>\n<p>Please use the locally configured Duix.avatar service to process the following two materials\u3002<\/p>\n<p>Reference Person Video<br \/>\nfinal embossed tape [project directory\/udio\/output\/narration_master.wav]<\/p>\n<p>only oral videos are generated that are fully consistent with the length of the tape. the output video remains silent and does not recompose the sound. checks if the image of the person is continuous, if the mouth is synchronized and if the output time is consistent with narration_master.wav\u3002<\/p>\n<p>The output of this step is a silent person video. The character has been tuned to the mother belt, but it has not been placed in the official picture\u3002<\/p>\n<p><strong>3.4 Synthesis of digitals into basic videos<\/strong><\/p>\n<p>both basic and silent audio videos are checked and are then called on to compose-avatar-video\u3002<\/p>\n<p>please use compose-avatar-video to synthesize silent digital person videos that are already good for mouth into basic video\u3002<\/p>\n<p>Basic Video<br \/>\nMute Digital Video<\/p>\n<p>Set the digital man into a circle of 300 x 300, put it in the lower left corner. Keep your face intact, without covering subtitles, titles and important information. Keeps the original audio of the base video and does not use the tracks in the digital video\u3002<\/p>\n<p>Post-synthetic check time, resolution, audio track, character tailoring, subtitle mask and front frame. Output check report and final video path\u3002<\/p>\n<p>The position, size and shape can be changed. There are only three that need to be kept\u3002<\/p>\n<ol>\n<li>People can't hide subtitles and key messages<\/li>\n<li>the final sound still comes from narration_master.wav<\/li>\n<li>Synthesizing and rechecking the sound drawings for length and mouth sync<\/li>\n<\/ol>\n<p><strong>Chapter 4<\/strong><\/p>\n<p><strong>4.1. Clarify the animation target<\/strong><\/p>\n<p>Codex can change the painting, but only if you're clear about the target\u3002<\/p>\n<p>The most effective types of information are five\u3002<\/p>\n<ul>\n<li>What scene<\/li>\n<li>Which element<\/li>\n<li>In the second<\/li>\n<li>From what state to what<\/li>\n<li>What must remain still<\/li>\n<\/ul>\n<p>4.2 Access to new illustrations Skill<\/p>\n<p>Animations and illustrations can continue to expand. For example, using open-source Ian Small Black Image to access the visual material phase, Skyll to make a fixed IP with its own authorized character material\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-55993\" title=\"18429511j00tjzv4n00atd000v9000edp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/08\/18429511j00tjzv4n00atd000v900edp.jpg\" alt=\"18429511j00tjzv4n00atd000v9000edp\" width=\"1125\" height=\"517\" \/><\/p>\n<p>Such changes allow Codex to analyse input and output\u3002<\/p>\n<p>Please analyze this next Skill directory, input, output and dependence\u3002<br \/>\nhttps:\/\/github.com\/helloianneo\/ian-xiaohei-illustrations\/tree\/main<\/p>\n<p>i want to get it into the visuals phase of angent-video-pideline. please indicate at which stage the call should be made, how the product should be written to the public-assets.json, how it should be reversed in case of failure, and how the neutral default value for the public water line should be avoided\u3002<\/p>\n<p>The list of access programmes and documents will be modified before I confirm it\u3002<\/p>\n<p>Make it anything. Don't mess with the timeline. In the end, the visual material will go back to the unified scene, subtitle and audio time frame\u3002<\/p>\n<p><strong>Chapter 5<\/strong><\/p>\n<p><strong>5.1 Record reusable content<\/strong><\/p>\n<p>When a video runs, the most valuable things tend to remain in the middle\u3002<\/p>\n<p>Where articles come from, how they can be translated into oral broadcasts, what styles fit for which platform, what animations work on, where readers can easily read, should be written back to the knowledge base\u3002<\/p>\n<p>I use Obsidian to control this part. At least the following are retained for each theme\u3002<\/p>\n<ul>\n<li>Original articles and sources<\/li>\n<li>Oral post-confirmation<\/li>\n<li>Video Item Path<\/li>\n<li>UsedProfile and Version<\/li>\n<li>Test and full film modification records<\/li>\n<li>Platform title, cover and publication<\/li>\n<li>Data and comment feedback after publication<\/li>\n<\/ul>\n<p>The next time you do the same, you do not have to start with a blank word. Reuse verified configurations and failed records before changing the period to different places\u3002<\/p>\n<p><strong>5.2 Reuse a multi-platform<\/strong><\/p>\n<p>The same article can be broken down into multiple products along a main line\u3002<\/p>\n<p>Long<br \/>\nIDEAS - X LONG<br \/>\nIdeas-Public Articles<br \/>\n3 to 5 short messages<br \/>\nIdeas-- 60 to 90 seconds<br \/>\nIdeas - 3 to 8 minutes of speaking video<br \/>\nIdeas - Short-Screen Video<br \/>\n\u2514-- VIDEO NUMBERS, TREMORS, B STATION AND LITTLE RED BOOK RELEASE<\/p>\n<p>EACH PLATFORM NEEDS A SEPARATE START, DURATION, PAINTING, SUBTITLE SECURITY ZONE AND CTA. CORE FACTS ARE SHARED AS MUCH AS POSSIBLE WITH THE MAIN LINE, AND EACH PLATFORM IS LESS LIKELY TO FIGHT EACH OTHER AFTER WRITING A SET\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-55996\" title=\"8d7205e1j00tjzv5800kzd000v90mzp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/08\/8d7205e1j00tjzv5800kzd000v900mzp.jpg\" alt=\"8d7205e1j00tjzv5800kzd000v90mzp\" width=\"1125\" height=\"827\" \/><\/p>\n<p><strong>5.3 How does this workflow materialize<\/strong><\/p>\n<p>Disbursement is the ultimate goal, and the following options can be realized\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-55995\" title=\"425 fdacfj00tjzv5i0073d000v9000dkp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/08\/425fdacfj00tjzv5i0073d000v900dkp.jpg\" alt=\"425 fdacfj00tjzv5i0073d000v9000dkp\" width=\"1125\" height=\"488\" \/><\/p>\n<p><strong>Final Thoughts<\/strong><\/p>\n<p>First run, step by step, run out of 60 to 90 seconds\u3002<\/p>\n<p>Keep your voice, voice, subtitles and basic images in order to keep them in your own area of work\u3002<\/p>\n<p>AND THEN WE ADD THE CHARACTER IP, AND WE FINALLY PICK UP THE LOCAL DIGITAL\u3002<\/p>\n<p>The first video could take hours to make it back to configuration, material and cache\u3002<\/p>\n<p>The more you do, the more the system is like your own production line\u3002<\/p>\n<p>Workflow open source address:<a href=\"https:\/\/github.com\/JayceHuang\/agent-video-pipeline\">this post is part of our special coverage syria protests 2011<\/a> (A video tutorial inside)<\/p>","protected":false},"excerpt":{"rendered":"<p>Local deployment: All digital population programmes seen before the zero-cost digital population programme were based on heygen, about $10 a minute, expensive. The programme is a local telecast digital programme, which requires only one Nvidia graphic card. Chapter 1 1.2 Without a NVIDIA graphic card, skip the local digital synthesis. Articles can continue to be published orally, audio, subtitles, illustrations, animations and video renderings. The digital person simply folds a good mouth into a basic video without affecting the production of the previous content. And you can choose the cloud-based digital program<\/p>","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[149,144],"tags":[165,8792,1252,2658],"collection":[],"class_list":["post-55989","post","type-post","status-publish","format-standard","hentry","category-jiaocheng","category-baike","tag-ai","tag-8792","tag-1252","tag-2658"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts\/55989","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/comments?post=55989"}],"version-history":[{"count":0,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts\/55989\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/media?parent=55989"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/categories?post=55989"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/tags?post=55989"},{"taxonomy":"collection","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/collection?post=55989"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}