{"id":56707,"date":"2026-09-03T17:35:28","date_gmt":"2026-09-03T09:35:28","guid":{"rendered":"https:\/\/www.1ai.net\/?p=56707"},"modified":"2026-09-03T17:35:28","modified_gmt":"2026-09-03T09:35:28","slug":"%e7%94%a8ai%e5%a4%8d%e5%88%bb%e6%8a%96%e9%9f%b3%e7%88%86%e6%ac%be%e8%a7%86%e9%a2%91%ef%bc%8ccodex-libtv-%e7%9a%84-ai-%e7%9f%ad%e8%a7%86%e9%a2%91%e5%88%9b%e4%bd%9c%e5%85%a8%e6%b5%81%e7%a8%8b","status":"publish","type":"post","link":"https:\/\/www.1ai.net\/en\/56707.html","title":{"rendered":"The full process of creating an AI short video with an AID retweeted sound boom video, Codex + LibTV"},"content":{"rendered":"<p>I've been brushing a fire on my voice<a href=\"https:\/\/www.1ai.net\/en\/tag\/%e5%8d%a1%e7%82%b9%e8%a7%86%e9%a2%91\" title=\"[Sees articles with [card video] labels]\" target=\"_blank\" >Card Video<\/a>Very rhythmic. I can't find the original\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56708\" title=\"69986 a05j00tks7ls004sd000ik00dqp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/69986a05j00tks7ls004sd000ik00dqp.jpg\" alt=\"69986 a05j00tks7ls004sd000ik00dqp\" width=\"668\" height=\"494\" \/><\/p>\n<p>And then you want to replace the characters with a similar video\u3002<\/p>\n<p>Much of the web failure is not that people are not good-looking, but rather that there is no right relationship between walking, graft, camera distance and blocking\u3002<\/p>\n<p>We only keep the action and space structure of the explosive video and replace it with our role:<strong>First, use Codex to abstract the action<a href=\"https:\/\/www.1ai.net\/en\/tag\/%e6%b7%b1%e5%ba%a6%e8%a7%86%e9%a2%91\" title=\"[Sees articles with [Deep Video] labels]\" target=\"_blank\" >Depth Video<\/a>, and finish the replacement in LibTV\u3002<\/strong><\/p>\n<p>Depth Video+Person Three View Workstream<\/p>\n<blockquote>\n<ul>\n<li><strong>Depth Video<\/strong>It is a type of video that contains in-depth information about the scenes, and by recording the distance from each pixel to the camera, achieves a more real, stereo visual experience than the traditional two-dimensional video\u3002<\/li>\n<li><strong>Number Three View<\/strong>It's through the complete display of the character's face, side, and three different perspectives on the back\u3002<\/li>\n<\/ul>\n<\/blockquote>\n<p>And my whole production process is based on deep video\u3002<\/p>\n<p>In-depth videos are for \"action and space\" and the third view of the person is for \"look and dress\", which is more stable than uploading the original video. However, it is not a universal key: the consistency of face and clothing still depends on the figure map, and complex fingers can also go wrong\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56710\" title=\"dc2937eaj00tks7mg00hgd000pv00ep\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/dc2937eaj00tks7mg00hgd000pv00eep.jpg\" alt=\"dc2937eaj00tks7mg00hgd000pv00ep\" width=\"931\" height=\"518\" \/><\/p>\n<p><strong>i. Allow Codex to generate deep video<\/strong><\/p>\n<p>I gave the original video to Codex, which is the only part of the core command:<\/p>\n<blockquote>\n<ul>\n<li>Please turn this into a deep video<\/li>\n<\/ul>\n<\/blockquote>\n<p>Codex read video parameters 6.25 seconds, 1080 x 1920, 60fps, 371 frames. The Python environment is then automatically created, with OpenCV, NumPy, ONNX Runtime DirectML, FFmpeg encoding and downloading about 99MB models. I don't need a handwritten code for the whole process. Just let us know\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56709\" title=\"036085e7j00tks7n50001ed000il00j1p\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/036085e7j00tks7n5001ed000il00j1p.jpg\" alt=\"036085e7j00tks7n50001ed000il00j1p\" width=\"669\" height=\"685\" \/><\/p>\n<p>And finally, we get the following deep video:<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56712\" title=\"33e66764j00tks7nx001dd000ih00dqp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/33e66764j00tks7nx001dd000ih00dqp.jpg\" alt=\"33e66764j00tks7nx001dd000ih00dqp\" width=\"665\" height=\"494\" \/><\/p>\n<p>Tips: In principle, you can use any Agent, whether it's WorkBuddy, ZCode, Trae or any other Agent, to generate deep video, mainly by calling your local programs and graphic cards\u3002<\/p>\n<p><strong>II. What algorithm does Codex use<\/strong><\/p>\n<p>The core model is the Vit-Small of Decth Anything V2, about 248 million parameters. It provides a single-spectrum relative depth estimate for each frame: the area close to the lens is whiter and darker at a distance\u3002<\/p>\n<p>A SIMPLE FRAME-BY-FRAME PROCESSING IS EASY TO GLITTER, SO THE SCRIPT TAKES THREE MORE STEPS: TO HARMONIZE THE GREYSCALE RANGE WITH THE 2%-98% FRACTION; TO SMOOTH THE DEPTH RANGE OF THE ADJACENT FRAME; AND TO GIVE PRIORITY TO THE CURRENT FRAME AT THE EDGE OF THE MOTION. ZOOM BACK TO THE ORIGINAL RESOLUTION AND OUTPUT BY H.264, CRF 10, NO AUDIO\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56711\" title=\"d8817bfj00tks7io00ogd000pz00eip\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/de8817bfj00tks7oi00ogd000pz00eip.jpg\" alt=\"d8817bfj00tks7io00ogd000pz00eip\" width=\"935\" height=\"522\" \/><\/p>\n<p><strong>III. Summary card requirements and actual time-consuming<\/strong><\/p>\n<p>Generating in-depth video has certain requirements for computer graphic cards\u3002<\/p>\n<p>My computer is RTX 3060 Laptop 6GB, the processor is i7-1270H, deduced from DirectML\u3002<\/p>\n<p>According to the file time record: first time to configure the environment, download the model to a piece of about 12 minutes; it took about 2 minutes 23 seconds to actually process the 371 frame\u3002<\/p>\n<p>Empirically, ViT-Small handles this 6-10 seconds short video, 4GB display can be tried, and 6GB is more adequate; to change a larger Base\/Large model or handle long videos, it is recommended above 8GB\u3002<\/p>\n<p>NO ONE CAN RUN CPU, JUST A LOT SLOWER. THESE ARE EMPIRICAL AND NOT A RIGID THRESHOLD FOR ALGORITHMS\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56713\" title=\"823d11daj00tks7p70jwd000p00efp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/823d11daj00tks7p700jwd000py00efp.jpg\" alt=\"823d11daj00tks7p70jwd000p00efp\" width=\"934\" height=\"519\" \/><\/p>\n<p><strong>IV. Generating a third view of the person<\/strong><\/p>\n<p>This is a relatively simple one, which provides a close image of the person, which is passed on to a large biographic model, allowing it to produce a three-view of the person, using the following instructions:<\/p>\n<blockquote>\n<ul>\n<li>reference is made to the image of the person who uploaded the picture, using a close-image and three-view reference layout with a clean and white background. on the left side is a close profile of the person and on the right is a third view of the person, with a full body view: a front view, a left view and a back view, all neutral positions with hands on the side. while the clothing is of a more general quality, the hairdressing and clothing adjustment is more or less simple, and the five-person costumes remain the same, with no other scenes, text, logo, etc. on the scene except for the person. the overall style is a clear, high-resolution print picture, not an illustration. each view produces a single picture, with a ratio of 9:16<\/li>\n<\/ul>\n<\/blockquote>\n<p>So we can get a third view of this character\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56714\" title=\"0b32059bj00tks7pq00bwd000pr00ehp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/0b32059bj00tks7pq00bwd000pr00ehp.jpg\" alt=\"0b32059bj00tks7pq00bwd000pr00ehp\" width=\"927\" height=\"521\" \/><\/p>\n<p><strong>V. Completion of replacement in LibTV<\/strong><\/p>\n<p>After the deep video was finished, we opened LibTV and uploaded it:<\/p>\n<ol>\n<li>Recent photos of the same person + three view pictures<\/li>\n<li>Handle the deep video\u3002<\/li>\n<\/ol>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56717\" title=\"12f03c7fj00tks7q600bd000pu00iop\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/12f03c7fj00tks7q600bdd000pu00iop.jpg\" alt=\"12f03c7fj00tks7q600bd000pu00iop\" width=\"930\" height=\"672\" \/><\/p>\n<p>A hint can be written directly:<\/p>\n<blockquote>\n<ul>\n<li>Using the black and white depth video provided as an in-depth guidance reference, strictly following the depth of space in the deep video, the attitude of the person, the moving trajectory of the camera, the image structure; replacing the characters in the deep video with the assigned character, the five officials and the costumes, removing the small black shoulder bag<\/li>\n<li>The structure of the deep video abstract is transformed into a short film of overwritten characters, high-profile, clean and seamless photography studios with clean background and ground, with no furniture, stools, wires, photographic equipment or other miscellaneous items. Reverting real environmental materials, natural light, person action, scene layout and the deep video provided, the image flowed unfailingly, the continuous matching of in-depth information remained unchanged, and the role image remained uniform\u3002<\/li>\n<\/ul>\n<\/blockquote>\n<p>As a result of this generation, the movement, moving and moving relationships of the original blast video were preserved and the person was replaced by a new role in the three views\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56715\" title=\"f360664ej00tks7qp004id000pr00dlp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/f360664ej00tks7qp004id000pr00dlp.jpg\" alt=\"f360664ej00tks7qp004id000pr00dlp\" width=\"927\" height=\"489\" \/><\/p>\n<p>VI. OTHER PROGRAMMES<\/p>\n<p>LibTV's been online lately, and if you don't have a graphic card, or if you feel the computer's production is slow, you can produce it fast, free\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56716\" title=\"f360664ej00tks7s6004id000pr00dlp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/f360664ej00tks7s6004id000pr00dlp.jpg\" alt=\"f360664ej00tks7s6004id000pr00dlp\" width=\"927\" height=\"489\" \/><\/p>\n<p>But I tested it, and it's not as accurate as my local graphic card, but it's fine\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56718\" title=\"707c08b2j00tks7sx0010d000ia008kp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/707c08b2j00tks7sx0010d000ia008kp.jpg\" alt=\"707c08b2j00tks7sx0010d000ia008kp\" width=\"658\" height=\"308\" \/><\/p>\n<p>The core of this workflow is the handling of action, slot and spatial relationships for explosive video, which can significantly increase the recalcitrantness of the person ' s replacement; if the core is a line, clip or emotional performance, its help is limited\u3002<\/p>","protected":false},"excerpt":{"rendered":"<p>Recently, a flash-fire card video was drawn on the shivering, which was very rhythmic. I can't find the original. And then you want to replace the characters with a similar video. Much of the web failure is not that people are not good-looking, but rather that there is no right relationship between walking, graft, camera distance and blocking. We only keep the action and the space structure of the explosive video, and then we replace it with our role: abstract the action into an in-depth video with Codex, then replace the person in LibTV. Depth Video + Personal Three View Workstream Depth Video is a video type that contains in-depth information on the scene, allowing for a more real, 3-dimensional visual experience than traditional two-dimensional video by recording the distance of each pixel to the camera. The third view of the character is by presenting the character's face, the side<\/p>","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[149,144],"tags":[956,8901,8902],"collection":[],"class_list":["post-56707","post","type-post","status-publish","format-standard","hentry","category-jiaocheng","category-baike","tag-ai","tag-8901","tag-8902"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts\/56707","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/comments?post=56707"}],"version-history":[{"count":0,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts\/56707\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/media?parent=56707"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/categories?post=56707"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/tags?post=56707"},{"taxonomy":"collection","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/collection?post=56707"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}