{"id":55296,"date":"2026-07-30T10:24:32","date_gmt":"2026-07-30T02:24:32","guid":{"rendered":"https:\/\/www.1ai.net\/?p=55296"},"modified":"2026-07-30T10:25:03","modified_gmt":"2026-07-30T02:25:03","slug":"%e8%ae%a9%e4%bd%a0%e6%88%90%e4%b8%ba%e4%b8%8d%e9%9c%b2%e8%84%b8%e7%9a%84%e7%9f%ad%e8%a7%86%e9%a2%91%e5%8d%9a%e4%b8%bb%ef%bc%8cheygen%e3%80%81minimax-%e4%b8%8e-codexai-%e7%9a%84%e7%bb%93%e5%90%88","status":"publish","type":"post","link":"https:\/\/www.1ai.net\/en\/55296.html","title":{"rendered":"You've become a faceless short video blogger, HeyGen, Mini Max, and Codexai's combined AID Population Production Workstream"},"content":{"rendered":"<p>&nbsp;<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-55300\" title=\"2efe8a8ej00tiyua90iqd000v90cip\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/07\/2efe8a8ej00tiyua900iqd000v900cip.jpg\" alt=\"2efe8a8ej00tiyua90iqd000v90cip\" width=\"1125\" height=\"450\" \/><\/p>\n<p>I sent one yesterday<a href=\"https:\/\/www.1ai.net\/en\/tag\/%e6%95%b0%e5%ad%97%e4%ba%ba\" title=\"[View articles tagged with [digital people]]\" target=\"_blank\" >Digital Human<\/a>SPEAK OF FDE<a href=\"https:\/\/www.1ai.net\/en\/tag\/%e5%8f%a3%e6%92%ad%e8%a7%86%e9%a2%91\" title=\"[View articles tagged with [oral video]]\" target=\"_blank\" >oral video<\/a>Many people's first reaction is not to ask me what tool I'm using, not to believe it's really digital. I made the whole digital population program a reusable one <a href=\"https:\/\/www.1ai.net\/en\/tag\/codex\" title=\"_Other Organiser\" target=\"_blank\" >Codex<\/a> Skill\u3002<\/p>\n<p>GitHub Repository: https:\/\/github.com\/Jingyi-Wu-Richael\/rachel-digital-human-protection<\/p>\n<p>My final work stream is:<\/p>\n<p>Material check <a href=\"https:\/\/www.1ai.net\/en\/tag\/minimax\" title=\"[View articles tagged with [MiniMax]]\" target=\"_blank\" >MiniMax<\/a> Generate sound <a href=\"https:\/\/www.1ai.net\/en\/tag\/heygen\" title=\"_Other Organiser\" target=\"_blank\" >HeyGen<\/a> 15 seconds of samples, \u2192 HeyGen full copy, \u2192 download, \u2192 status archive<\/p>\n<p>THIS PACKAGE IS SUITABLE FOR KNOWLEDGE BROADCASTING, SHORT VIDEO BULK PRODUCTION, COURSE DELIVERY, PRODUCT INTRODUCTION, MULTILINGUAL CONTENT AND FIXED IP DIGITAL ACCOUNTS\u3002<\/p>\n<p><strong>I. How to divide the whole chain<\/strong><\/p>\n<p>At the heart of the package is the sound and image being removed\u3002<\/p>\n<p>Mini Max: Responsible for voice cloning and TTS\u3002<\/p>\n<p>HeyGen: The person in charge is like a driver, mouth, face and action\u3002<\/p>\n<p>Codex: Responsible for inspection of materials, call processes, recording of status, risk control\u3002<\/p>\n<p>The complete link is:<\/p>\n<p>Preparation of a live recording, using Mini Max overseas cloned sound\u3002<\/p>\n<p>2. Hand over the file to Mini Max TTS to produce a cloned sound MP3\u3002<\/p>\n<p>Prepare a clear and positive image\u3002<\/p>\n<p>4. Upload the image and the MP3 generated by Mini Max to HeyGen\u3002<\/p>\n<p>5. Mr. 15 seconds to make a sample, confirm and then generate a complete video\u3002<\/p>\n<p>The easiest mistake here is that the voice has been cloned by MiniMax and has been reprogrammed by HeyGen\u3002<\/p>\n<p>If the goal is to preserve MiniMax 's cloned sound, the audio generated by MiniMax should be uploaded to HeyGen, using this audio-driven image. Do not cut back the `script + voice_id ' scheme of HeyGen, otherwise sound likeness will be covered\u3002<\/p>\n<p><strong>Two, HeyGen and Minimax<\/strong><\/p>\n<p>For the first operation, it is proposed to use `Image-to-Video ' directly, just api key, and a 2-min video consumes approximately two knives\u3002<\/p>\n<p>It is light: Avatar is not trained at the outset, but is able to quickly see the availability of the mouth, expression, face stability and image by uploading an image and an audio\u3002<\/p>\n<p>My suggestions:<\/p>\n<p>Single test: Use Image-to-Video first\u3002<\/p>\n<p>Long-term production of fixed account numbers: Photo Avatar\u3002<\/p>\n<p>Don't rush training until it's certified. Avatar, the lighter the test, the better\u3002<\/p>\n<p>And again, Mini Max, my advice:<\/p>\n<p>First test: clone sound with an overseas version of Voice Clane and generate MP3 with TTS\u3002<\/p>\n<p>`speech-2.8-hd '\u3002<\/p>\n<p>search speed and bulk preview: the first run test can be used for `speech-2.8-turbo ' \u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-55299\" title=\"7857501ej00tiyubc0086d000v9000h3p\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/07\/7857501ej00tiyubc0086d000v900h3p.jpg\" alt=\"7857501ej00tiyubc0086d000v9000h3p\" width=\"1125\" height=\"615\" \/><\/p>\n<p><strong>Preparation of materials: no complexity, but clean<\/strong><\/p>\n<p>Sound cloning was not working well, and often it was not a model that failed, but a sample that was not clean\u3002<\/p>\n<p>In the current official demand, Mini Max, sound samples need to be met:<\/p>\n<p>FORMAT: MP3, M4A OR WAV\u3002<\/p>\n<p>Length: at least 10 seconds and at most 5 minutes\u3002<\/p>\n<p>FILE SIZE: NO MORE THAN 20 MB\u3002<\/p>\n<p>Optional alert audio: less than 8 seconds and to provide text that fully corresponds to the audio\u3002<\/p>\n<p>I suggest 30 to 90 seconds: single, backgroundless, non-obvious mix and current, volume stable, speed of speech close to the daily mouth, not to overnomic, mutation or compression of severe second-hand audio\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-55298\" title=\"dd9a9110j00tyuc7007id000v9000q0p\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/07\/dd9a9110j00tiyuc7007id000v900q0p.jpg\" alt=\"dd9a9110j00tyuc7007id000v9000q0p\" width=\"1125\" height=\"936\" \/><\/p>\n<p>Same thing with people. Clean up:<\/p>\n<p>Face-to-face, five officials with no cover\u3002<\/p>\n<p>SHOW YOUR HEAD, SHOULDER AND UPPER BODY, AND THE CHARACTER IS ON THE SCENE ABOUT 50% TO 70%\u3002<\/p>\n<p>The mouth is clear and not covered by hair, hands, filters\u3002<\/p>\n<p>The vertical short video is preferred to a 9:16 or close to 9:16\u3002<\/p>\n<p>HeyGen Assets API supports PNG, JPEG pictures; the upper limit for flyer files on general material is 32 MB\u3002<\/p>\n<p>The clearer the mouth, the more stable the mouth; the cleaner the sound, the more like it is\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-55297\" title=\"a4ad99f9j00tiyu and 0064d000v9000wp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/07\/a4ad99f9j00tiyucm0064d000v900owp.jpg\" alt=\"a4ad99f9j00tiyu and 0064d000v9000wp\" width=\"1125\" height=\"896\" \/><\/p>\n<p><strong>IV. HOW TO PUT IT IN CODEX<\/strong><\/p>\n<p>I'll put every digital project in the same directory structure:<\/p>\n<p>project\/<br \/>\ninputs\/<br \/>\nideas -portrait.jpg<br \/>\nimage-source.mp3<br \/>\nscript.md<br \/>\n_work\/<br \/>\nimageover-full.mp3<br \/>\n-preview-15s.mp3<br \/>\n_other organiser<br \/>\n\u2514-outputs\/<br \/>\n- preview-15s.mp4<br \/>\n\u2013 final-1080p.mp4<\/p>\n<p>And then to Codex said:<\/p>\n<p>\u201c`text<\/p>\n<p>Please use Mini Max for overseas voice cloning and HeyGen API for digital population\u3002<\/p>\n<p>Input:<\/p>\n<p>documentation: inputs\/script.md<\/p>\n<p>image: inputs\/portrait.jpg<\/p>\n<p>sound samples: inputs\/voice-source.mp3<\/p>\n<p>Mini Max Key from Environmental Variable MINIMAX_API_KEY<\/p>\n<p>HeyGen Key from Environment Variable HEYGEN_API_KEY<\/p>\n<p>Process:<\/p>\n<p>1. Check material first\u3002<\/p>\n<p>2. Mr. 15 seconds\u3002<\/p>\n<p>3. The full version of the sample will be generated after confirmation\u3002<\/p>\n<p>4. All tasks ID are written to work\/job-state.json\u3002<\/p>\n<p>Do not display the full key in the log, script or document\u3002<\/p>\n<p>&#8220;`<\/p>\n<p>This step focuses on standardization\u3002<\/p>\n<p>You don't have to believe you remember every time. You just have to put the right process in Skill and get Codex to do it every time\u3002<\/p>\n<p><strong>V. This Skill fixed something<\/strong><\/p>\n<p>Warehouse address:<\/p>\n<p>https:\/\/github.com\/Jingyi-Wu-Richael\/rachel-digital-human-production<\/p>\n<p>Call:<\/p>\n<p>\u201c`text<\/p>\n<p>This video is made using $rachel-digital-human-producation, with a 15-second sample\u3002<\/p>\n<p>&#8220;`<\/p>\n<p>It fixed:<\/p>\n<p>Fixed directory structure\u3002<\/p>\n<p>Check first the case, the image and the sound samples\u3002<\/p>\n<p>Mini Max is in charge of the voice, Heygen is in charge of the video\u3002<\/p>\n<p>Mr. Forever becomes 15 seconds of a sample\u3002<\/p>\n<p>The sample is not explicitly confirmed and does not produce a full version\u3002<\/p>\n<p>each step records `voice_id ' , `asset_id ' , `video_id ' and state\u3002<\/p>\n<p>Failure is not blindly retried to avoid double deductions\u3002<\/p>\n<p>EVENTUALLY MP4 MUST BE FULLY CHECKED\u3002<\/p>\n<p>Do not write API Key, authorized header, temporary download links to a log or status file\u3002<\/p>\n<p><strong>vi. little tips<\/strong><\/p>\n<p>Digital video is the easiest place to step on\u3002<\/p>\n<p>Common problems include: sound that does not match itself, Chinese accent, slight face deformation, strange mouth angles and jaw movements, irregular blinking and shoulder movements, and inadequate configurations for short video platforms\u3002<\/p>\n<p>So I wrote \"15 seconds sample\" as a mandatory step. 15 seconds to confirm sound, mouth, image and image. Make sure it's okay. Run the whole video again\u3002<\/p>\n<p>Related Links:<\/p>\n<p>GitHub Repository: https:\/\/github.com\/Jingyi-Wu-Richael\/rachel-digital-human-protection<\/p>\n<p>Mini Max Voice Cline Official Document: https:\/\/platform.minimax.io\/docs\/guides\/speech-voice-cline<\/p>\n<p>HeyGen Image to Video Official Document: https:\/\/developmenters.heygen.com\/image-to-video<\/p>\n<p>HeyGen Assets Official Document: https:\/\/developmenters.heygen.com\/assets<\/p>","protected":false},"excerpt":{"rendered":"<p>&nbsp; Yesterday, a digital video was posted about FDE, and many people's first reaction was not to ask me what tool I used, not to believe it was really digital. I made the whole digital population stream a reusable Codex Skill. GitHub Repository: https:\/\/github.com\/Jingyi-Wu-Richael\/rachel-digital-human-prodition: material check \u2192 MiniMax produces \u2192 HeyGen 15 seconds sample manual confirmation \u2192 HeyGen full copy \u2192 download test \u2192 state archive<\/p>","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[149,144],"tags":[5321,6623,1025,1490,2883,1252],"collection":[],"class_list":["post-55296","post","type-post","status-publish","format-standard","hentry","category-jiaocheng","category-baike","tag-ai","tag-codex","tag-heygen","tag-minimax","tag-2883","tag-1252"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts\/55296","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/comments?post=55296"}],"version-history":[{"count":0,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts\/55296\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/media?parent=55296"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/categories?post=55296"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/tags?post=55296"},{"taxonomy":"collection","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/collection?post=55296"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}