
GENERATE FIRST FROM CLOUD APIDigital HumanSpeaking of
THIS ARTICLE WANTS TO SHARE SOMETHING I JUST RAN THROUGH. CONNECT TO THE API OF SOME CLOUD SERVICESBean curdFrom now on, give it a short video link, a voice of its own and a well-trained digital image, so that the bean bag can continue to produce the script, audio and..Digital Human Video, and lastly save the piece to the specified folder。
I'll write this whole thing down today from the perspective of bean buns. API is used to extract files, Mini Max is responsible for the cloning of sound and the generation of audio, and mirrors are responsible for opening up digital people. The bean bag is responsible for calling these interfaces, organizing intermediate files and checking whether each step is actually completed。
The division of labour among the four platforms in the process set out below。
|
platform |
What's in this process |
|---|---|
|
API |
Turn short video sharing links into text |
|
MiniMax |
Turn words into their own cloned voices |
|
Cicada Mirror |
Let the digitals talk to this voice |
|
Bean curd |
Call the first three platforms and connect files, audio and video |
The process is most easily contained between platforms. The file, audio and digital video are scattered in three services, and the complete task is stopped in the middle of any interface where the document or parameter is erroneously filled. In order to facilitate positioning issues, three platforms are registered, their respective vouchers are obtained, and digital persons and sound are prepared. The link transfer, text transfer sound and audio-driven digital persons were then tested, all of which were passed before the legumes were produced in full。
The operation follows the following order。
- The platform is ready to document, document, material, three single measurements, complete pieces, manual acceptance
I finally ran a 10.24 seconds, 1080 x 1920 stand-up digital video. The whole job was carried out by the bean bag, and the real API Key did not appear in the alerts, articles, Skill files and results catalogue. Time-consuming, document size and error reporting for later articles are also derived from this test。

Text Guide
Zero-one recognizes the whole process
First look at what's in charge, what's in charge, what's in charge, what's in place, what's in the account and what's in place。
02 | Complete platform preparation
REGISTER THREE CLOUD-END PLATFORMS IN TURN, FIND INTERFACE DESCRIPTIONS, CREATE API VOUCHERS, AND PREPARE THEIR OWN CLONED SOUND AND DIGITAL IMAGE。
03-Opium Bean Packing Project
Three Skills were created in a stand-alone project, six messages were entered into the security configuration and the environment was confirmed to be available through a non-costed read-only pre-test。
04 is going from single to one
Tests of linked transliteration, text-to-sound and audio-driven digital people. After the adoption of the three, the bean bag will read the official text and generate a full video。
05 accepted and miscalculated
Opens the result folder to check the file, audio and video, retrys the failure and decides whether to enter the batch。
I. A first reading of the entire digital line
1. What is the responsibility of each of the three platforms
It would be easier to understand the division of labour between these three services if the digital population were conceived as a live anchor。
Ji-Ling is in charge of “Looking for the script”. You give it the shared link to the public video, which returns the title, subtitle or text. The only material obtained here is raw material, which does not amount to a new text that can be published directly。
Mini Max is in charge of talking. You create a cloned sound and then submit the confirmed oral text to the TTS interface, which generates MP3 audio。
THE MIRROR IS FOR "SHOW." YOU CREATE AND TRAIN YOUR OWN DIGITAL IMAGE, THEN YOU UPLOAD MP3 TO IT, AND THE PLATFORM ALLOWS DIGITAL PEOPLE TO FOLLOW THEIR MOUTHS, FACES AND MOVEMENTS, AND EVENTUALLY RETURN TO MP4。
The bean bag is responsible for connecting the three platforms. It reads links, calls Skill, saves intermediate files, continues to wait for a walk-in mission and downloads it into a film after the mission has been completed. The success of each step is judged by the acceptance conditions in the hint。
2. What needs to be prepared before we begin
These materials will be ready prior to formalization。
- A public access short video-sharing link
- A clean recording for the cloning of sound
- A live video for training digital figures
- API Key
- Mini Max API Key and Voice ID
- App ID, Secret Key and Digital ID
- A folder dedicated to this product
- A bean bag client and an independent soy bag digital person project。
Record and train video as much as possible. The recording is to reduce background music, echo and talk to others; the training video is to remain face-to-face, the light is stable and the picture is continuous. The material is of poor quality, and the best models can only magnify the problem。
3. Speed-up of the Platform ' s web site
The sites actually used are below. Log in and log in, then back to work。
|
platform |
Official entrance |
Use in this curriculum |
|---|---|---|
|
Bean curd |
Create Skill, perform testing and link complete processes |
|
|
API |
Open short video file extraction, create API Key |
|
|
Mini Max Open Platform |
Cloning sound, acquiring Voice ID and API Key |
|
|
Open Mirror Platform |
CREATE DIGITAL PEOPLE, GET API VOUCHERS AND GENERATE VIDEOS |
If you need to check the interface parameters, you can look directly at MiniMax, MiniMax Synchronized Speech Synthesis, Spectrum AccessToken, and Mirror Synthetic Digital Video. The platform page may be recast later; if the link is opened for login, you can search for it by the menu name given here。
4. Security and authorized borders
API Key corresponds to a paid pass for a third-party platform. Do not post it into regular conversations or shared items, or SKILL.md, scripts, articles or Git warehouses in Skill。
This tutorial records only the configuration entry name and not the true value。
ZHILING_API_KEY
MINIMAX_API_KEY
MINIMAX_VOICE_ID
CANJING_APP_ID
CANJING_SECRET_KEY
CANJING_PERSON_ID
The main course is to obtain vouchers from each of the three platforms and then fill in the real values for the security configuration of the soybag project。
Saves the bean bag only to check if the configuration exists and the format is reasonable, and does not allow it to print the whole content. Only variable names are written in the normal conversation, and the real values only appear in the security configuration area。
THREE TESTS CALL FOR REAL INTERFACES, AND DIGITAL GENERATION USUALLY INCURS A SMALL COST. FOR THE FIRST TIME, “AS-NEEDED CONFIRMATION” IS MAINTAINED AND THE RANGE IS CONFIRMED AND CONTINUED WHEN AN EXTERNAL API, UPLOADING AUDIO OR STARTING VIDEO SYNTHESIS IS SEEN。
ii. Prepare to get the short video file first
1. Registration and access to the control desk
OPENS THE API USER CONSOLE TO REGISTER AND LOG IN AS PROVIDED ON THE PAGE. AFTER ENTERING THE CONSOLE, A LIST OF PRODUCTS IS FOUND CONFIRMING THAT SHORT VIDEO RESOLUTION, FILE EXTRACTION OR VOICE-TO-TEXT-RELATED PRODUCTS HAVE APPEARED UNDER ACCOUNT NUMBERS。
The interface name may be adjusted later, and the judgement is simple. This interface receives short video sharing links and returns readable text results。

2. Finding a file extraction interface
Click on the file to extract the product, read the interface description and do not copy the example code immediately. Four things in the file first。
- What is the address of the request
- THE REQUEST IS GET OR POST
- (a) The parameters in which the shared link is placed
- The task is to return simultaneously or to continue the practice after submission。
Short video clips are often taken in a different process. The first request will only return the task number and in a few seconds the task status will be checked. The bean bag, Skill, will be in charge of the inquiry. You don't have to refresh it manually. When the interface shows “successful submission”, the file is usually not generated。

3. Access and safe storage of API Key
Go back to the console and open the Key Management. Click to create if the account number does not have a key; do not create again if a key is already available。
ZHILING_API_KEY. THE FULL VALUE SHOULD NOT APPEAR IN ARTICLES, SCREENSHOTS AND BEAN BAG RESPONSES. SOME PLATFORMS ONLY DISPLAY A COMPLETE KEY ONCE THEY ARE CREATED, CONFIRMING THAT THEY HAVE BEEN SAFELY STORED BEFORE CLOSING THE BULLET WINDOW。

4. Write down interface rules and not call them for the time being
Prepare an open short video-sharing link, not a long 10-minute video for the first time. At this point, it is only confirmed that the product is open, that the interface can receive shared links and that the task query is recorded。
The real interface test is in chapter VI. Once the Skill, voucher and output directories in the bean bag are ready, a minimum test will be performed. This makes it possible to judge whether the problem arises from the platform interface or from the bean bag configuration when there is an error。
iii. Prepare Mini Max to clone and call its own voice
1. Confirmation of sites and accounts upon registration
Enter Mini Max Open Platform and register by page. Domestic accounts use open domestic platforms to the extent possible and do not mix key, sound and account balances created at different sites。
AFTER LOGIN, CHECK IF " VOICE " " SOUND CLONING " AND " API KEY " CAN BE SEEN IN THE CONSOLE. SOUND CLONING REQUIRES REAL NAME CERTIFICATION FOR INDIVIDUALS OR BUSINESSES. IF AN ENTRY POINT IS NOT SHOWN, CHECK THE AUTHENTICATION STATUS, ACCOUNT PRIVILEGES AND WHETHER THE PRODUCT IS OPERATIONAL。
2. Preparation of audio recordings and creation of cloned sound
Enter “sound cloning”, upload or directly record a clean human voice. According to the current interface of MiniMax, the clone audio should be MP3, M4A or WAV, with a duration of between 10 seconds and 5 minutes, and the file should not exceed 20MB. The following should be noted when recording。
- Just one person to talk to。
- Don't play background music。
- (b) Avoiding visible spraying of wheat, echoes and electric currents。
- Use normal speech speed and natural emotions。
- The content is as continuous as possible and less fusion of debris。
A name that can be identified with the sound before submission, such as “My Daily Oral”. Once the platform has been processed, it is tested on the web site to confirm that there is no apparent mechanized, transliterated or volume mutation。

3. Voice ID found and recorded
Turn on "My Sound" and find the sound you just created and tested through. Sound names are easy to confirm. API calls for Voice ID。
Save Voice ID to MINIMAX_VOICE_ID, do not write directly to Skill text. If there are more than one sound, the bean bag should be chosen by name or in the order specified by you, not at random。

4. Create and save API Key
Enter the "API key " , click to create a new key and fill out a name for an identifiable use, such as "doubao-digital-human". Created to copy to the security configuration MINIMAX_API_KEY。
Key and Voice ID have different roles. Key is responsible for monitoring, Voice ID is responsible for designating sound. When 401, without permission or sound color is not available, check the account number and site and confirm whether the two fields belong to the same MiniMax account。

5. TEST SOUND, DO NOT CALL API FOR THE TIME BEING
Try the newly created sound on MiniMax and confirm that there is no apparent mechanical sound, type word or volume mutation. The API individual test is placed in Chapter VI, at which time only three to five seconds short sentences are generated。
iv. Preparation of mirrors to create their own image of digital people
1. REGISTER AND ACCESS THE API ACCESS PAGE
OPENS THE MIRROR API ACCESS PAGE AND LOGS INTO THE "API ACCESS " OR OPEN PLATFORM PAGE. THE BEAN BAG IS GOING TO CALL THE MIRROR FROM THIS ENTRANCE, AND THE REGULAR WEB PAGE INTO A FILM EDITOR CANNOT ACCESS THE MISSION。
Checks whether the account is open for API and whether there are limits to the amount available to confirm that the access points for App ID, Secret Key, Digital Person List and Task Query are available. The platform package, billing and available models vary, with the actual cost based on the pre-generation page tips。

2. CREATION OF API VOUCHERS AND RECOGNITION OF THE MEANS OF ASSURANCE
KEY AREA FOUND ON API ACCESS PAGE. THERE'LL BE TWO VOUCHERS HERE。
|
Configure Name |
use |
|---|---|
|
|
Apply or account identification |
|
|
API Key for signature or authentication |
Both must come from the same application. After copying, insert the security configuration, do not intercept the complete key, and do not allow the bean bag to print out all the requests when you are wrong。
AccessToken documents according to a mirror will not be directly used by the business interface. Skill needs to request AccessToken first with App ID and Secret Key, and then put Token at the head of the subsequent interface. AccessToken's default is valid for one day, after retake, the old Token will expire immediately. So do not save AccessToken as a long-term configuration, and do not allow multiple tasks to be refreshed at the same time。

Create your own digital person
Enter a "digital person" or "my digital person", click to create. The platform provides, for example, “video-breed digital person” and “Vinson digital person” and, in order to use their own image of the person, choose “videobreed digital person” and upload the material they are entitled to use。

The training video meets the following conditions as far as possible。
- (a) Placing material, not too small for the subject
- (a) The person ' s face is intact and unshielded and the view is as visible as possible
- The light and background remain stable
- Do not cut the camera, turn around or suddenly leave the picture
- (b) Be visible in the mouth and avoid hand, mask or cup cover
- Only the images and voices of people who have their mandate are uploaded。
Pending completion of Platform training after submission. Do not continue to call the video interface while the status is still in process。
4. FIND DIGITAL ID
WHEN THE TRAINING IS COMPLETED, OPEN THE DETAILS OF THE DIGITAL PERSON, CONFIRM THAT THE PREVIEW, THE NAME AND THE CONFIGURATION ARE CORRECT, RECORD ITS DIGITAL PERSON ID AND KEEP IT AS CHANJING_PERSON_ID。
HERE'S THE NAME AND ID. THE NAME IS EASY TO CONFIRM THAT API USUALLY RELIES ON ID. IN ORDER TO AVOID LEAKAGE OR MISUSE, ONLY THE NAME IS SHOWN AND THE COMPLETE ID IS NOT DISPLAYED。
5. Recording of audio-drive requirements and harmonization of tests later
The mirror cloud interface cannot directly read the local path on the computer. When chapter VI starts testing, Skill uploads the MP3 generated by MiniMax to mirror file management, waiting for the file to be available, and then hands back the file ID to the digital video interface。
Using mirrors to synthesize digital video files, 8000 Hz or 16000 Hz can be uploaded on a single-channel file if you want the video to be subtitled. The program uses 16000 Hz, Monophonic MP3. The basic model is selected for testing, with 1080 x 1920, 1080p and audio only 3 to 5 seconds. Digital video is a walker and continues to be checked and downloaded after successful submission。
V. Establishment of a digital person project in bean bags
1. New directories of stand-alone projects and results
Opens a bean bag, creates a new soy bag digital person on the left, and prepares a separate output folder for the first exercise. Original audio recordings, training videos, test files and official films should not be mixed with desktops or download the base of the directory。
It is recommended that each intermediate result be numbered。
01-zhiling-copy.txt 02-minimax-voice.mp3 03-chanjing-test.mp4 04-final-script.txt 05-final-voice.mp3 06-final-digital-human.mp4
Order numbers are easily miscalculated. When you look at the last file name, you know where the mission stopped。
2. Three independent bean buns, Skyll
There is no need to download compression packages in advance or to prepare complex codes. Let the bean bag read the interfaces of the three platforms first, and then the three APIs become Skill, one thing. The interface addresses and parameters are based on the platform document and are stopped when unwritten fields are encountered, so that the beans are not subject to empirical guessing。
|
Skill |
enter |
Output |
Configuration to read |
|---|---|---|---|
|
Take it out in the Jigsaw case |
Open video-sharing links |
TXT FILE |
|
|
Mini Max Text-to-Speech |
Text, model, speed of speech |
MP3 AUDIO |
|
|
Mirror Digital Man |
MP3 AUDIO, IMAGE SPECIFICATION |
MP4 VIDEO |
|
Three Skills with the same set of rules. The certificate is only read from the security configuration and is not written in Skill. The input and output path should be clear, and the staggered task will only rotate with the same task number. When the interface returns successfully, the file is checked, and the error is kept and stopped at the current step if the interface fails。
Copy the following hint, and the bean bag will create three Skills in the current project along this lines。
Please create three separate local Skill:
1. JIGGLING: RECEIVING PUBLIC VIDEO-SHARING LINKS, CALLING THE JIGSAW INTERFACE AND EXPORTING PURE TEXT TXT。
2. Mini Max Text-to-Speech: Receive text, model and speed, call the specified cloned sound, output MP3。
3. MIRROR DIGITAL PERSON: RECEIVES LOCAL MP3 AND VIDEO SPECIFICATIONS AND GENERATES AND DOWNLOADS MP4 FROM DESIGNATED DIGITAL PERSONS。
Common demands:
– Read the current interface documents of the three platforms before preparing the call logic; stop and explain the interface address, request parameters or signature rules when they are not clear。
– CERTIFICATES CAN ONLY BE READ FROM A SECURE CONFIGURATION, WITH VARIABLES ZHILING_API_KEY, MINIMAX_API_KEY, MINIMAX_VOICE_ID, CHANJING_APP_ID, CHANJING_SECRET_KEY, CHANJING_PERSON_ID。
– No certificate or complete ID may be written to Skill, script, log or reply。
– Each Skill is responsible for only one capability and writes down input parameters, output paths, interface failure information and acceptance methods。
– A different task must be kept with a task number and the same task should be rotated and duplicates prohibited。
– Mini Max is allowed to specify a sampling rate and acoustics when exporting MP3, and the tutorial defaults 16000 Hz, one channel。
– Skill first uses App ID and Secret Key to get AccessToken. Token expires or returns 10400 only once, refreshing and then continuing the task without writing Token into configuration for a long time。
– When Skill receives a local MP3, upload the file and wait until the file is available before submitting a digital video job. Sets the source =1。
– DO NOT DIRECTLY OVERWRITE THE OUTPUT FILE WHEN IT ALREADY EXISTS; STATE FIRST; STOP WHEN THE API CALL FAILS AND DO NOT AUTOMATICALLY CHANGE TO ANOTHER PLATFORM。
Creates only static checks after they are finished and does not call any paid API. The table ends with the three Skill names, input, output, required configuration and check results。
3. API VOUCHERS FOR THREE PLATFORMS PROVIDED TO BEAN BUNS
When you have completed the registration and configuration of the first three platforms, you should have six messages. API Key, Mini Max provides API Key and Voice ID, mirrors Apple ID, Secret Key and Digital ID. They all have to be copied from your own platform account, not using examples of articles or other people's vouchers。
|
platform |
The configuration of the bean bag required |
|---|---|
|
API |
|
|
MiniMax |
|
|
Cicada Mirror |
|
Make sure that the current bean bag is accessible only to you, and then open the configuration area of three Skills to fill the corresponding fields with six real values. The true document is only filled in in the secure configuration area and is not pasted to the following hints, regular conversations, articles or screenshots。
After the six configurations have been completed, the following check hint is copied。
PLEASE CHECK ONLY THE SECURITY CONFIGURATION OF THE CURRENT SOYBAG DIGITAL PERSON AND DO NOT CALL ANY EXTERNAL OR PAID API IN THIS ROUND。
ZHILING_API_KEY, MINIMAX_API_KEY, MINIMAX_VOICE_ID, CHANJING_APP_ID, CHANJING_SECRET_KEY, CHANJING_PERSON_ID。
No real value may be read, displayed, repeated or cut off. Do not modify the configuration or automatically create or replace a voucher. Only variable names and problems are reported when the missing item or the format is clearly abnormal。
When completed, report in a form whether the six configurations are ready and whether the subsequent individual test conditions have been met。

Let the bean bag check the environment first
Before being formally called, the bean bag will be pre-screened at no cost. The following paragraph can be reproduced directly。
PLEASE IMPLEMENT A READ-ONLY PRE-SCREENING OF THE CURRENT SOYBAG DIGITAL PERSON PROJECT, WHICH IS PROHIBITED FROM CALLING ANY EXTERNAL OR PAID API。
In turn:
1. Three Skills exist and can be loaded
2. Whether the six security configurations exist but do not read or display real values
3. Availability of the required operational dependencies
4. output catalogues exist and can be written
5. whether output/01-zhiling-copy.txt, output/02-minimax-voice.mp3 and output/03-chanjing-test.mp4 have the same name files to avoid miscovering。
Please use the table to output " Check Items, Results, Pending Issues " . If any one of them fails to pass, the problem is clearly identified and stopped; when all is adopted, only the reply “may begin with three single tests”。
Only three Skills, configurations and output directories are normal for the next step。
VI. Testing three capabilities on a case-by-case basis
FOR THE FIRST TIME, NO LINKS, REWRITES, TTS, AND DIGITALS WILL BE HANDED OVER AT ONCE. THE WHOLE JOB WENT WRONG, AND IT'S HARD TO JUDGE WHETHER THE LINK IS INVALID, THE SOUND IS DEAD, THE KEY IS WRONG OR THE MIRROR IS STILL IN LINE. ONLY ONE INPUT AND ONE OUTPUT ARE MEASURED AT A TIME, WHICH IS THE FASTEST。
Test one. Can the link become a clean file
the input is an open shared link and the output is output/01-zhiling-copy.txt. copys the following hint, replacing the shared link only。
Please carry out a single test of the "Whispering" case。
Input:
– Open video-sharing link: [Pick shared link here]
– output file: unput/01-zhiling-copy.txt
Implementation requirements:
1. USE ZHILING_API_KEY IN THE SECURITY CONFIGURATION TO CALL FOR THE SMARTPHONE。
2. ONLY THE ORIGINAL TEXT OF THE MESSAGE OR SUBTITLES IN THE VIDEO WILL BE RETAINED, AND THE ENTIRE INTERFACE WILL NOT BE REWRITED, SUMMARIZED OR PRESERVED, JSON。
3. If the interface returns the task number, keep searching for the same task until it is successful or clearly fails。
4. If a temporary network error is encountered, a maximum of automatic retry after several seconds. Stop when it fails again and do not continue to create mandates。
5. API Key shall not be shown in the response, log or output file。
Checks that the output file exists after completion and is not empty. Only reports: execution results, number of body characters, time-consuming and file path。
Checks that the file is not empty, readable, and does not use JSON, the interface returned, as the main text. The first and second independent tests I met. After removing the influence of the machine's agent, I manually initiated the third test, which succeeded in obtaining 363 characters in 5.4 seconds, file size 955 bytes. The third time here is a new test after manual rehearsing, which is not an automatic retest allowed by a hint。
If there's a temporary network reset, make sure that the service is available, and then only deal with the one thing, not with Mini Max and the mirror。
Test two. Can the case be a sound-usable one
enter is a fixed short text with output output/02-minimax-voice.mp3. the following paragraph can be reproduced directly。
Please perform a single "Mini Max Text-to-Voice" test。
Input:
– Test: Hello, this is my first digital voice test。
– output file: output/02-minimax-voice.mp3
Implementation requirements:
1. READ MINIMAX_API_KEY AND MINIMAX_VOICE_ID, USING THE SPECIFIED CLONED SOUND。
2. The model uses speech-2.8-turbo, which is set at 1.0, output 16000 Hz, mono-channel MP3。
3. Read the test file as it is, without amplification, without opening remarks and with one voice interface。
4. API Key and full Voice ID shall not be shown in the response, log or output file。
UPON COMPLETION, CONFIRM THAT MP3 CAN BE PLAYED OVER A PERIOD OF 2 SECONDS WITHOUT A SIGNIFICANT CUT-OFF, BUT ONLY REPORT: EXECUTION RESULTS, AUDIO LENGTH, FILE SIZE, TIME-CONSUMING AND FILE PATH。
Plays it in person at the time of acceptance, confirming that the sound is correct, that the sound is complete and that it reads the correct keyword. This live sound is 4.05 seconds, 67,956 bytes, approximately 1.4 seconds。
The long case will be tested after the short sentence is passed. Because long audio is more expensive, it also makes it easier to expose breaking sentences, abbreviations in English and multi-sounding problems。
Test three. Can audio drive a digital man
Enter is the second MP3 that has been heard, and the output is output/03-chanjing-test.mp4. The following hint applies to the individual digital person created at the mirror main station。
Please perform a single "Silver Audio Driver Digital Man" test。
Input:
– driver audio: output/02-minimax-voice.mp3
– output file: output/03-chanjing-test.mp4
Implementation requirements:
1. CHANJING_APP_ID, CHANJING_SECRET_KEY AND CHANJING_PERSON_ID READ FROM THE SECURITY CONFIGURATION。
2. AccessToken is first obtained with App ID and Secret Key, without display or permanent preservation of Token. Only refresh once returns 10400。
3. UPLOAD LOCAL MP3 TO MIRROR FILE MANAGEMENT, WAIT FOR FILE STATUS TO BE AVAILABLE AND USE RETURNED FILE ID TO CREATE THE VIDEO。
4. A digital person designated using CHANJING_PERSON_ID. This digital person comes from the main mirror station, so source set to 1。
5. select the basic model, generate 1080 x 1920, 1080p vertical video, and maintain the default subtitle processing。
6. SEARCH FOR THE SAME TASK NUMBER ON AN ONGOING BASIS AFTER TASK SUBMISSION AND DOWNLOAD MP4 IMMEDIATELY UPON COMPLETION. DO NOT REPEAT TASKS DURING THE WAITING PERIOD。
7. App ID, Secret Key, AccessToken and full digital person ID should not be shown in responses, logs or output files。
After completion, confirm that the video can be opened, with 1080 x 1920 images, and contains both video and audio streams, which are basically the same length of the video as the input audio. Report only: execution results, video hours, file size, time-consuming and file path。
THE FILE EXTENSION ONLY INDICATES THAT IT IS CALLED MP4 AND THAT THE FOLLOWING ELEMENTS ARE EXAMINED ITEM BY ITEM AT THE TIME OF ACCEPTANCE。
- Is it normal to open it
- Is it 1080 x 1920
- There's no sound
- Whether the time is close to the input audio
- Is the face, mouth and subtitle normal。
The test I've generated this time is 4.04 seconds, the file size 3,055,955 bytes, waiting around 42.6 seconds. The bean bag did not download the document until after the mission had been completed and was not considered “successful” as the end result。
I'm running out of data down there。
|
project |
result |
Time-consuming |
Output |
|---|---|---|---|
|
Take it out in the Jigsaw case |
successes |
5.4 sec |
955 B / 363 CHARACTER |
|
Mini Max Voice |
successes |
1.4 sec |
4.05 SEC / 67,956 B |
|
Mirror Digital Man |
successes |
42.6 sec |
4.04 SEC / 3,055,955 B |

All three outcome documents can be opened and the material prepared is completed through their respective acceptances。
VII. Generation of the first full digital person video with a bean bag
1. Compilation of the results of Ji-jin into an official oral report
after the three tests were passed, output/01-zhiling-copy.txt was organized into its own official broadcast. it's a raw material that you can compress and rewrite. the final version is still to be confirmed。
The first complete piece is controlled in about 10 seconds, with 45 to 65 Chinese words left. Duplicate the following hint, so that the bean bag gives a candidate and does not call the voice and digital interface。
read output/01-zhiling-copy.txt and organize it into a text suitable for about 10 seconds of digital population。
Rewriting request:
1. Use only information that can be confirmed in the original language and does not include unverified data, experiences and effects。
2. Maintain a primary message, contained in 45 to 65 Chinese words, using natural words and complete sentences。
3. Not to reproduce the long sentences in the original text, to reorganize the expression, to delete account name, curricula and unconnected opening。
4. This round only gives a copy of the submission, does not save the document, and does not call Jiji, Mini Max or Mirror API。
The list of candidates for the final report is pending confirmation。
The following sentence is reproduced after reading the candidate and confirming that the facts, tone and length are appropriate。
Confirms the use of the previous version of the candidate. Save as output/04-final-script.txt, not rewritten, not calling any external API. Only number of characters and file paths are reported after saving。
in order to separately validate the interface, i used a 58-word self-written test, also saved as output/04-final-script.txt。
- I'M GOING TO USE THE AI TOOL TO ORGANIZE THE FILE, GENERATE SOUND, DRIVE THE DIGITAL PERSON TO COMPLETE A FULL VIDEO, THE WHOLE PROCESS IS SIMPLE AND EFFICIENT, AND EVEN THE FIRST OPERATION CAN START QUICKLY。
once confirmed, the subsequent retest will only deal with failed audio or video steps and will no longer change the script。
2. Clearing the whole mission at once
Confirms that after all three individual tests have been completed, the following hints can be copied. It uses the sound and digital person in the secure configuration, without filling in real Key or ID. Actual API calls and minor costs will arise after implementation。
Please generate a treaty of 10 seconds for an official digital person, using MiniMax and Skyll, which have already been tested individually in the current project。
Fixed parameters:
– output directory: output
– oral programme: read output/04-final-script.txt
– MiniMax model: speech-2.8-turbo
– Speed: 1.0
– Audio specification: 16000 Hz, mono-channel MP3
– SOUND: A CLONED SOUND THAT IS SPECIFIED WITH MINIMAX_VOICE_ID AND HAS BEEN MEASURED SEPARATELY
– NUMERIC PERSON: A DIGITAL PERSON WHO HAS BEEN IDENTIFIED AND HAS BEEN TESTED SINGLE-SCALE USING CHANJING_PERSON_ID
– video specification: basic model, vertical 1080 x 1920, 1080p
Implementation order:
1. confirm first that output/04-final-script.txt exists and is not empty and does not modify the content of the document。
check output/05-final-voice.mp3 and output/06-final-digital-human.mp4; stop and report when any document exists and do not overwrite。
3. Call Mini Max once to generate output/05-final-voice.mp3; confirm after generation that the file is available and complete。
4. output/06-final-digital-human.mp4 is only called once after the audio acceptance has been completed。
5. Skill upload MP3 first, then submit video assignments using valid AccessToken. Saves the task number and keeps searching for the same task, and does not allow the submission to be repeated for waiting。
Security and failure management:
– The voucher is read only from the security configuration, does not modify the configuration, does not show API Key, Voice ID, App ID or digital ID。
– An immediate halt to any failure, the retention of a successful document and a summary of the failed steps and original errors。
– Allow only automatic re-testing of failed steps where a temporary network error occurred; no rerun of successful steps。
Final acceptance:
- Confirmation of the existence of all three official documents
– Reporting number of oral characters, length of audio, length of video, file size and total time
– CONFIRM VIDEO AT 1080 X 1920 AND CONTAIN BOTH H264 AND AAC AUDIO STREAMS
– Only the result table and three file paths will be exported and no other versions will be generated。
Why does it have to be written “Only retrying the failure when it fails”
In a waterline, the case is almost free, the cost of voice is low and digital video is usually slower and more visible。
If only the final downloading of the document fails, but the steps ahead are fully reruned, there will be additional audio, video and additional costs. The hint is to let the bean bag first find the failed step and only rerun this part。
Open result folder acceptance
This bean bag eventually produces six sequentially numbered documents, and both test and official products are kept in the same separate folder。

The final acceptances are presented below。
|
project |
Actual results |
|---|---|
|
Official broadcast |
58 Words (included points) |
|
Official Audio |
10.27 SECONDS, 167,604 B |
|
Official video |
10.24 SECONDS, 8,250,407 B |
|
Image |
H.264, 1080 X 1920 |
|
Sound |
AAC AUDIO STREAM |
|
Total time-consuming |
64.7 sec |
Once the video is turned on, a manual check is made to see if the person is correct, if the mouth is synchronized, if the subtitles are complete, if the beginning and end are cut off。
The article was accompanied by this reality film that can be opened directly in the computer。
VIII. Common issues and sequencing
1. Jiggling returns connection reset or timeout
Check if the public link can be opened in the browser and then check if the service is accessible. When services are available and vouchers exist, only one retest with spacing is made. Don't regenerate the sound and video because you failed。
Mini Max tip 401, unlicensed or no sound found
See if API Key comes from the current site before confirming that Voice ID belongs to the same account. Check then whether the sound is still available and whether the account has access rights. Do not put the whole Key in the wrong screenshot。
Sound generation, but not its own
Note that the bean bag is selected for the system or another clone. Go back to "My Voice" to confirm the name, specify the correct Voice ID in the security configuration, and only run back to Mini Max。
4. Lens always show processing
Digital video is a walker, and short films may take dozens of seconds. Let the bean bag continue to search for the same task and do not repeat new tasks. Check the mission status and level after the normal waiting time of the platform。
5. Mirror returns 10400 or AccessToken authentication failed
Confirm first that App ID and Secret Key are from the same application. Let Skill retake AccessToken and continue to search for or submit the current steps. Do not use old values after refreshing Token and do not allow multiple tasks to duplicate Token。
6. MP4 IS OPEN, BUT NO SOUND
THE EXISTENCE OF A DOCUMENT DOES NOT MEAN THAT IT IS COMPLETE. BEAN PACKS SHOULD CHECK THE MEDIA FLOW, AT LEAST BOTH VIDEO AND AUDIO. THIS TEST IS H.264 + AAC. ONLY THE VIDEO STREAM IS FOLLOWED BY A REHEARSING OF THE UPLOAD AUDIO OR COMPOSITE PARAMETERS。
7. Subtitles are too high, masked or cut
These kinds of problems come from templates and video layouts, and API Key stays still. Confirm resolution and direction, and reposition subtitle safe areas, fonts and numbers。
8. Repeated analysis of bean bags but delayed implementation
Shorten the task and use a defined sentence. Mini Max, when this step is delayed, copy this one directly。
Please stop planning and only MiniMax voice generation is currently being implemented。
enter the text file to read output/04-final-script.txt, and output to output/05-final-voice.mp3 using the cloned sound, speech-2.8-turbo and 1.0 speeds that have been measured singlely. do not rewrite the text, do not call the mirror or create other documents。
If the output file already exists, stop and report; if the call fails, keep the error summary and stop, do not automatically replace the model, sound or platform. command to report progress when running, and only the path, length and size of the file when completed。
After the bean bag wrote the plan, the mission was not completed. This is the end of the process when documents are dropped, open and accepted。
The most appropriate course of upgrading for White
The first video will not be generated in bulk as soon as it runs. It's easy to locate where there's a problem。
- Fixed short sentences and repeated testing of the same sound
- Fixed audio, comparing different digital persons or basic, high-quality models
- (b) Add a rewrite but confirm the final draft at each time
- SERIALIZE LINK EXTRACTION, REWRITE, TTS, DIGITAL PEOPLE INTO SINGLE STABILIZATION TASKS
- Once a single task has stabilized, it is then done in bulk, in volume and on time。
Three customs are maintained at each level of promotion。
- Intermediate documents are sequentially numbered
- Retrying the failure chain only when it fails
- Once the bean bag report is completed, the final document is opened for acceptance。
After the first ten seconds of the clip, leave the six numbered files. The next exchange of letters or figures will be followed by a return to the corresponding document at which point the error occurred, without the need to start over。