Traditionalskit 8 MINUTES, COST 50 TO 1.5 MILLION AI SHORTS, AND COST LESS THAN 200,000JellyNo, no, no, no

Don't study the hints. I pulled out the technology for the production of short operas
- Promotional messages: AI skitsIT LOOKS LIKE A "ONE-KEY GENERATION" WITH 13 UNDISCOVERED PROCESSES, 500 SCRAPS, AND MANUAL REFINEMENTS OF 40% HOURS. I TOOK ONE OF EIGHT MINUTES OF 117 SHOTS BY FRAME, AND I WROTE AN AUTOPSY REPORT。
- This paper is based on an independent analysis of the open information, all of which is speculative and welcomes the error of the producer. Yellowo shortshow reverses the reasoning: a well-spreading water line, at least three judgements need to be corrected

Let's start with what's right
Recently, a current-line map has been circulated in the AI short play entitled " AI Short-Study " (Reverse Hypothesis) " . The author did a solid job: a frame-by-frame analysis of a piece of an 8-minute, 117 lenses was carried out, a 24-fps-30-fps patch was detected, a complete 13-step process flow chart was prepared, and a hardware configuration was recommended。
This map has generated much discussion in the technological community. But when I searched myself for open information on the short opera and cross-referenced the technical parameters of each video-generated model, it was found that at least three key judgments needed to be corrected. Some are not “mistakes”, but “conclusions are made when there is insufficient evidence”。
Here's an analysis with Claude。
First shot: Huang Ko is not a play. It's a platform
This picture makes a hypothetical mistake: it uses "yellow fruit" as a specific piece of work to decipher and draw a linear 13-step water line. But it's actually one of themMulti-author content distribution platform.
The evidence is direct. Yellow fruit clearly says on its Telegram official channel and recruitment page:
“A GLOBAL RECRUITMENT OF CONTRACT CREATORS. WHETHER YOU'RE A MATURE AI DIRECTOR, WRITER, OR A VISUAL ARTIST WITH AN OPEN BRAIN HOLE, AS LONG AS YOUR WORK MEETS OUR STANDARDS, WE'LL BUY IT."
Their business model is that the platform provides distribution channels (the official network + Telegram + TikTok), provides access to “facing models”, the creator produces content and the platform acquires copyright。
What does that mean? It means a different short play on a yellow fruitProbably from different creators or small teams, using different models, different work streams, different quality standardsI don't know. You can't deduce from the frame-by-frame analysis of a play, the "cruise line of the fruit" because there is no single line of water。
In other words, the painting may be the production process of a particular work on a yellow fruit, but the title of it is “The fruit makes a stream of water”, which is tantamount to treating the characteristics of a sample as an overall feature。
Yellow fruit is more like YouTube in the AI short play world than Netflix。 Netflix has uniform production standards and internal processes, and each channel on YouTube is divided. And when you're going to analyze YouTube's "making water lines," the right way is not to take down a video, but to see what common infrastructure is at the platform level。
In the case of the fruit, this common infrastructure is the “delimiter model”. It'll spread back。
Second shot: 24fps does not equal Wan, Feedance and 24fps
One of the strongest technical arguments in that picture was the 30fps of a plate-covered frame, but the frame-by-frame test found a frame repetition of 1 for every 5 frames, which is typical of the 24:30 pulldown mark. And the conclusion is: 24fps raw output points to Wan 2.1/2.2。
the first half of this chain of reasoning is fine. the 24-30fps patch marks do prove that the original material was not generated by 30fps. but the conclusions of the latter part went too fast。
becauseNot just Wan Output 24fps.
I cross-referenced the primary output frame of the mainstream video generation model in 2026:

Key data:The output of Seedance 1.0 Pro is also 24fpsI don't know. ByteDance's technical document clearly labels its output as "up to 1080p at 24 fps, with statistics from 2 to 12 seconds"。
in other words, "24fps marks" evidenceIt's also pointing to the two model families, Wan and SeedanceCan't lock Wan alone。
The other indirect evidence given in that picture (scenario length distribution, consistent drift feature, I2V activation sense) is also not exclusive. The lens of 72% is distributed in 5 seconds, which is the same as the "2-12 seconds" window generated by Seedance 1.0. Consistency drift is a problem of varying degrees in all current models. I2V's sense of activation is a common feature of all Image-to-Video tubes, not unique to Wan。
More interestingly, Kling's original frame rate is 30fps or more (2.6, 30-48fps, 3.0 60fps). If some of the yellow fruit's works have found traces of 24-30, then..At least it can be ruled out that these cameras are directly generated by Kling, because Kling output is in itself 30fps+, does not need to do pulldown。
So the more reasonable inference is:
The creators on the yellow fruit platform probably mix models. Kling is used for regular lenses (30fps straight out), Wan or Seedance is used for those that require more customised control (24fps generated and 30fps time line unified). And Breakout Content almost always depends on the Open Source Model (Wan), because closed source API cannot allow you to fine-tune NSFW content。
This leads to a third amendment。
Third shot: The real moat is not a 13-step process, it's a "delimiter model."
That's the 13 of the graphAdult contentThe treatment was placed in step 8, entitled “Adult lens processing (adult fine-tuning/LoRA)”, which was downplayed into the process。
But in terms of the business logic of the fruit, it's not a step in 13 steps。This is at the heart of the whole business model。
One of the key words in the yellow fruit's creators' recruitment post is: “Providing access to a limited model for high-quality creators on a paid basis so that every drop of your inspiration will be perfect”
“Property to provide access to a restricted model”, which contains several layers of information:
First, the break-through model is an asset of the platform party and is not made by the creators themselves。 Yellow fruit, as the platform, invests resources to train or maintain these models, which are then made available to contract creators in the form of tool use rights. This is a “infrastructure is service”。
Secondly, these models are fine-tuned on the open source base model。 To generate content not permitted in closed-source commercial API (e. g. Seedance or Kling), you must fine-tune your own model with your own data. This can only be done by open source models. In the current open-source ecology (2026, August 2026), the most suitable base model for micro-modulation of video generation is Wan 2.2 (Apache 2.0 license, 14B reference number MoE structure, with a mature LoRA training process in place in the community)。
Thirdly, this is the real barrier to yellow fruit vis-à-vis individual creators。 Personal creators can generate regular content by Kling API or Seedance API, but to generate "facing" content, they need to collect data sets, train LoRA, and maintain model versions, much more than the threshold of "prescriptive." Huanggo packed this threshold into a platform service。
So if you're going to draw a more accurate "business map of the fruit," it's not a linear 13-step water line, but a layer structure:
Bottom: Breakout Model(possessed by platform, based on Wan fine-tuning, including role LoRA library and scene preset)
Middle level: Creator tool chain(The creators choose, Kling/Seedance/Wan mix, respective later processes)
Upper level: distribution and liquidation+ Telegram + TikTok, ad + Subscription + Copyright Acquisition)
IT'S THE MIDDLE OF THE PICTURE, AND ONLY ONE OF THE CREATORS. WHAT REALLY DETERMINES IS THAT THE FRUIT IS A FRUIT, NOT AN "AI SHORT PLAY," AND IT'S THE BOTTOM AND UPPER。
Infer from the bottom: a more rational technical hypothesis
Combining the above three amendments, I re-engineered a hypothesis about the short opera technology. It was stressed that this was still speculation, but based on more open information and more careful speculation。
Hypothesis: two-track model + lateral harmonization
GENERAL CAMERA (NON-NSFW): The creator uses Kling or Seedance for commercial API generation. These APIs have mature role coherence tools (Kling's Elements feature supports a maximum of four reference diagram locking characters) that generate high quality, speed, and do not require local computing. The dialogue lens can be used to save the late Lip-sync process by using the original mouth sync function of Seedance 1.5。
BREAK-IN CAMERA (NSFW): Use local models based on Wan 2.2 fine-tuning provided by the platform. These models are trained by the LoRA of the NSFW data sets to generate content that is not allowed by the commercial API. Local GPU algorithms are required, reasoning is slow, but running on own hardware is not subject to content clearance。
Later harmonization: The two tracks converge at a later stage. Kling Original 30fps+ camera and Wan/Seedance Native 24fps camera, uniform to 30fps time line. This explains why the frame-by-frame tests were conductedPart of the footage shows 24-30 patches, while the other footage doesn'tI don't know. If all the lenses are generated by the same model, the patch tracks should be distributed evenly; but if they were mixed models, the trace distribution would be uneven。
The hypothesis explains the known evidence better than "all use Wan":
- Why did Huanggo's recruitment post say Seedance and Kling instead of Wan
- Why is there a 24fps frame patch
- Why do different lenses sometimes have different quality
- WHY WOULD HUANGGO PROVIDE A "REMUNERATED" BREAK MODEL? BECAUSE IT'S A CLOSED SOURCE, API
Year 2026 Mainstream video generation model full parameters comparison
In order to validate the two-track hypothesis, the hard parameters of each model must be distributed. The following data is derived from model official documents and third-party baseline tests:
One of the most critical columns in this table is “Can self-defined fine-tuning”。Only Wan 2.2 allows you to train LoRA with your own data set。 ALL CLOSED-SOURCE APIS HAVE CONTENT SECURITY CLEARANCES, WHICH DO NOT ALLOW YOU TO FINE-TUNE THE NSFW MODEL. THIS LOCKS OUT THE BASE MODEL SELECTION FOR THE BREAKOUT。

Another point that deserves attention is the “native audio”. Kling 2.6+ and Seedance 1.5+ both support video and audio synchronization, which is a significant efficiency advantage in dialogue-intensive short dramas. The traditional process requires a Mr.'s silent video, a separate TTS, an oral synchronization of at least three steps. Models with primary audio are in place. This further supports the hypothesis of a closed-source API for conventional lenses, because the creator has no reason to give up this efficiency advantage。
Five perjury predictions
A good hypothesis must be perjured. Here are five testable predictions of the two-track model hypothesis, which can be verified by anyone with a piece:
Projection 1: Spectrum patches are unevenly distributed。 If the lens-by-scenario marks the presence or absence of a 24-30 frame patch, the proportion of the NSFW lens that shows a frame patch should be significantly higher than that of the non-NSFW lens. Since the NSFW lens is an output of 24fps, the non-NSFW lens does not need a frame if you use Kling (30fps+)。
PROJECTIONS 2: THE RESOLUTION CEILING FOR NON-NSFW LENSES IS HIGHER。 Kling 3.0 for original 4K, Wan 2.2 upper limit is 720p. In the 100% zoom down the detail to compare different lenses, the non-NSFW lens (perhaps with Kling) should be clearer than the NSFW lens (perhaps with Wan)。
Projection 3: Conversation lens oral synchronous mass dichotomy。 Using the dialogue lens of the Seedance/Kling primary audio tube, the mouth synchronisation should flow naturally. Using the late Wa + Lip-sync conversation lens, there may be minor mouth delays or no match for voice。
Forecast 4: There are systemic differences in movement flow。 Kling 3.0/3.5 The motion generated under 60fps is very fluid, and even down to 30fps timelines keep better movement information. The movement generated by Wan 24fps has a more obvious frustration with fast movement. A lens-by-scenes comparison of the flow of motion should divide into two groups。
Projection 5: Consistency of the same role in different types of lens varies。 The Kling’s Elements (reference lock) and Wan’s LoRA (microlock lock) fail in consistency. Kling prefers to keep face but lose body details, and LoRA prefers to keep the whole style but to shift below the extreme angle. If two different modes of drift are observed in the same play, it is a direct evidence of the mixing of models。
Any of the five above-mentioned articles would enhance the credibility of the double-track hypothesis; any explicit denial (e.g., complete parity in the frame marks of all the lenses) would weaken it。
The part of the picture that's still valuable
AFTER THE THREE JUDGEMENTS HAVE BEEN CRITICIZED, SAY WHERE THE PICTURE WAS DONE WELL. THE FOLLOWING POINTS ARE OF VALUE TO ANYONE WHO WANTS TO UNDERSTAND AI ' S SHORT PLAY INDUSTRIALIZATION。
Camera density analysis
117 SHOTS / 8 MINUTES = AVERAGE OF 4.1 SECONDS PER SHOT, MEDIUM 3.2 SECONDS. THIS DATA ITSELF IS MORE PERSUASIVE THAN ANY QUALITATIVE DISCUSSION. IT DIRECTLY QUANTIFYS THE REASONS WHY THE SHORT PLAY LOOKS SO FAST:It's not the director who likes to cut fast, it's the model that lasts five to eight seconds。
THE TRADITIONAL SHORT PLAY IS ABOUT 40-60 SHOTS, AND AI DOUBLED. THIS MULTIPLICATION WILL DIMINISH AS THE MODEL EVOLVES (LONGER SINGLE-TIME GENERATION), BUT IN 2026 IT REMAINS A HARD INDICATOR OF THE DIFFERENCE BETWEEN AI AND REALITY。
The concept of “assets”
The figure classifies process 1-3 (the script spectroscopy, role LoRA, scene template library) as an “assets layer”, which is very perceptive. It points to a fact that has been ignored by the majority of the entrants:AI THE CORE COMPETITIVENESS OF THE SHORT PLAY IS NOT THE ABILITY TO WRITE NOTES, BUT THE ABILITY TO BUILD AN ASSET BANK。
Roles LoRA, scene templates, style presets, all digital assets that can be reused across projects. A team did 10 plays and accumulated an asset bank that was completely different from a new hand's experience from scratch. The larger the asset bank, the lower the marginal cost of each new play。
This also explains why the yellow fruit has to be a platform rather than a content studio. If you have a good set of default model asset banks, the ability to do your own drama is linear (many bottlenecks), but opening it to multiple creators turns it into a multiplier relationship。
Artificially refined percentage judgement
THE FIGURE IMPLIES THAT MANUAL FINE-TUNING (HAND REPAIR, FACE REPAIR, DRESS REPAIR, REPAIR) ACCOUNTS FOR MORE THAN 401 TP3T OF THE TOTAL WORKING HOURS. THIS JUDGEMENT IS HIGHLY CONSISTENT WITH PUBLICLY AVAILABLE INDUSTRY DATA. ACCORDING TO INDUSTRY STUDIES, ROLE CONSISTENCY IS THE BIGGEST TECHNICAL PAIN IN AI SHORT-TIME PRODUCTION, WHICH WAS IDENTIFIED BY THE CREATORS OF 76% AS THE FIRST CHALLENGE。
Even in August 2026, the model would still require a large number of hand-generated, facial consistency and clothing stability. The role coherence solution has evolved to the third generation (PuLID + Adetailler + ControlNet + IP-Adapter FaceID) and the consistency can reach 95%, but the inconsistency of 5% is still captured by the audience in successive clips。
Models are progressing, but fine working hours are not reduced to zero. Declining is the simplistic density, not its very existence。
Cost truth: cheaper than you think, but less than some people blow
Industry data provide a number of reference cost benchmarks:
- FUNNY EXPRESSION BAG TYPE AI COMICS: $50-150 PER MINUTE
- FINE AI COMIC: 1000-3000 YUAN PER MINUTE
- AI HUMAN SHORTPLAYS (THE CATEGORY OF YELLOW FRUIT): COST ABOUT 301 TP3T
- PREFECT, AI, HUMAN SHORTPLAY: WITHIN 200,000
- Traditional reality shorts of equal duration: $5-1.5 million
REDUCTIONS EXCEEDING 901 TP3T. IT'S REAL DATA。
But one of the key variables is..ProductionI DON'T KNOW. IN THE SPRING 2026 SLOT, THERE WAS AN OPEN CASE: 3 PEOPLE TEAM COMPLETED 75 DAYS OF THE FULL-PROCESS PRODUCTION OF THE " AIR DELTA " BY THE FIRST 75-DAY SERIES OF THE " "AIR TRANSPORT DELTA " , WITH 16 HOURS ON THE LINE, BREAKING HUNDREDS OF TIMES。
What does this efficiency mean? It means that the big cost is not a factor, not a model callTime cost of manpowerI don't know. If an creator produces one in a week, monthly income can be right as long as it covers human costs and computing leases。
THE YELLOW FRUIT MODEL FURTHER LOWERS THE START-UP THRESHOLD FOR CREATORS: YOU DON'T NEED TO REFINE YOUR OWN CUT-OFF MODEL (PLATFORM PROVISION), YOU DON'T NEED TO DO YOUR OWN DISTRIBUTION CHANNELS (PLATFORM PROVISION), YOU JUST NEED TO HAVE SPECTROSCOPY CAPABILITY + MODEL OPERATION + LATE EDITING. IT TURNS "A MAN DOING AN AI SHORT PLAY" FROM A THEORETICAL POSSIBILITY TO A PRACTICAL IMPLEMENTABLE THING。
Legal risk: Not a “grey zone”, a solid criminal red line
It must be written here that the risk is clear. The nature of the content of the Cyclops determines that it does not wipe the edges in the grey area, but operates near a clear legal red line。
The production and dissemination of pornographic material for profit (article 363 of the Criminal Code)。 ANYONE WHO MANUFACTURES, REPRODUCES, PUBLISHES, SELLS OR DISTRIBUTES PORNOGRAPHIC MATERIAL FOR PROFIT SHALL BE SENTENCED TO UP TO THREE YEARS ' IMPRISONMENT, DETENTION OR CONTROL AND FINE. IN SERIOUS CASES, FOR THREE TO TEN YEARS. THE OBSCENITY GENERATED BY AI HAS BEEN RECOGNIZED IN JUDICIAL PRACTICE AS “OBSCENITY” AND IS NOT EXEMPTED FROM IT。
Deep-false and portrait rights。 The role of LoRA is suspected of violating the right to portrait if it is trained in the face of the person without authorization (art. 1019 CCS). If used for the production of pornographic material, there may be a number of criminal offences。
Operational risks of the Platform。 Huanggo's servers and domain names are currently outside China, but Chinese criminal law has personal jurisdiction over Chinese citizens. If the operator or creator is a Chinese citizen, even if the server is outside the country, it may be prosecuted。
Huanggo's own location is also interesting: they use multiple domain names such as “huangguai.com” “huangguai.ai”, the Telegram group as the main community, TikTok as the channel of diversion. Such a decentralized domain name and channel strategy in itself implies active circumvention of compliance risks。
For the creators, it needs to be very clear that:Writing for yellow fruit and making side videos from media platforms themselves are completely different levels of legal risk。 The former are organized content production and distribution, while the latter may be only personal。
Transport value: Discarding content to read process
Leaving aside the content properties of the yellow fruit, the proven process capability has a direct migration value for the compliance track。
There are three core transferable capacities:
Engineering capability for multi-model mixing。 Different models are used at different stages (Kling for regular lens, Seedance for mouth sync, Wan for depth customization) and then at later stages. This capability of “best modeling” applies equally to compliance scenes such as brand TVC, knowledge dramas, and the Civil Service promotional film。
Role asset management。 Training and maintenance of role LoRA, creation of a template of scenes, management of cross-scenes consistency. The asset management methodology does not depend on content compliance, but addresses the consistency of all AI image industrialization。
QC ACCEPTANCE STANDARDS。 "What kind of footage can and must be redone." This set of standards will enable individual creation to become a scalable production. Yellow fruit as a platform necessarily has a set of internal quality-checking standards for its contents, which in themselves are valuable methodological assets。
Specific application scenarios are inactive and are similar to those presented in industry reports: brand advertising, knowledge-based literacy, voice-based visualization, business training, and civil service promotion. The only thing I want to add is:WHEN THE SAME PROCESS WAS MOVED TO THE COMPLIANCE TRACK, THE BIGGEST CHANGE WAS NOT TO REMOVE STEP 8 (NSFW CONTENT), BUT TO ADD TWO NEW PROCESSES: CONTENT CLEARANCE AND COPYRIGHT COMPLIANCE。 These two processes do not exist in grey content production, but are the bulk of the costs and time in compliance production。
Conclusion: Analysis framework is more important than conclusions
THIS ARTICLE SPENDS A LOT OF TIME QUESTIONING A WELL-SPREADED WATER FLOW MAP, NOT BECAUSE IT'S WORTHLESS, BUT RATHER BECAUSE IT'S THE MOST DETAILED PUBLIC ANALYSIS I'VE EVER SEEN OF THE IA SHORT PLAY PRODUCTION PROCESS。
But a good analytical framework should be able to withstand cross-check. When I took the conclusions of the map to compare the open information of the fruit, the technical parameters of the models and the industry data, three judgements were found that needed to be corrected. Rather than overturning the original map, those amendments made reasoning more prudent。
In 2026, the AI short play was at a special stage. Model capabilities are rapidly evolving (from 720p/24fps in Wan 2.2 to Kling 3.5, with only a few months between 1080p/60fps in the middle), industrial processes are rapidly iterative (from doing everything by one person to platform division of labour), and business models are rapidly tested in error (from free to subscription fees to copyright acquisitions)。
In this rapidly changing environment, specific findings are short-lived (the data of the article may be obsolete one month after its publication), but the analytical framework can be reused: a frame-by-frame test to validate the frame rate, cross-reference to model parameters to exclude alternative interpretations, different logic to distinguish the platform layer from the creator layer, and a technical selection from the business model。
This methodology is much more useful than the specific answer, "Wan or Kling."。
THE ANALYSIS IS BASED ON PUBLIC INFORMATION FROM AUGUST 2026. AI VIDEO GENERATION FIELDS ARE HIGHLY ITERATIVE AND THE MODELLING CAPABILITIES AND COST DATA INVOLVED IN THE TEXT MAY HAVE CHANGED AT THE TIME OF PUBLICATION。