ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

IT'S NOT PARTICULARLY DIFFICULT TO USE AI TO CREATE A 10-SECOND PICTURE。

The real difficulties are:
(a) How to keep the characters, drawings, dramas and emotions consistent when a story requires dozens of shots, multiple scenes and multiple productions
When a draw card fails, how to modify it at the smallest cost is to solve only the part of the problem, rather than to reverse the whole video。

(THE ARTICLE ENDS WITH AN APPENDIX. IF THIS IS TOO LONG TO READ, YOU CAN COPY THE APPENDIX IMMEDIATELY TO THE GPT.)

THIS TIME, WE'RE TRYING TO TURN A MINISTER'S SCRIPT INTO AN AI NARRATIVE OF ABOUT FOUR AND A HALF MINUTESShort Film.

Because the original script is long and not original, we did not start by looking for the whole story, but by choosing one that is relatively complete and independent. Eventually, this part of the story was broken into 20 separate video clips, which were then made up of a complete short clip by editing。

In the process, we have gradually developed a reusable production method。

If only to quickly understand the whole approach, it would be sufficient to remember the next five phases
If you are prepared to actually start production, use the full working stream of the later text, confirm nodes and start-up hints。

The whole process can be understood in five stages

ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

Figure: First determine how much you do, then turn the script into a productionable director's camera script, then prepare the assets, batch generation, and then finish the film by editing and local refurbishment。

These five phases are:

  1. Set production range
  2. The script became the director's camera script
  3. Production and binding of visual assets
  4. Test with Agent Batch Generation
  5. Cuts, fixes and pieces

They are “cognitive frameworks” that help readers to understand quickly。

The real implementation will also include judgement, confirmation, freezing, retrofitting and retreat within each phase. As a result, a more complete implementation flow will be provided at the end of the article。

What tools are used for this project

This approach does not depend on a fixed software, but requires a distinction between four types of work:

  • Screenplay analysis, director planning and quality review
  • Visual asset production for characters and scenes
  • Video generation and mass task distribution
  • Cuts, sounds, subtitles and final output

The main tools for this project are set out below。

1. Planning end: ChatGPT Work+GPT-5.6 Sol

ChatGPTWork, a working model, is responsible for managing long work processes, reading scripts, dismantling missions, recording confirmation results and awaiting manual review at each stage。

Officially, Work has been positioned as a working model that allows sustainable review of results, such as documents, tables, presentations, etc., to be completed in combination with objectives, documents and context。

GPT-5.6 Sol is used for script analysis, director lens planning, continuous inspection, preparation of tips and quality judgement to produce results. It is officially positioned as a front-line model for complex professional work。

ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

Figure: Planning end-use work model and GPT-5.6 Sol, responsible for dismantling, planning, inspection and stage confirmation。

2. Asset end: ChatGPT image generation

Personal setting, clothing reference, scene reference and key prop to generate and adjust directly in ChatGPT。

My practical experience is that a working model is appropriate to manage the project as a whole, but if a person chart or scenario map is repeatedly unsatisfied, an identified asset description and reference chart can be taken to an ordinary chat-in-the-mode picture。

This is not to suggest that there must be a fixed quality difference between the two models, but rather that single images tend to be more easily concentrated in ordinary chats and are not disturbed by the complex context of the project as a whole。

3. Generation end: Dream + Feedance

The video generation component of the project uses the Agent model in the dream canvas and uses the Seedance 2.0 Mini for video generation。

Here's the difference between "Agent" and "Video Model":

  • Video models actually generate a video segment。
  • Agent is more like a dialogueable task assistant who reads requests, matches assets, distributes assignments and organizes multiple generation。

In other words, we can give Agent multiple tasks at a time, but it does not directly generate an ultra-long video, but rather assigns the content to several separate video-generated tasks。

ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

Figure: The generation end uses the dream-painter Agent model, sets 16:9 and selects the appropriate video-generated model。

It should be noted that the present document records the status of the tool at the time of production of the project. The model version, the duration of the individual period, the volume, the credit price and the interface function may continue to change。

Therefore, before a new project can really start, it is necessary to reconfirm the boundaries of the capabilities of the current instrument and not to regard the specific figures in this paper as a permanent rule。

Phase I: determination of production scope

Many projects were not lost in the quality of generation, but in the beginning they did not decide “how much to do”。

Once the extent of production is not determined, the next lens script, the number of assets, the cost of generation and the production cycle will be in a changing state。

ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

Figure: First, the project objectives, original core, tool capacity and cost of generation are identified, then the decision is made whether to complete the project, to complete it in stages or to select a relatively complete segment。

A ministerial script can usually be produced in three ways。

Method I: complete at once

Suitable for:

  • The script itself is not long
  • The story is perfect
  • Total time and budget control
  • I hope I finally get a full piece

The term “one-time completion” here refers to the completion of the entire script within the same production cycle, not to the generation of an ultra-long video by video models at a time。

In the case of self-generated, mature scripts, priority can normally be given to the production of the entire project, as long as the tools, time and budget allow。

Mode two: complete phased completion

Suitable for:

  • It's worth the whole story
  • But it's too expensive, long or too difficult to manage
  • It can be phased in by chapter, scene, or feature

This approach is not to abandon the whole work, but to break it down into multiple production cycles。

Mode III: Select a relatively complete segment

Suitable for:

  • The original script is very long
  • The current project is mainly for testing methods
  • Limited budget
  • The original wasn't his own work
  • All you have to do is present one of the complete drama units

This is the way in which a relatively complete section is chosen from a long script and eventually produced about four and a half minutes of narrative footage。

The point is: this decision must be made before the camera is removed。

If the budget is not enough and the model is not long enough, the previously written lenses, assets and hints are affected。

Also confirm the tool boundary

In determining the scope of the content, it is also necessary to identify together:

  • How long can the current model generate as much as a single time
  • Which banners and resolution are supported
  • Support the frame of the front, the picture reference or the full reference
  • How many assets can be quoted for each assignment
  • Agent, how many assignments can you send at once
  • Agent will automatically change the hint
  • How much is consumed per segment
  • How much is required to set aside for test and re-engineering

The project uses a stand-alone video clip of about 5-15 seconds, but this is only the project programme at the time, and not all models must use this length。

Cost estimates, scope decisions

At the time of the project, Feedance 2.0 Mini produced a 15-second video segment at no discount to consume 130 points。

This is only a historical record of the cost of the project at that time, and does not mean that the same price is still used, either now or in other models。

A rough budget can be prepared in the following way:

  • Basic generation cost = projected number of video segments x unit fraction
    Project budget = base generation cost + test cost + back-up provision

If 20 segments are expected to be generated, there will be no preparation of only 20 segments, and space will be set aside for test, failure and partial re-shooting of high-risk segments。

The first stage, which ultimately needs to be confirmed, is not a vague “I'm going to make a video”, but a clear statement of scope:

  • Full, phased or out
  • Target is a piece of time
  • Target platform and painting
  • Visual style
  • Sound and lines
  • Models and tools used
  • Number of approximate video segments
  • Cost and retrofit

Phase two: turning the script into the director's camera script

The original script cannot be directly equivalent to video-generated hints。

The script mainly describes people, events, lines and emotions; video models need to know:

  • What's in the current picture
  • What's the status
  • What did the character do
  • What about the camera
  • When will the camera be switched
  • What's the end of this section

So we need to translate the script into the director's camera script。

ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

Figure: Split the story into a separate video section and continue to break it into a specific lens; allocate time to each shot and check the continuity of actions, lines, emotions and clips。

Distinguishing the video segment from the lens

The two concepts cannot be confused。

Video session

A video segment corresponding to an independent video generation mission。

For example, paragraphs 1, 2 and 3 are generated separately from video models。

Lenses

The camera is an internal image unit of a video segment。

A 15-second video segment, for example, can contain:

  • Camera 1: People walk into the room
  • Camera two: Cut another person's reaction
  • Camera three: Two people talking
  • Camera 4: A key props

Therefore, one video generation does not necessarily have only one shot。

Why does every video section continue to tear down the camera

There are three main reasons。

One, control the rhythm

If the model is only told “to generate a 15-second two-person conversation”, it is likely that the model will determine its own scenery, rhythm and actions, and that the results will not necessarily meet narrative needs。

The focus and time allocation of the images will be clearer when the first, second and third shots are removed。

Second, let the different video segments naturally connect through the clip

The traditional method of continuous generation often uses the end of the previous paragraph as the chapeau of the next paragraph, as a constraint on the location of the person and the continuity of the image。

This approach is appropriate for:

  • One shot at the end
  • Continuous exercise
  • Accurate camera extension
  • We have to keep the same track

However, long narratives usually include a large number of camera switching, and each paragraph does not need to be strictly sequential。

This project follows another approach:

  • Each video segment is produced independently and does not use the end of the previous paragraph as the frame of the next paragraph; the interface between the video segments is done by normal lens switching and state continuity。

We have abandoned the strict “framework continuum”, but we retain:

  • Personal identity
  • Clothes continuum
  • Scene and time continuum
  • Props are continuous
  • Action result continuous
  • Emotional
  • The sight and the axis are continuous
  • Light follows the whole style

This reduces both the rigidity of the front-end frame relay and the working methods of the traditional video clips。

Third, leave the minimum correction unit for later repairs

If there's four shots in a 15-second video clip, only the third of which has a problem, we can:

  • Remove Third Camera
  • Shorten third shot
  • Replace with other material
  • Just one new shot

The whole 15-second video need not be regenerated。

This is why the splitting of cameras not only serves generation, but also directly determines the cost of returning to work。

By what rules

The following factors can be considered in combination:

  • Did the scene change
  • Change in time
  • Whether or not a whole unit of characters is formed
  • Did the mood finish a turn
  • Is it possible to finish the line within the limit
  • Time allowed for single generation of the current model
  • Whether this paragraph has a clear start and end state
  • Whether it is easy to edit the material after generation with the material

Each video section should answer two questions:

  1. What status does this start
  2. After the end of this paragraph, what is the status of the next paragraph

The director's camera script can use the following format:

PARAGRAPH X: ESTIMATED DURATION: 15 SECONDS

Targets of the plot:
What is the narrative task to be accomplished in this paragraph

Start state:
What's the status of characters, scenes, props, emotions and movements

ONE SHOT FOR X SECONDS:
The scenery, the flight space, the movement of people, the expression, the lines, the environment。

CAMERA TWO, X-SECOND:
The scenery, the flight space, the movement of people, the expression, the lines, the environment。

CAMERA THREE, X-SECOND:
The scenery, the flight space, the movement of people, the expression, the lines, the environment。

End state:
Where does this end up? What is the next one to inherit

Asset binding:
What figures, costumes, scenes and props are used

Prohibited:
What people, movements, objects, subtitles or images can't change

Each camera needs to be allocated a time frame for the recommendation, and the time frame cannot be greater than the total length of the video segment。

After this, we have to examine each paragraph:

  • Are the actions and lines really fit
  • Is there any contradiction between the person's direction and the person's direction
  • Whether the props suddenly appear or disappear
  • Is the action contrary to space relations and basic physical logic
  • Is the mood turning too fast
  • Whether the lens switch has a cut basis
  • Can the end state automatically connect to the next one

After the director's script is confirmed, the version should be frozen before the asset is produced。

Phase III: development and binding of visual assets

Don't start to produce a large number of figures and scenes as soon as you get the script。

The more secure order is:

  • The director's camera script is completed and the necessary assets are pushed back from the camera script。

Because it's only when the camera is sure that we really know what's going on。

ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

Figure: The minimum necessary assets are extrapolated from the director's camera script that has been frozen, and the numbering, binding and freezing of the version is consolidated。

Assets comprise mainly four categories

Personal assets

include:

  • Face and hair
  • Age and quality
  • Body %
  • Heads, sides and back
  • Main expression
  • Clear identifiers

Clothing assets

If the same person changes clothes at different times or in a different scene, a separate number is required, which cannot be written only “on black”。

Site assets

include:

  • Overall space structure
  • Entry and exit
  • Main furniture position
  • Light source direction
  • Time and weather
  • Different perspectives for lenses

Key props assets

Only the real impact of the drama or the need for special props would be worth making alone。

Common items such as general cups, tables and chairs do not have to generate all of the assets, otherwise the larger the number of assets, the higher the management costs。

Assets can come from three directions

  • Reference Picture Existing
  • Stable key frame intercepted from a successful video
  • Minimum necessary assets regenerated

Upon confirmation, a clear name and version of each asset is required, such as:

P01_THOUSIAN_BLACK
P02_CHOOSE _CHOOSE _BLANK _V2
S01_CONTROL CENTRE_NIGHTVIEW_V2
S02_ APARTMENT BEDROOM_NIGHTVIEW_V4
D01_DIVORCE AGREEMENT_V1

Asset numbers are then tied to specific video segments and cameras。

Instead of just "reference the lead man's picture", it should read:

  • PARAGRAPH 6 USES P01_V3, P02_V2 AND S02_V4。

This would provide an immediate understanding of which segments would be affected, even if a particular picture were replaced at a later stage。

ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

Figure: Person and scene assets that were eventually locked in by the Argument project. Assets are identified with a single number and tied to the corresponding video segment。

Once assets are identified, the version should be locked。

The reference maps should not be changed frequently during bulk generation unless fundamental problems with the identity of the person, clothing logic or scene structure have subsequently been identified。

Phase four: test first, then generate the Agent mass

When lens scripts and assets are ready, do not generate all videos immediately。

A high-risk footage is selected for testing。

ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

Figure: Test lens scripts, assets, tips and model capabilities with high-risk footage; after the test, Agent distributes multiple independent generation tasks。

What's a high-risk footage

For example:

  • Two characters are in the same frame and interact
  • There are clear lines and words
  • It's complicated
  • You need to lie down, hold, embrace, etc
  • Existence of mirrors, screens or special props
  • There's a lot of emotional change
  • It's complicated
  • Also contain multiple lens switches

If the model could be substantially completed even for the most difficult segments, the risk of bulk generation would be much smaller。

If the test fails, the problem can be identified in advance:

  • The lens script is overcrowded
  • It's too difficult
  • Asset instability
  • The scene is not structured
  • There is a ambiguity in the hint
  • Current model capacity is inadequate

ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

The diagram uses a 15-second high-risk segment to verify the performance of the person, scene, action and lens, and then decides whether to generate the mass。

Agent, it's not a video model, it's a mission assistant

When using the dreampaint, it is understood as a production assistant who can talk。

It can help us:

  • Read multiple hints
  • For each match of character and scene assets
  • Add uniform painting and style requirements
  • Create multiple independent video generation tasks
  • Assistance in managing the generation of results

But because it can understand and process the text, it may also be “smartly” reworded。

For example:

  • Compress content
  • Change camera order
  • Merge Actions
  • Rewrite lines
  • Delete information it considers duplicate
  • Reassemble the confirmed footage

The instructions given to Agent must therefore be clear:

  • Agent is responsible only for distribution and additions and may not modify the principal phrase already confirmed。

Directly Available Agent Batch Generation Command

Please generate, in parallel, 10 separate 16:9 live video paragraphs corresponding to paragraphs 1 to 10。

Do not use the end of the previous paragraph as the frame for the next paragraph. Each paragraph must be created independently of the person, scene and props, in accordance with his or her initial state。

In assigning tasks to each paragraph, please place the original text of the principal phrase of the paragraph that I provided in the corresponding video-generation task。

The camera, action, line words and order of execution shall not be subtracted, summarized, polished, rewritten, merged or rearranged。

All you need to do is:

1. To place at the beginning of the introductory words the assets of persons, clothing, scenes and props to be used in the paragraph
2. Complementing the uniform painting, style, quality, sound and global rules
3. Checks whether paragraph numbers, asset numbers and hints are correct
4. Establish a separate video generation mission for each paragraph。

In the event of a conflict between asset requirements, global rules and the principal phrase, do not modify the principal phrase. Please indicate the conflict and wait for my confirmation。

Special attention needs to be paid to:

  • The submission of 10 paragraphs at a time does not mean that the model directly generates a long video containing 10 paragraphs。

What actually happened was:

  • The user submits a set of tasks →Agent creates an independent task for each segment The video model generates each video segment。

ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

Figure: A task submitted to Agent at a time, assigned by Agent and generated 10 separate videos, instead of a very long one。

Prior to formal submission, it is recommended that another pre-screening be performed:

  • Correct number of each paragraph
  • Completeness of each primary hint
  • Is there a match between character and scene asset
  • Is the start state clear
  • Does the end state leave space for subsequent clips
  • Can lines and actions be completed within a limited time frame
  • Was it wrong to inherit the ending of the previous paragraph
  • Agent is not allowed to rewrite the main message

The project completed a high-risk footage test, followed by the generation of the remaining video clips in batches and the retention of assets, tests, batch results and complementary material in the same canvas。

ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

Figure: Full canvas of the Arsenic project: left-hand man-and-scene assets, with high-risk footage tested and the rest generated in batches, right-hand retention of the results, resulting in 20 paragraphs of material for editing。

The value of this plaque management is not just “many segments generated at a time”, but it allows the Asset, Mission, Results and Refurbishment Records Office to be in the same project space to track what reference is used, why failure, and how it is subsequently modified。

Phase V: editing, re-engineering and production

AI VIDEO GENERATION IS COMPLETE AND DOES NOT MEAN THAT THE WORK HAS BEEN COMPLETED。

The result is just material. The real form still needs to be edited。

ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

Figure : Take a paragraph-by-paragraph look, position the problem to a specific lens, then remove, replace, recapitulate or regenerate according to the problem level。

Let's start with the actual material

Do not mechanically edit according to the expected length of time in the camera script。

The anticipated duration is only pre-generated planning and should ultimately be judged on the basis of actual material:

  • Which move is really done
  • What second of character is most natural
  • Which shot is worth keeping
  • What pauses can be shortened
  • Which material, though successful, did not help the story

Use the material available to finish the rough cut and then observe the overall rhythm。

Refurbish by level of problem

If the problem is only one shot, you can:

  • Remove Directly
  • Shorten lens
  • Replace with other material
  • Intercept the available part from the successful video
  • I'm going to take an alternative shot

If the problem affects the entire video section, for example:

  • Starting state is completely wrong
  • The character is clearly moving
  • The scene is not consistent
  • Core action not completed
  • Wrong line order
  • The whole space relationship is a mess

This is when the whole video section needs to be regenerated。

This is also why we insisted on “continue to tear down the camera inside the video section”: We would like to reduce the number of returns from the whole video to a single shot。

Extra footage can be deleted directly

Not every shot generated must be used。

If a shot does not help the story, it can be deleted even if the picture itself is not clearly wrong。

The clip is not a passive collage of the resulting results, but a final narrative creation。

Finally, we'll deal with subtitles and sound

Once the image structure is basically locked, it is then:

  • The lines and the images match
  • Subtitle production
  • Environmental sound supplement
  • Sound Adjustments
  • Musical and emotional relations checks
  • Volume Balance
  • Final Export

Finally, an overall quality check was conducted:

  • Is the character consistent
  • Continuity of clothing, scenery and props
  • Is it consistent
  • Whether or not the axes and sights are misleading
  • Complete line
  • Is the subtitle correct
  • Whether there are any excess figures, words, watermarks or abnormal limbs
  • Compliance of painting, resolution and sound specifications with the release platform

What is the relationship between five phases and full workflow

Five phases are a methodological framework to help people quickly understand:

  • Find out how much to do, how much to do, how much to write into the resulting camera script, how to prepare for visual reference, how much to shoot and how much to produce, how much to cut into a complete film

However, there is a large number of identified points that need to be addressed when implementing them:

  • Adoption of the adaptation programme
  • Whether the video segment contains actions and lines
  • Is the director's camera script frozen
  • Asset version confirmed
  • Is the high-risk test passed
  • Agent whether or not to keep the full hint
  • The question is whether the camera, the asset or the rhythm
  • Whether the modification will affect what has previously been confirmed

Therefore:

  • Five phases are cognitive models that can be easily understood by the reader; a complete workflow is an implementation manual for those who are really prepared to make them。

Just understand the methodology, remember five stages。

Prepare to formally start the project and then use the full flow chart below。

ONE TEAM OF AI VIDEO COMMERCIALS RUN A FIELD COURSE, RUNNING FROM 0 TO 1 THROUGH THE ENTIRE PROCESS OF PRODUCING AN AI SHORT FILM

Note: It is recommended that downloads be scaled up, as shown belowTHE CODE FOR A MASTER FLOWCHART, IF IT IS NOT VISIBLE, CAN BE COPIED DIRECTLY TO GPT

What does this really solve

In retrospect, the most important part of the process is not how many steps have been added, but the resolution of several practical issues。

1. Make sure how much it takes to avoid losing control of the project in half

Full, phased and extracted decisions are not post-generated decisions, but production decisions before the project begins。

2. Translating literary scripts into director lens scripts that can be implemented by models

Models need to know not only what happened, but also where the picture begins, how it changes, where it ends。

3. Do not press all video segments for end frame interface

Long narratives can be used to switch from normal lenses to connect. Each segment is generated independently, while maintaining continuity of character, state, emotions and space, often with greater flexibility。

4. Reduction of the back-office to the camera

The lens is removed and replaced or supplemented at a later stage without having to repeat the whole video。

5. Treat Agent as an assistant, not as a director

Agent may distribute tasks and additional assets, but the principal hint, the sequence of the lens, the actions and the lines must be controlled by the confirmed director's camera script。

6. Need to identify and freeze at each critical stage

The lack of confirmation continues to push downwards, with assets, hints and videos being based on instability。

INSTEAD OF LETTING AI FINISH EVERYTHING AT ONCE, IT STOPS AT THE RIGHT NODE AND WAITS FOR A DECISION。

Attach: When a new item is opened, you can send a hint directly to ChatGPT Work

The following hint can be used directly at the start of a new script。

It is recommended that a new ChatGPT Work project be created, that GPT-5.6 Sol be selected, and that this hint be sent with the Full Implementation Flow。

  • You will serve on this project:

    1. THE DIRECTOR OF THE AI NARRATIVE FILM
    2. Directed the camera script planner
    3. Visual asset planning and version management assistant
    4. Video-generated alert writer
    5. Generation of quality inspectors
    6. Consultant for editing and partial refurbishment。

    YOUR MISSION IS TO ASSIST ME IN THE GRADUAL TRANSFORMATION OF AN ORIGINAL SCRIPT INTO AN IMPLEMENTABLE, AUDITABLE, BATCH-GENERATED, PARTIALLY RETROFITTED, SHORT-POSTED VERSION OF AI。

    In the next message, I will provide scripts, references and tools that are currently being prepared。

    Please strictly observe the following modus operandi。

    I. General principles

    1. Do not rewrite, without confirmation, the preponderance of the original, the relationship between the person, the core line, the temperament and the confession。
    2. If deletions, re-allocations, consolidations or visual adaptations are required, the reasons and effects must be stated, pending confirmation by me。
    3. The project as a whole uses only the following five phases and does not create a system of steps that are duplicative:

    Phase I: determination of production scope
    Phase 2: the script becomes the director's camera script
    Phase III: development and binding of visual assets
    Phase IV: Tests and Agent bulk generation
    Phase V: editing, re-engineering and production。

    4. One phase at a time. I have not confirmed this stage and cannot automatically proceed to the next phase。
    5. After confirmation at each stage, the current version is recorded and considered frozen。
    6. If subsequent changes affect what has been frozen, the affected phase, video segment, lens or asset must be listed before I can confirm that the lock has been unlocked。
    7. You are responsible for making professional recommendations, identifying risks and examining errors, but the final decision is confirmed by me。

    [II. Phase 1: Determination of the scope of production]

    Upon receipt of the script, the following tasks will be completed:

    1. To extract the circumstances, persons, lines, air and words that must be retained
    2. Estimating the total duration of the original script suitable for generation
    The judgement is more appropriate:
    – Complete at once
    – Complete phasing
    – Select a relatively complete clip
    4. If it is not possible to generate a full set of proposed content boundaries, explain why this begins and ends here
    5. To identify target platforms, banners, visual styles, sound programmes and rescale scales
    6. Identify the capacity boundaries of video tools, including:
    - The maximum generation time of a single segment
    – Available photo or video references
    – Agent's mass task capacity
    – Is it possible for Agent to rewrite the hint
    – A single-part cost and a back-up budget
    7. Recommend the expected length of the film, the number of video segments and the scope of production。

    Stop after completion and wait for my confirmation。

    Phase Three: The script becomes the director's camera script

    Following the confirmation of the first phase:

    1. Disassembly the scripts within the confirmed range into separate video segments
    2. A video segment corresponds to an independent video generation mission
    3. Each video segment continues to be split into one, two and three
    4. The recommended duration of each camera allocation is continuous and the addition is checked to exceed the total length of the video segment
    Each video clip must state:
    – The plot target
    – Commencement
    – The scenery, the flight, the movement, the expression, the lines and the sound of the camera
    – End state
    – Relates to the editing of the back and forth
    – Asset requirements
    – Prohibition
    By default, using the method of “independent video segment + continuous camera clipping” without the mandatory use of the end of the previous paragraph as the frame for the next paragraph
    7. There is still a need to examine the continuity of the identity of the person, the dress, the scene, the props, the results of the action, the mood, the vision, the axis and the light
    8. A paragraph-by-paragraph examination of whether action, lines and emotions can be completed within a limited period of time
    9. If they cannot be installed, preference will be given to compression, resize the camera or split the video segment, so that the model does not force too much。

    Stop after completion, wait for me to confirm and freeze the director's script。

    [IV. Phase 3: Production and binding of visual assets]

    After the director's script confirms:

    1. The least necessary assets from the frozen camera script
    2. Separately:
    - The person
    - Clothing
    – scenes
    – Key props
    3. Do not generate redundant assets that are not actually present in the lens
    Categorizing existing images, key frames for successful videos and assets still to be generated
    5. Establish a clear number, name and version for each asset
    6. Tie assets to the corresponding video clips and cameras
    7. If a single image of a person or scene is produced repeatedly and repeatedly in the working model, I may be advised to bring the confirmed description and reference diagrams to the ordinary chat and to create them separately, but without any unauthorized change in the asset setting
    8. Releases frozen after asset recognition may not be changed at will during bulk generation。

    Stop after completion and wait for my confirmation。

    [Phase V, IV: Test with Agent Mass Generation]

    After asset recognition:

    1. Select an action, frame, line, space relationship or more emotional video segment to be tested
    2. Check for character consistency, scenes, actions, lines, camera execution and overall style
    3. When the test is not passed, the question arises first from lens scripts, assets, hints or model capabilities
    4. Batch generation tasks will be organized only after the test
    5. Tie the right assets for each video segment
    6. In preparing the Agent General Directive, it must be clear that:
    – Each segment is a separate video mission
    – Do not use the end of the previous paragraph as the starting frame for the next paragraph
    – Each segment is independent in the form of characters, scenes and props based on its starting state
    – Each paragraph of the main introduction is retained as it is
    – Agent may not delete, summarize, colour, rewrite, merge or rearrange cameras, actions, lines and sequences
    – Agent is responsible only for supplementing assets, global style, painting, painting and rules, and for establishing independent missions
    7. Check paragraph numbers, asset numbers, start-up status, end-state, length of time, lines and prohibitions before the batch is submitted
    8. If a conflict is detected, it is not permitted to modify the main subject message that has been frozen on its own and must be reported first。

    Stop after completion and wait for me to confirm whether to submit for generation。

    "Six, Five: Cuts, Retorts and Scripts."

    Upon receipt of the material generated:

    1. Planning for rough cuttings based on material actually available and non-mechanical reproduction of the projected duration
    2. Take a paragraph-by-paragraph look and position the problem to a specific lens
    3. If only one shot is affected, priority is recommended:
    – Delete
    - Shortening
    – Replacement
    - Intercept other available materials
    – Separately
    4. It is recommended that the entire video segment be regenerated only when the start state, identity of the person, scene, core action, line order or entire section is logically wrong
    5. To propose solutions for asset drift, error in action, line and rhythm, respectively
    6. Processing of subtitles, sound, music, end-checking and export once the image structure is locked
    7. Final inspection of persons, clothing, scenes, props, emotions, axes, lines, subtitles, sound, unusual images and release specifications。

    [VII. Output format for each stage]

    Only the current phase is exported each time and organized according to the following structure:

    1. Current phase objectives
    2. Analysis and judgement
    3. Results at this stage
    Risks and issues to be identified
    5. Decisions required of me。

    At the end of each phase it must be stated that:

    “The above is the result of the current phase. Please review; if there are no questions, please answer `Acknowledge' and I will proceed to the next stage.”

    When you first receive this tip and a workstream, do not pre-analyze the script that has not yet been sent, but respond:

    “Workstreams are included. Please send scripts, reference pictures, target-forming styles, as well as video tools and limitations already identified. I'll start with the first phase after receiving it.”

flowchart TD
["PROVISION OF ORIGINAL SCRIPTS AND REFERENCES"] &gt; B ["DEFINITION OF PROJECT OBJECTIVES"]<br/>Platforms, banners, styles, target duration, sound, calibration of scales]

B &gt; C[“EXTRACT ORIGINAL CORE]<br/>"Circumstances, persons, lines, aromatics and words to be retained"]

C &gt; C1 ["ESTIMATING THE DURATION OF THE COMPLETE SCRIPT AND THE COST OF PRODUCING IT]<br/>Check the length of the model, bulk capacity, mode of reference and weights”]
C1 > C2{"IS IT ALL GENERATED?” ♪ I'M SORRY ♪

C2 - "YES" - &gt; D["VISUAL APPLIANCE AND RHYTHM PLANNING"<br/>Narrative nodes, attention curves, core and optional segments"]
C2 - NO - &gt; C3 [ "SELECT A PHASED PRODUCTION"]<br/>"or intercept relatively complete footage"
C3 > D

D-E {“REFORMULATION CONFIRMATION?” ♪ I'M SORRY ♪
E - NO - > D
E - YES - F ["DISASSEMBLE INTO 5-15 SECONDS INDEPENDENT VIDEO SEGMENT"]<br/>A corresponding video generation"]

F &gt; G["CURRENT CAMERA SCRIPT"<br/>Each segment continues to be split into one, two and three."
G &gt; H["SETTLING RECOMMENDED DURATION FOR EACH SHOT<br/>Add to check the total length of the paragraph']

H-I'M "CAN ACTION, LINES AND EMOTIONS FIT?" ♪ I'M SORRY ♪
I - NO - > J ["COMPRESSION ACTION, ADJUSTMENT OF LENGTH OR SPLIT PARAGRAPHS"]
G

I - YES - K ["SPEECH-BY-SPEECH"]<br/>Direction, motion, space, physics, props, lenses, lines"
K-L{L”DIRECT CAMERA SCRIPT CONFIRMED AND FROZEN? ♪ I'M SORRY ♪

L - NO - > M ["REFORM QUESTION PARAGRAPH OR QUESTION LENS ONLY"]
M-H

L - YES - N ["REVERSE AND COMPRESS THE LIST OF ASSETS"]<br/>Roles, costumes, scenes, key props]
N–O{ASSET SOURCE?}

O - "PHOTO AVAILABLE" - > P ["ARRANGE EXISTING ASSETS"]
O - "SUCCESSFUL VIDEO" - Q ["STANDING KEY FRAMES"]
O - "ALWAYS MISSING" - &gt; R["GENERATING MINIMUM NECESSARY ASSETS<br/>Turn-chat mode generation when work effects are not ideal]

P–>S[“UNIFORM TITLE, VERSION AND RULES OF SUCCESSION OF ASSETS”)
& S
R-S

S->T{“ASSET CONFIRMATION AND FREEZING?” ♪ I'M SORRY ♪
T - NO - > N
T - YES - U ["STOW ASSETS TO CORRESPONDING PARAGRAPHS AND LENSES"]

U-&gt; U1 ["SELECTION FOR CONTINUITY OR HIGH OPERATIONAL RISK<br/>A video clip to test."
U1 > U2{"IS THE TEST PASSED? " ♪ I'M SORRY ♪

U2 - NO - &gt; U3 ["PERFORM AND MODIFY AFFECTED<br/>Scripts, assets or hints for lenses"]
U3 > U1

U2 - Yes &gt; V ["Integration is DreamAgent General Instructions"]<br/>Global rule + asset binding + separate hints<br/>Disable Agent from rewriting camera maintips']

V &gt; W ["PREHEART CHECK BEFORE SUBMISSION"]<br/>Independent start-up status, continuity of editing, assets, lines, sound, length of time”)
W > X{"CAN SUBMIT? '}

X - NO - > Y ["MODIFICATION ONLY OF AFFECTED HINTS OR ASSET MAPS"]
Y-V

X-Yes Z-Z-Z<br/>Assigned to multiple separate video generation tasks

Z-> ZA ["PHOTO-BY-SPEE"]
ZA – ZB{“WHAT LEVEL IS THE PROBLEM?” ♪ I'M SORRY ♪

ZB - &gt; ZC ["DELETE, SHORTEN, SUPPLEMENT OR REPLACE PROBLEM LENSES"]<br/>Regenerated corresponding paragraphs as necessary]
ZB - "ASSET DRIFT" &gt; ZD ["ENHANCED ASSET DESCRIPTION OR PHOTO REFERENCE"]<br/>Regenerated affected paragraphs]
ZB - &gt; ZE ["REGULAR LENS LENGTH, DENSITY OR PARAGRAPH DIVISION"]<br/>Regenerated affected paragraphs]

ZC ZA
ZD-ZA
ZE ZA

ZB - "ALL THROUGH" - ZF [CLIP MERGE, SOUND CORRECTION]<br/>Subtitle Text and Final Films

How to use

  1. Creates a new ChatGPT Work project。
  2. Select GPT-5.6 Sol。
  3. Sends the above start-up hint together with the full implementation flow。
  4. Send scripts, reference pictures, target styles and tool limits。
  5. AI COMPLETES ONLY ONE STAGE AT A TIME。
  6. Answer “confirm” without a question; point out those parts that need to be amended。
  7. Once all five phases have been completed, the assets already identified and the creation of the hints are handed over to the video-end for execution。

Thus, the main task of a person is to move from “to deal with all the details in person” to:

  • Decision direction, audit results, confirmation version, and correction at key nodes。

Conclusion

AI NARRATIVE SHORTS ARE NOT COPIED INTO THE GENERATOR AND WAIT FOR THE MODEL TO BE AUTOMATICALLY COMPLETED。

It's more like a little video production:

  • What does this decision say
  • How does this decision of the director's foot show
  • Visual asset binders and the world
  • Agent is responsible for distribution and organization of tasks
  • Video models generate material
  • Cutting and remodeling ultimately determines whether or not the story works

At the core of this approach is not to add complexity to the process, but rather to place it at the right stage。

Just to understand the approach, it is enough to remember five stages。

When really prepared to do so, the full workflow and stage validation mechanism is used: first to determine the scope, then to disassembly, first to test, then to generate in bulk, and in case of problems, to fix only the wrong camera or video segment。

WHEN THERE ARE CLEAR INPUT, OUTPUT AND VALIDATION STANDARDS FOR EACH PHASE, AI IS NO LONGER JUST A RANDOM IMAGE-PRODUCING TOOL, BUT WILL GRADUALLY BECOME A MANAGEABLE, AUDITABLE AND ITERATIVE NARRATIVE PRODUCTION SYSTEM。

statement:The content of the source of public various media platforms, if the inclusion of the content violates your rights and interests, please contact the mailbox, this site will be the first time to deal with.
TutorialEncyclopedia

THE AI TIP WRITES SECTION 38: HOW CAN IT BE COLORED WITHOUT A HINT? HEX COLOUR METHOD

2026-10-4 9:27:55

TutorialEncyclopedia

SECTION 39 OF THE AI TIP: GOODBYE PLASTIC AI! TWO BIG PROSPECTS BLOCKING THE MAGIC. ONE KEY TO THE MOVIE

2026-10-4 10:31:24

❯
Search