Al-Ai-Phrasing Section 14: Progress! Three steps to teach you how to get the AI jsons right

Do you think it's easy for AI to map, but it's hard for "control"? Even a nano banana or a dream that can understand people's words often leads to a shift in style due to vague descriptions. This paper will reveal a set of Hollywood-level "JSON Structured Thinking" that will teach you how to use logic to lock up the aesthetics, eliminate deep differences in video generation, and create real video- and video-level work。

INTRODUCTION: WHY IS YOUR PICTURE "ASSEMBLE"

In this time of the explosion of AI tools, tools like Nano Banana, Midjourney, or dream, already understand very complex natural languages. But most of the creators are still at the “tick card” stage:

Lucky, come out with a map。

Bad luck. I can't feel it。

And the deadliest thing is, when you want to make a series of graphics or videos, the style of the picture is always left and right, and it's not uniform。

The problem lies not in the tools, but in the way you communicate. Natural languages are scattered, while industrial standards are stringent。

WHAT WE'RE SHARING TODAY IS NOT A SIMPLE SPELL, BUT A "JSON PUNISHMENT." THIS IS A LOGIC THAT “CODES” THE DIRECTOR'S THINKING, AND IT CAN FORCE THE ATTENTION OF AI TO THE CORE PARAMETERS OF PHOTOGRAPHY, PHOTOS, IMAGES, ETC., WHICH, REGARDLESS OF THE MODEL YOU USE, CAN PRODUCE STABLE, UNIFORM, AND HIGHLY FILM-SENSITIVE IMAGES。

Phase 1: The Aesthetic Reverse

Al-Ai-Phrasing Section 14: Progress! Three steps to teach you how to get the AI jsons right

"Shawshank's Redemption" screenshot and image analysis

Before using any tool, we must establish a common “visible vocabulary”. The vocabulary is not dependent on specific software, but rather on the photography patterns of the physical world。

Physical media: granular sense of the film

Natural language models often paint images too clean, leading to a “plastic sense”. We need to introduce the concept of media:

Film Look (film sense): Keywords such as Kodak Vision 3, Halation (happiness), Film Grain (particle). It'll give the image a sense of air。

Digital Look (digital sense): Keywords such as Arri Alexa 65, Clean Sharp Focus. It's for science fiction or modern business。

Al-Ai-Phrasing Section 14: Progress! Three steps to teach you how to get the AI jsons right

Film Perceptions/Digital Perceptions

2. Lens language: key to breaking the plane

Anamorphic Lens, a film-sense "nuclear weapons." It brings Oval Bokeh and lateral dizziness, which opens the gap between ordinary photographs at short notice。

Al-Ai-Phrasing Section 14: Progress! Three steps to teach you how to get the AI jsons right

Telephoto Lens (long focus): Compress space, highlight the subject, and create an advanced sense of alienation。

3. Photologic logic: mood containers

Volumetric Lighting: To give dust and media in the air, the light has shape。

Rembrandt Lighting (Lembrandt Light): Classic triangle light, which gives people a very dramatic face。

Al-Ai-Phrasing Section 14: Progress! Three steps to teach you how to get the AI jsons right

Phase 2: Construction of the JSON middle stage (The Logic Lock)

This is at the heart of the curriculum. We don't feed JSON directly to the drawing AI, but to ChatGPT/Gemini as a generator Prompt "The Logical Bones."。

GENERIC JSON TEMPLATE

Please keep the following structure as your "director console":
{
Project_Settings: {)
"Style_Anchor": 'Cyberpunk Neo-Noir' // style anchor
Aspect_Ratio: "21:9 UltraWidescreen"/ / Widescreen
},
Subject_Core: {
"Character": "...", / / Subject Description
"Attire":'..., / / clothing details
"Action": "..."/ specific actions
},
"Environ_Layer": {
"Location":'..., / / scene
"Weather": "..." // Weather and atmosphere
"Background_Details": "..."/ Background elements
},
"Cinematography_Lock": {/ / [stylelock] core area
"Camera_Gear": "IMAX 70mm Film Camera",
“Lens_Type”: “Panavision Anamorphic Lens”,
"Lighting_Scheme": "Neon sidelighting mixed withrain fog,",
"Color_Grading": "Teal and Orange, High Contracting, Black Bypass"
}
}
Why would you do that

For nano banana/midjourney: They're good at understanding semantics. JSON structures force LLM to naturally integrate the parameters in Cinematography_Lock into the environment, rather than omit details, when writing “small essays”。

For FLUX/SD: They're good at tagging. JSON structure allows LLM to extract a high weight Tag at the top of Prompt。

It's like putting a string on AI, no matter how you change the lead, the quality of the image will never change。

Phase III: Full-model operational adaptation (The Exchange)

WITH THE STRUCTURE, WE JUST NEED TO LET LLM BE THE "TRANSLATOR."。

Operational Command (Prompt Engineering)

Enter your AI assistant (ChatGPT/Gemini):

"YOU'RE A FILM MASTER. PLEASE WRITE A HINT FOR THE [TARGET MODEL NAME] BASED ON THE JSON DATA I HAVE PROVIDED。

Reads the parameters in Cinematography_Lock to ensure that they dominate Prompt。

Read Subject_Layer and integrate it into the scene。

In the case of nano banana, write a very graphic paragraph; in the case of Stable Diffusion, output the English Label Group.”

Operational effects

in a narrow, rain-drenched ally of futistic Kowloon, a small family-surgeon against a Wet brick wall

Al-Ai-Phrasing Section 14: Progress! Three steps to teach you how to get the AI jsons right

FLUX Output: “Cinematic still, IMAX 70mm, Panavision air, 50mm, oval bokeh, Thai and Orange, black by paper, solid ground

Al-Ai-Phrasing Section 14: Progress! Three steps to teach you how to get the AI jsons right

You see, the core logic is perfectly consistent, and we have managed all the tools with a set of ideas。

Phase 4: Art of moving the picture - the verb

WHEN WE FEED A HAPPY STATIC IMAGE TO THE VIDEO AI, DO NOT REPEAT THE PREVIOUS HINT。

ALL YOU HAVE TO DO IS TELL AI IN THE SIMPLEST WHITE WORD: HOW THE CAMERA GOES, HOW THE SUBJECT MOVES。

The core here is “deductive”

THE STATIC MAP ALREADY HAS THE LIGHT, THE IMAGE, THE COLOR, THE VIDEO AI. IF YOU REPEAT THE "SIBBUNK, NEON LIGHT" AGAIN, IT WILL INTERFERE WITH IT. WHAT YOU NEED TO ENTER IS A PURE ACTION ORDER。

An all-powerful “natural language” formula

Formula: [horizon motion] + [subject micromotion] + [environment climate]

Just fill it in and get a big sense:

The camera movement:

Want to show the big picture? Enter "Slow zoom out " (slowly pull out)。

Want to show your emotions? Enter "Slow zoom in " (slow approach)。

Want to show space? Enter " Pan right " 。

Subject micromove (rejection of ghost animals):

Person: “Hair blowing in wind”, “Looking around”。

"Running" or "Fighting." It's hard for AI now to deal with big moves, to write the more exaggerating, the faster it goes. Subtle movement is the real thing。

Environmental climate (increased sense of mobility):

Input: " Dust floating " , " Rain falling " , "Smoke rising " 。

3. Field demonstration

Like that western cowboy picture:

WRONG COMMAND: "A COWBOY RIDES A HORSE IN THE GRAND CANYON, SUNSET, MOVIE SENSE..."

Correct natural language instruction:

"Slow cinematic zoom out, wind blowing the dust, worse breathing, subtle coat movement." I'm not sure

Summarizing: Do this video step, forget complex parameters. As the director speaks to the photographer, the simplest English/Chinese command image is "move." The original painting determines the quality, and your natural language determines life。

Concluding remarks: from operator to director

THIS JSON WORKSTREAM ESSENTIALLY FREES YOU FROM YOUR CUMBERSOME “TICK CARD” AND FOCUSES YOUR ATTENTION ON THE REAL CORE OF CREATION — AESTHETIC AND NARRATIVE。

Aesthetics is your limit。

JSON IS YOUR FENCE。

LLM IS YOUR IMPLEMENTER。

AS LONG AS THE FUTURE AI TOOLS ARE LANGUAGE-DEPENDENT, THE STRUCTURED DIRECTOR THINKING WILL NEVER BE OBSOLETE. NOW, GO BUILD YOUR MOVIE UNIVERSE。

statement:The content of the source of public various media platforms, if the inclusion of the content violates your rights and interests, please contact the mailbox, this site will be the first time to deal with.
TutorialEncyclopedia

AI TIP CREATION SECTION 13: BREAKING THE "AI PLASTIC SENSE" - USING THE "FAKE CAUSE AND EFFECT" TO GIVE THE IMAGE TO THE SOUL

2026-9-23 9:34:15

Information

The novel multimodal recommendation system paradigm DiffMM allows the diffusion model to recommend short videos!

2024-7-8 8:55:43

Search