{"id":57264,"date":"2026-09-23T10:20:57","date_gmt":"2026-09-23T02:20:57","guid":{"rendered":"https:\/\/www.1ai.net\/?p=57264"},"modified":"2026-09-20T14:42:17","modified_gmt":"2026-09-20T06:42:17","slug":"ai%e6%8f%90%e7%a4%ba%e8%af%8d%e5%88%9b%e4%bd%9c%e7%ac%ac%e5%8d%81%e5%9b%9b%e8%8a%82%ef%bc%9a%e9%a9%be%e9%a9%adai%e7%9a%84%e5%b9%bb%e8%a7%89%e9%93%be%e5%8f%8d%e5%ba%94","status":"publish","type":"post","link":"https:\/\/www.1ai.net\/en\/57264.html","title":{"rendered":"Al-Ai-Phrasing Section 14: Progress! Three steps to teach you how to get the AI jsons right"},"content":{"rendered":"<p>Do you think it's easy for AI to map, but it's hard for \"control\"? Even a nano banana or a dream that can understand people's words often leads to a shift in style due to vague descriptions. This paper will reveal a set of Hollywood-level \"JSON Structured Thinking\" that will teach you how to use logic to lock up the aesthetics, eliminate deep differences in video generation, and create real video- and video-level work\u3002<\/p>\n<p><strong>INTRODUCTION: WHY IS YOUR PICTURE \"ASSEMBLE\"<\/strong><\/p>\n<p>In this time of the explosion of AI tools, tools like Nano Banana, Midjourney, or dream, already understand very complex natural languages. But most of the creators are still at the \u201ctick card\u201d stage:<\/p>\n<p>Lucky, come out with a map\u3002<\/p>\n<p>Bad luck. I can't feel it\u3002<\/p>\n<p>And the deadliest thing is, when you want to make a series of graphics or videos, the style of the picture is always left and right, and it's not uniform\u3002<\/p>\n<p>The problem lies not in the tools, but in the way you communicate. Natural languages are scattered, while industrial standards are stringent\u3002<\/p>\n<p>WHAT WE'RE SHARING TODAY IS NOT A SIMPLE SPELL, BUT A \"JSON PUNISHMENT.\" THIS IS A LOGIC THAT \u201cCODES\u201d THE DIRECTOR'S THINKING, AND IT CAN FORCE THE ATTENTION OF AI TO THE CORE PARAMETERS OF PHOTOGRAPHY, PHOTOS, IMAGES, ETC., WHICH, REGARDLESS OF THE MODEL YOU USE, CAN PRODUCE STABLE, UNIFORM, AND HIGHLY FILM-SENSITIVE IMAGES\u3002<\/p>\n<p><strong>Phase 1: The Aesthetic Reverse<\/strong><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-57265\" title=\"2103734bj00tl70iq00jhd000uxrp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/2103734bj00tl70iq00jhd000u000xrp.jpg\" alt=\"2103734bj00tl70iq00jhd000uxrp\" width=\"1080\" height=\"1215\" \/><\/p>\n<p>\"Shawshank's Redemption\" screenshot and image analysis<\/p>\n<p>Before using any tool, we must establish a common \u201cvisible vocabulary\u201d. The vocabulary is not dependent on specific software, but rather on the photography patterns of the physical world\u3002<\/p>\n<p><strong>Physical media: granular sense of the film<\/strong><\/p>\n<p>Natural language models often paint images too clean, leading to a \u201cplastic sense\u201d. We need to introduce the concept of media:<\/p>\n<p>Film Look (film sense): Keywords such as Kodak Vision 3, Halation (happiness), Film Grain (particle). It'll give the image a sense of air\u3002<\/p>\n<p>Digital Look (digital sense): Keywords such as Arri Alexa 65, Clean Sharp Focus. It's for science fiction or modern business\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-57268\" title=\"739d03c9j00tl70j40048d000v9009op\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/739d03c9j00tl70j40048d000v9009op.jpg\" alt=\"739d03c9j00tl70j40048d000v9009op\" width=\"1125\" height=\"348\" \/><\/p>\n<p>Film Perceptions\/Digital Perceptions<\/p>\n<p><strong>2. Lens language: key to breaking the plane<\/strong><\/p>\n<p>Anamorphic Lens, a film-sense \"nuclear weapons.\" It brings Oval Bokeh and lateral dizziness, which opens the gap between ordinary photographs at short notice\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-57266\" title=\"be 4e2286j00tl70ji01b4d000u0011bp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/be4e2286j00tl70ji01b4d000u0011bp.jpg\" alt=\"be 4e2286j00tl70ji01b4d000u0011bp\" width=\"1080\" height=\"1343\" \/><\/p>\n<p>Telephoto Lens (long focus): Compress space, highlight the subject, and create an advanced sense of alienation\u3002<\/p>\n<p><strong>3. Photologic logic: mood containers<\/strong><\/p>\n<p>Volumetric Lighting: To give dust and media in the air, the light has shape\u3002<\/p>\n<p>Rembrandt Lighting (Lembrandt Light): Classic triangle light, which gives people a very dramatic face\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-57267\" title=\"b82f621cj00tl70jx00j7d000u00140p\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/b82f621cj00tl70jx00j7d000u00140p.jpg\" alt=\"b82f621cj00tl70jx00j7d000u00140p\" width=\"1080\" height=\"1440\" \/><\/p>\n<p><strong>Phase 2: Construction of the JSON middle stage (The Logic Lock)<\/strong><\/p>\n<p>This is at the heart of the curriculum. We don't feed JSON directly to the drawing AI, but to ChatGPT\/Gemini as a generator <a href=\"https:\/\/www.1ai.net\/en\/tag\/prompt\" title=\"[View articles tagged with [Prompt]]\" target=\"_blank\" >Prompt<\/a> \"The Logical Bones.\"\u3002<\/p>\n<p><strong>GENERIC JSON TEMPLATE<\/strong><\/p>\n<p>Please keep the following structure as your \"director console\":<br \/>\n{<br \/>\nProject_Settings: {)<br \/>\n\"Style_Anchor\": 'Cyberpunk Neo-Noir' \/\/ style anchor<br \/>\nAspect_Ratio: \"21:9 UltraWidescreen\"\/ \/ Widescreen<br \/>\n},<br \/>\nSubject_Core: {<br \/>\n\"Character\": \"...\", \/ \/ Subject Description<br \/>\n\"Attire\":'..., \/ \/ clothing details<br \/>\n\"Action\": \"...\"\/ specific actions<br \/>\n},<br \/>\n\"Environ_Layer\": {<br \/>\n\"Location\":'..., \/ \/ scene<br \/>\n\"Weather\": \"...\" \/\/ Weather and atmosphere<br \/>\n\"Background_Details\": \"...\"\/ Background elements<br \/>\n},<br \/>\n\"Cinematography_Lock\": {\/ \/ [stylelock] core area<br \/>\n\"Camera_Gear\": \"IMAX 70mm Film Camera\",<br \/>\n\u201cLens_Type\u201d: \u201cPanavision Anamorphic Lens\u201d,<br \/>\n\"Lighting_Scheme\": \"Neon sidelighting mixed withrain fog,\",<br \/>\n\"Color_Grading\": \"Teal and Orange, High Contracting, Black Bypass\"<br \/>\n}<br \/>\n}<br \/>\n<strong>Why would you do that<\/strong><\/p>\n<p>For nano banana\/midjourney: They're good at understanding semantics. JSON structures force LLM to naturally integrate the parameters in Cinematography_Lock into the environment, rather than omit details, when writing \u201csmall essays\u201d\u3002<\/p>\n<p>For FLUX\/SD: They're good at tagging. JSON structure allows LLM to extract a high weight Tag at the top of Prompt\u3002<\/p>\n<p>It's like putting a string on AI, no matter how you change the lead, the quality of the image will never change\u3002<\/p>\n<p><strong>Phase III: Full-model operational adaptation (The Exchange)<\/strong><\/p>\n<p>WITH THE STRUCTURE, WE JUST NEED TO LET LLM BE THE \"TRANSLATOR.\"\u3002<\/p>\n<p><strong>Operational Command (Prompt Engineering)<\/strong><\/p>\n<p>Enter your AI assistant (ChatGPT\/Gemini):<\/p>\n<p>\"YOU'RE A FILM MASTER. PLEASE WRITE A HINT FOR THE [TARGET MODEL NAME] BASED ON THE JSON DATA I HAVE PROVIDED\u3002<\/p>\n<p>Reads the parameters in Cinematography_Lock to ensure that they dominate Prompt\u3002<\/p>\n<p>Read Subject_Layer and integrate it into the scene\u3002<\/p>\n<p>In the case of nano banana, write a very graphic paragraph; in the case of Stable Diffusion, output the English Label Group.\u201d<\/p>\n<p><strong>Operational effects<\/strong><\/p>\n<p>in a narrow, rain-drenched ally of futistic Kowloon, a small family-surgeon against a Wet brick wall<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-57269\" title=\"8755aa36j00tl709010ld000v90hfp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/8755ae36j00tl70l9010ld000v900hfp.jpg\" alt=\"8755aa36j00tl709010ld000v90hfp\" width=\"1125\" height=\"627\" \/><\/p>\n<p>FLUX Output: \u201cCinematic still, IMAX 70mm, Panavision air, 50mm, oval bokeh, Thai and Orange, black by paper, solid ground<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-57270\" title=\"7d8ec17bj00tl70ll00umd000v9000hfp\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/09\/7d8ec17bj00tl70ll00umd000v900hfp.jpg\" alt=\"7d8ec17bj00tl70ll00umd000v9000hfp\" width=\"1125\" height=\"627\" \/><\/p>\n<p>You see, the core logic is perfectly consistent, and we have managed all the tools with a set of ideas\u3002<\/p>\n<p><strong>Phase 4: Art of moving the picture - the verb<\/strong><\/p>\n<p>WHEN WE FEED A HAPPY STATIC IMAGE TO THE VIDEO AI, DO NOT REPEAT THE PREVIOUS HINT\u3002<\/p>\n<p>ALL YOU HAVE TO DO IS TELL AI IN THE SIMPLEST WHITE WORD: HOW THE CAMERA GOES, HOW THE SUBJECT MOVES\u3002<\/p>\n<p><strong>The core here is \u201cdeductive\u201d<\/strong><\/p>\n<p>THE STATIC MAP ALREADY HAS THE LIGHT, THE IMAGE, THE COLOR, THE VIDEO AI. IF YOU REPEAT THE \"SIBBUNK, NEON LIGHT\" AGAIN, IT WILL INTERFERE WITH IT. WHAT YOU NEED TO ENTER IS A PURE ACTION ORDER\u3002<\/p>\n<p><strong>An all-powerful \u201cnatural language\u201d formula<\/strong><\/p>\n<p>Formula: [horizon motion] + [subject micromotion] + [environment climate]<\/p>\n<p>Just fill it in and get a big sense:<\/p>\n<p>The camera movement:<\/p>\n<p>Want to show the big picture? Enter \"Slow zoom out \" (slowly pull out)\u3002<\/p>\n<p>Want to show your emotions? Enter \"Slow zoom in \" (slow approach)\u3002<\/p>\n<p>Want to show space? Enter \" Pan right \" \u3002<\/p>\n<p>Subject micromove (rejection of ghost animals):<\/p>\n<p>Person: \u201cHair blowing in wind\u201d, \u201cLooking around\u201d\u3002<\/p>\n<p>\"Running\" or \"Fighting.\" It's hard for AI now to deal with big moves, to write the more exaggerating, the faster it goes. Subtle movement is the real thing\u3002<\/p>\n<p>Environmental climate (increased sense of mobility):<\/p>\n<p>Input: \" Dust floating \" , \" Rain falling \" , \"Smoke rising \" \u3002<\/p>\n<p><strong>3. Field demonstration<\/strong><\/p>\n<p>Like that western cowboy picture:<\/p>\n<p>WRONG COMMAND: \"A COWBOY RIDES A HORSE IN THE GRAND CANYON, SUNSET, MOVIE SENSE...\"<\/p>\n<p>Correct natural language instruction:<\/p>\n<p>\"Slow cinematic zoom out, wind blowing the dust, worse breathing, subtle coat movement.\" I'm not sure<\/p>\n<p>Summarizing: Do this video step, forget complex parameters. As the director speaks to the photographer, the simplest English\/Chinese command image is \"move.\" The original painting determines the quality, and your natural language determines life\u3002<\/p>\n<p><strong>Concluding remarks: from operator to director<\/strong><\/p>\n<p>THIS JSON WORKSTREAM ESSENTIALLY FREES YOU FROM YOUR CUMBERSOME \u201cTICK CARD\u201d AND FOCUSES YOUR ATTENTION ON THE REAL CORE OF CREATION \u2014 AESTHETIC AND NARRATIVE\u3002<\/p>\n<p>Aesthetics is your limit\u3002<\/p>\n<p>JSON IS YOUR FENCE\u3002<\/p>\n<p>LLM IS YOUR IMPLEMENTER\u3002<\/p>\n<p>AS LONG AS THE FUTURE AI TOOLS ARE LANGUAGE-DEPENDENT, THE STRUCTURED DIRECTOR THINKING WILL NEVER BE OBSOLETE. NOW, GO BUILD YOUR MOVIE UNIVERSE\u3002<\/p>","protected":false},"excerpt":{"rendered":"<p>Do you think it's easy for AI to map, but it's hard for \"control\"? Even a nano banana or a dream that can understand people's words often leads to a shift in style due to vague descriptions. This paper will reveal a set of Hollywood-level \"JSON Structured Thinking\" that will teach you how to use logic to lock up the aesthetics, eliminate deep differences in video generation, and create real video- and video-level work. Introduction: Why is your picture \"Assemble\"? In this time of the explosion of AI tools, tools like Nano Banana, Midjourney, or dream, already understand very complex natural languages. But most of the creators are still at the \"tick card\" stage: luck, a map. Bad luck. I can't feel it<\/p>","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[149,144],"tags":[802,3149,491,192,6481],"collection":[],"class_list":{"0":"post-57264","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"hentry","6":"category-jiaocheng","7":"category-baike","8":"tag-ai","11":"tag-prompt","12":"tag-6481"},"acf":[],"_links":{"self":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts\/57264","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/comments?post=57264"}],"version-history":[{"count":0,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts\/57264\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/media?parent=57264"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/categories?post=57264"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/tags?post=57264"},{"taxonomy":"collection","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/collection?post=57264"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}