
THE GOAL OF THIS LESSON: A COMPUTER WITH ONLY A SYSTEM AND A GRAPHIC CARD, FOLLOWED BY AN AI VIDEO WITH STEREO, AND LEARNED THE OFFICIAL HINT -- H3 IS GOOD AND BAD, AND MOST OF IT IS ON IT. AT EACH STEP, IT WAS WRITTEN “WHAT TO DO, HOW TO DO, WHAT TO SEE AFTER”, WITHOUT ANY CODE BASIS。
one thing you have to know before you start: the native canvas of the local model is 768p。 The model was only trained on the short side of a canvas of 768, and the official 2K came out of another closed-source module. But 768p is not a local ceiling: the resolution can be hard and bigger (someone runs 1920 x 1088 on 5090, the painting is really better at a cost of 10 seconds and 58 minutes) and the community already has a ready-made local overscore stream of 768p to pull 1080p - a two-way cut-off cut-off. First, the division of the three modules is clear, and the official source is only the middle:
- H3-Context-IR - UNDERSTANDING AND STRUCTURING OF THE HINTS (REFORM YOUR DESCRIPTION INTO A MODEL-FRIENDLY FORMAT). ❌ CLOSED SOURCE, API ONLY
- H3-Base - 33B generate the subject, output 768p level audio video. Zenium Open Source, principal of the program
- H3-Regenerate-2K - Reborn 768p to 2K. ❌ Closed source, API only

So the phrase "local running H3 at no cost" basically works: drafts, 768 p is free, 1080 p can be solved locally by community super-points; only "official 2K" must go API. When to use it, there's a clear account at the end。
The data is from a RTX 4090 48GB, but it's not the threshold - the officially confirmed floor is 3060 12GB + 32GB RAMI don't know. From 3060, 5070 to 3090, 5090, how much can be quantified, how fast can you run, there is a table for step 0. The small amount of the presence only affects how much you can specify and how long you can wait, and does not affect whether you can run。
catalogs
- Knows a few terms (30 seconds)
- Step 0: Web version inspection + hardware self-check
- Step 1: Install ComfyUI
- Step 2: Download Model File
- STEP 3: RUN THROUGH THE FIRST VIDEO (T2V)
- Step 4: Understand parameters - hard rules of frame and resolution
- STEP 5: TUSHENG VIDEO (I2V)
- STEP 6: REFERENCE VIDEO (R2V)
- Step 7: Full Guide to Phrasing
- Step 8: Accelerate - SageAttention 24%, stack Cache up to 58%
- Check for common problems
- Sound effect, alignment ahead of expectations
- Local vs API, account
1. Understanding several terms (30 seconds)
- ComfyUI: RUN FREE OPEN-SOURCE WORKSTATION FOR AI GENERATION MODEL. THE INTERFACE IS NODE-CONNECTED, BUT YOU DON'T HAVE TO DO IT YOURSELF -- OFFICIAL TEMPLATES ARE SET, YOU CHANGE PARAMETERS。
- Workflow: a lined node map, equivalent to a “formula”. Loading template = Opens a ready-to-use formulation。
- Weight / Model File: Model body, suffix .safetensors Big paper. H3 needs four: DiT (painted), text encoder (readers with hints), video VAE and audio VAE (translating internal data into pixels and sound waves)。
- Quantitative: Technology to minimize the model。bf16 It's full of precisionint8,nvfp4 It is a compressed version, with very small loss of paint quality and a significant decline in the demand for visibility. Local deployments are mostly in quantitative form。
- T2V / I2V / R2VThree modes of generation - pure text generation, giving pictures to move, and targeting characters or styles for reference material (chart/video/audio)。
Step 0: Web version inspection + hardware self-check
Before we do it, it takes 10 minutes to check on the web page
40 GB weight first. H3 In the conch AI page version (domestic hailuoai.com, overseas hailuoai.video) you can direct the test: login → video generation → model selection MiniMax H3 POACH A PICTURE OF THE PERSON IN THE PICTURE WAVING AT THE CAMERA, WITH THE SOUND OF THE STREET ENVIRONMENT, AND GENERATE IT。
Look at three things:The subject's unstable, the camera's moving naturally, the sound and the image, rightI DON'T KNOW. THESE THREE ARE SATISFIED, THEN GO DOWN; IF NOT, YOU DON'T HAVE TO READ THE LESSON, SAVE THE REST OF THE DAY. THE SIZE OF THE PAGE VERSION AND THE API ARE TWO SETS OF ACCOUNTS, AND THE PAGE SIZE DOES NOT AFFECT THE MOVEMENT OF API。
Hardware self-check
- GraphicsMinimum NVIDIA 12GB display, recommended 24GB+. How to find: Windows press Ctrl+Shift+Esc → PERFORMANCE GPU → "PILOT GPU RAM"; OR COMMAND LINE LOSS nvidia-smi.
- MemoryMINIMUM 32GB, RECOMMENDED 64GB. TASK MANAGER & PERFORMANCE & MEMORY。
- Disk FreeMINIMUM 60 GB, RECOMMENDED 150 GB+. THE WEIGHT OF ABOUT 40 GB, PLUS THE GENERATION AND R2V WEIGHTS, WILL CONTINUE TO RISE。
- systems: Windows 10/11 or Linux.
- reticulation: For access to Hugging Face or its mirror image, see step 2 mirror scheme。
Right: What's your graphic card in
The CofyUI core developer officially confirmed the floor line:3060 12GB + 32GB RAM + a nice NVME solid and can run 480pI don't know. 12GB has a 12GB run, find your own line of action:


Data source: 5090 and 12GB cards from community measurements, 4090 mobile versions from 20-step measurements on SageAttention
Two generic cross-slotting reminders:
- System storage and solidity are as important as graphic cards。 H3 Runs on a consumption-grade card only with a layer load of "notable memory " , 32 GB memory is the bottom line, 64 GB from the face — 16 GB memory + large memory is not moving. The model goes from disk to memory, and the mechanical hard drive will keep the first run waiting for a few minutes, and the weight must be on NVME。
- 30/40 and 50 are a hidden difference: NVFP4 is only 50 (Blackwell) supported by hardware。 When loading nvfp4 weights on 3090/4090, ComfyUI follows a "simulator" - saves disks and visible occupancy, but before counting, depresses back to high accuracy, without taking time. So 50 users are bold enough to choose a community NVFP4 version of DiT (smaller and faster), and 30/40 users are more cost-effective to choose INT8/FP8。
Step 1: Installation of ComfyUI
Version shall be thallium 0.30.0, this is a hard threshold: the original support of H3 (4 dedicated nodes + official templates) was merged with Comfy-Org/CommyUI #15224 on 3 August 2026, without these nodes in the old version and without any working flow。
STATUS A: COMPLETE NEW INSTALLATION (RECOMMENDED FOR WHITE)
- Open the ComfyUI official network and download the corresponding systemdesktop versionInstallation package。
- Double-click installation, all the way by default. The installation automatically handles Python, PyTorch, CUDA dependency, which is why the desktop is friendly to new hands。
- Finish loading start. The first start will be initialized, and when it's finished。
Case B: ComfyUI
- desktop version: Check in the menu for updates to the latest stable version。
- manual guit deployment_Other Organiser don't pullonce again pip install -r requirements.txt.
- Integration pack (Autumn leaf pack, etc.): When the integration package is updated by the author or the kernel is updated in the starter. Note: Desktop and integration packages followSteadyrelease, individual nightly features may arrive a few days later。
Verify installation successful
Open Post Startup Browser http://127.0.0.1:8188(desktop directly pops up window) See node canvas. Confirmed version: Settings (low left corner gear) → On, version number ≥ 0.30.0. If you're lower, go back and update. Don't tryThe most high-frequency error of the MiniMaxH3ImageToVideo node, 99% is a problem of the version.
Step 2: Download model files
The model hosts the Comfy-Org/ MiniMax-H3 repository in Hugging Face。
warning: do not close the entire warehouse。 THE ORIGINAL WAREHOUSE IS ABOUT 318 GB, AND YOU ONLY NEED FOUR FILES。
Domestic download speed up
Link any Hugging Face i don't know Replace cf-miror.comAnd it's a mirror image of the country, and it's a magnitude difference. The large 20GB file proposes a download tool (IDM, aria2, or browser to bring back the download) to support the break-up, without having to repeat it。
We'll use the order line. One order will be precise, and four files will be executed set HF_ENDPOINT=https://hf-miror.com, Linux/macOS export):
pip install-U hugglingface_hub
hf download Comfy-Org/ MiniMax-H3 \
diffusion_models/minimax_h3_fl2va_pruned_int8_convrot.safetensors\
text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors \
vae/minimax_h3_video_vae_fp16.safetensors \
vae/minimax_h3_udio_vae_fp32.safetensors \
–local-dir ComfyUI/models/
4 REQUIRED T2V / I2V (APPROXIMATELY 39.6 GB)

TWO VAES All must go down- VIDEO VAE OUT, AUDIO VAE OUT, LESS VOICE VAE YOU'LL GET A SILENT VIDEO。

After the release, the directory should be long:
I don't know
ideas - models/
_diffusion_models/
│ -minimax_h3_fl2va_pruned_int8_convrot.safetensors
ideas -text_encoders/
│ qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
> vae/
_minimax_h3_video_vae_fp16.safetensors
─-minimax_h3_audio_vae_fp32.safetensors
desktop user note: the model directory may be searched for "model paths " in the settings under the path selected for installation。
More lazy: Skip manual download, direct step 3 - When loading the template, ComfyUI will pop a window to list the missing model and give the download button and click automatically down to the right position. The value of manual downloads is that they can be sent through mirrors and breakpoints, and people with a volatile network recommend it manually。
An overview of all official variants (options as required, most not available)
Diffusion model (choose one; want to play R2V to play ref2va):
- minimax_h3_fl2va_pruned_int8_convrot(19.5 GB) - ✅ recommend,T2V/I2V BEST BALANCE
- minimax_h3_fl2va_pruned_fp8_scaled(19.5 GB) - ALTERNATIVE QUANTITATIVE FORMAT
- minimax_h3_fl2va_int8_convrot(31.7 GB) - STANDARD INT8 (UNCUTED)
- minimax_h3_fl2va_bf16(61.7 GB) - FULL PRECISION, FINE CALL
- minimax_h3_ref2va_pruned_int8_convrot(19.5 GB) - R2V MODE SPECIALStep 6 will be used
- minimax_h3_ref2va_bf16(61.7 GB)— R2V FULL PRECISION
Text encoder (select):
- qwen3vl_32b_minimax_h3_nvfp4_awq(14.6 GB) - ✅ recommendANY GPU CAN RUN
- qwen3vl_32b_minimax_h3_int8_convrot(25.3 GB) - INT8
- qwen3vl_32b_minimax_h3_bf16(48.0 GB) - FULL PRECISION
Why do you choose this(jumping does not affect operations):
- H3 Eats a visible head of encoder, not 33B's Dit。 The encoder directly uses the full weight of Qwen3-VL-32B (62.13 GiB under bf16, larger than the DiT of 61.73 GiB), where a 32B visual language model is only a "problem." So the encoder has to select the unvfp4 version to be quantified. The online circulation of “codifier” 51.5GB is incorrect, based on the size of the official warehouse file。
- pruned = cutting down the AdaLN branch of 13B。 6173 - 37.46 ≈ 24.3 GiB, exactly the 13B. The official confirmation that this part of the modem output can be expected to be a cache and that pure reasoning does not need to be loaded at all. Conclusion: Pruned for reasoning alone and fine-tuned for full weight。
Quantification of community selection by card slot (optional extra meals)
Official pruned_int8+nvfp4_awq All-eat combinations. Any slot is recommended to run first. After running, think faster, save it and come here for dinner. The following are community-based third-party conversions (unofficial publication, drawings and permissions for self-checking of warehouses README), which are directly loaded by the original ComfyUI:

When changing community weights, remember one:Two VAEs, always in the official original- They're small, and quantifying them can only hurt paint and sound。
5. STEP 3: RUN THROUGH THE FIRST VIDEO (T2V)
Load official templates
- Open CommyUI Top MenuWorkflow → Browse Templates → videoClassification, search "Mini Max H3"。
- You'll see 6 templatesDon't get me wrong
- Mini Max H3 Text to Video - Local, Vincent videoStart with this
- Mini Max H3 Image to Video - Local, Tusheng video (first frame/tail frame)
- Mini Max H3 Reference to Video – local, reference-based video, with an additional ref2va weight
- api_minimax_h3_t2v / r2v / flf2v – API Edition to Mini Max Cloud, to fill API key, ignore first
- click on Text to VideoI don't know. If the bullet window hint is missing the model, download it by hint; step 2 is manually set and is ready (to restart the ComfyUI - model list scanned on startup without recognition)。
You know the key nodes in the work stream
There's a line of nodes on the canvas when the template is loaded. All you need to know is these:

- UNETLoaderSelect the DiT weight to confirm the bottom box as minimax_h3_fl2va_pruned_int8_convrot.safetensors.
- Text Encoder Load Node: Confirm Selected qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors.
- Two VAE LoadersVIDEO VAE AND AUDIO VAE ARE SELECTED SEPARATELY。
- Resolutory Selector: Control output resolution (the width ratio, Megapixels, two parameters calculated as wide). Template default is a quick preview specificationDon't move for the first time.
- promptword box (positive)Write what you want。
- Length/ Frame Input: determine the length of the video, see step 4。
- SaveVideo: OUTPUT END OF MP4。
Run
In the prompting box, write an English scene (not to mention grammar, step 7 is a positive lesson):
A little walkers along a long street at night, no light lights reflecting in puddles, soft rain sounds.
point (in space or time) Run.
You should see something
- THE PROGRESS BAR IS FIRST LOADED IN THE MODEL (APPROXIMATELY 34 GB FROM DISK TO DISPLAY/RAM) AND THEN GRADUALLY SAMPLED。
- It's normal for the first time to be very slow: 4090 48G LIVE HEAD RUN, 644 SECONDS (INCLUDING LOADING), 425 SECONDS WITH THE SAME SPECIFICATION; 768 X 432 SMALL SIZE HEAT RUN, ONLY ABOUT 99 SECONDS. OTHER SLOT ALIGNMENTS ARE EXPECTED: 5090 OUT OF 5 SECONDS APPROXIMATELY 108 SECONDS; 16 GB CARDS (E.G. 4090 MOBILE VERSION) OUT OF 960 X 540 OUT OF 5 SECONDS APPROXIMATELY 182 SECONDS; 12 GB CARDS ARE LOADED BY MEMORY STRATIFICATION AND 864 X 480 OUT OF 90 FRAMES ABOUT 6 MINUTES. THE FIRST RUN IS TO ADD A FEW MORE MINUTES OF LOADING TIME TO THESE NUMBERS, MAKING TEA, ETC。
- After the video preview of the SaveVideo node:There's pictures, there's voicesI DON'T KNOW. SOUND AND IMAGE ARE GENERATED IN THE SAME FRONT-TO-FACE TRANSMISSION, WHICH IS THE FUNDAMENTAL POINT OF THE H3 DISTINCTION FROM A POST-GENERATION VOICE MODEL。
- If the console is brushed Input tensors must be in dtype of torch——It's not a mistakei don't know. the official document confirms that this is a normal hint for some layers back to standard attention, and that it is not affected。
It's over. The deployment phase is over. It goes from "can run" to "can use."。
Step 4: Hard rules for understanding parameters - frames and resolutions
These two rules read the code of H3 node of ComfyUI directly, not experiential。
rule i: the frame number must fall on the 17k+5 grid and be silently adsorbed
Source logic is a line:
♪ When n n % 17! ♪
n + = 1
The frame you fill out, if it's not legal, it willSniff up to the nearest legal value, no error, no hintI don't know. Measure 61 frames, get 73 frames。

the training area is 124-362 frames, 24fps below 5.17 to 15.08 seconds;all legal slots in this compartment:

60 seconds and only 15.08 seconds, the longer content needs to be filmed and edited. the node tooltip is ~124-362, longer is unsettled - less than 124 frames can run (the grid starts from 5, 22, 39 ...), but quality is not assured outside the training distribution. the day-to-day picks it out of this table, so don't let it suck -- the six seconds you think and the actual 6.58 seconds, in a card-faced scene, is a disaster。
Rule II: The resolution's "native canvas" is short edge 768
Model PressShort edge 768, area limit 768 x 1344, multiple of 32 per axis"Train." Operational recommendations:
- Try the hint: The template default fast preview specification (e.g. 768 x 432), approximately 99 seconds/bar, fast iterative。
- Spectrum: Resolute Selector’s Megapixels 1.0,16:9 1344 x 768- It's the original canvas, the best quality。
- floor 384p:256p and lowerTotal failureIt's official confirmation, not your configuration。
- It's not impossible, it's uneconomical: the model has not been trained in a larger painting, but the hard drive can come up with a piece — a 10-second piece of 1920 x 1088, which is actually better than 768p, at a cost of 58 minutes. community consensus is..Low Resolution Generation + Local Excess: 768p scale is followed by a LTX 2.3 or Wan 2.2 5B ultra-spectrum workflow to 1080p, and the same 1920 x 1080 finished product is reduced from approximately 700 seconds to 161 seconds, 6 times faster, and the paint is close to the status profile. Want the official original 2K, take the API regeneration -- three paths to the end。
Two realistic patterns of presence and time-consuming
It's a model, not a painting。 The peaks of 768 x 432 and 1152 x 640 were almost identical (the difference was < 1%). Process-level detail: The sampling phase is stable on 19 GB, the coding/decode phase is running at once to 32 GB - this is the reading of 48 G card unrun, the dynamic load of ComfyUI adapts to the empty visibles, and the small visible cards can run by switching to and from, but only slower. Meaning: The model can be installed, the resolution can be driven to the original canvas; it can't be installed, the lower resolution can't save you, but a smaller quantitative version. OOM does not waste time changing sizes。
Time-consuming patterns: Pixel dimensions are calculated and time-consuming。 1152 x 640 relative 768 x 432 is 2.22 pixels, which takes only 1.86 times more time, sublinearity, and high resolution is more intuitive than intuitive; but the length of time is the opposite - attention is ultra-linear, double the length, more than double the time。

Specification scale measurements (4090 48G, SageAttention cuda++ + acceleration):

Read the table three times:Five seconds, five minutes, 10 seconds, 12 minutes, 15 seconds, one 19 minutesI don't know. Note also that 1536x864 relative 1152x640 is 1.8 pixels, time-consuming close to double the same time — the next linear dividend of the painting is gone, and it is confirmed once again that the supernatural canvas is not available. The replicability of the data has been tested: reruns with the specifications on the next day, 320.11s against 32.05s, error 0.6%。
STEP 5: TUSHENG VIDEO (I2V)
Purpose: To move a ready-made image, or to replace the middle motion with two at the end。
- Template Library Loading Mini Max H3 Image to VideoI don't know. The weight and T2V are identical (fl2va) and do not need to be downloaded again。
- use LoadImage We'll upload your map MiniMaxH3ImageToVideo Node first_frame Enter。
- last_frame It's optional: ONLY FOR THE FRAME = FROM THIS MAP ONWARDS; ALL AT THE END = THE MODEL DISPLAYS A CONSISTENT MOVEMENT BETWEEN THE TWO GRAPHS (THE "A TO B" LENS SUITABLE FOR PRODUCT ROTATION, ATTITUDE CHANGE, ETC.); ONLY FOR THE FRAME = THE MODEL REVERSES A REASONABLE OPENING, AND EVENTUALLY FALLS ON YOUR GRAPH。

- THE HINT IS WRITTEN IN THE 7-STEP I2VA FORMAT (NEEDS A LINE FIRST ALIGNMENT COMMAND)。
- The input diagram will be adapted to generate resolution and too much difference will be processed --Before uploading, customize the map to match the width of the outputSaved by accidental cutting。
STEP 6: REFERENCE VIDEO (R2V)
USE: LOCKING A CHARACTER, A PAINTING STYLE, AN ACTION, A MIRROR OR A SOUND TO MAKE THEM APPEAR IN A NEW VIDEO. THIS IS THE STRONGEST AND MOST COMPLEX PATTERN OF H3。
Ready
R2VAnother weight ref2va, AND T2V/I2V fl2va Not universal:
- Go back to Comfy-Org/ MiniMax-H3 download minimax_h3_ref2va_pruned_int8_convrot.safetensors(19.5 GB), SAME Photo by Flickr user ComfyUI/models/diffusion_models/.
- ENCODERS AND TWO VAES ARE REUSED WITHOUT RESET。
- Template Library Loading Mini Max H3 Reference to Video, confirm UNETLoader cuts the ref2va weight。
Use rules
- Quantity ceiling: Up to 9 reference charts, 3 reference videos (each with its own track), 3 independent reference audio。
- Use tab references in connect order: The first is, the second is Take that kind of push. These labels must be used to name names in the message。
- For each reference "Piracy": specify which reference line is what — look, style, action, mirror or sound. The authorities have made it clear that the visible assignment is much more effective than putting a pile of material behind it。
- {\bord0\shad0\alphah3d}ref_image_size parameter:watch(Default) zoom in to generate resolution, fastmax keeps a maximum of 2048 px short edges, and the role is more solidbut token, every step of the sample is completeI don't know. Daily watchOnly when your face can't be locked max.
STRUCTURAL DIFFERENCES IN R2V HINTS
R2V 'S COMPLETE HINT IS A SIX-PART STYLE (THREE MORE THAN THE T2V FIELD) WITH A FIXED ORDER:subject_definitions(Defines labels and features for each reference) summary(Summary paragraph beginning with the type of task in square brackets) retention_analysis(label-by-label declaration of retention) detailed_description(Major, 350-500 English, opening style before [Shot 1]) overall_sundscape → no, no, no.

Each of the four labels:
- - What to repeat in a film: people, scenes, costumes, styles, actions
- - A map is directly used as a frame or structure anchor
- - Total level of relationship: edited source, starting point of renewal, source of cut rhythm
- - Audios copied or referenced
One of the easiest things to get wrong:Pictures are only used to define roles or styles , write in the source In the definitionI don't know. This is the "This is the frame" scene。
For the other two paragraphs:summary Click task type in square brackets at the beginning, multiple + Company--keepframe command / i'm sorry / i'm sorry / video conversion / audio return / audio referencei don't know. adjudication: reference video provides only mirrors and rhythms = reference promotion, really changing the video eding。retention_analysis Gives each label a fixed relationship word: the visible content is used full_preserved / i'm sorry / attribute_transfer / weak_referenceAudio full_copy / i'm sorry / reference / weak_reference.
It's the first time that this format has receded, but it's the output format of the official pay module Context-IR -- you're a handwritten whore. The rookies don't have to come up with six paragraphs at a time: first, replace the material with the illustrative hints that the template contains, then sew the entire case-by-case version of the official R2V guide, and write a template at a time, with a few words at a time. Just remember:Every reference in subject_definitions It is marked with a label and a characterization, which is used in all subsequent paragraphs to describe how much it has been retained。
Step 7: A guide to the integrity of the narrative
H3 IS A STRUCTURED SET OF OFFICIAL DEFINITIONS (ORIGINAL VERSION OF THE OFFICIAL GUIDE). A SINGLE WORD CAN ALSO COME OUT, BUT MULTIPLE LENSES, CHARACTER PAIRS, SPECIFIED MIRRORS, CARD TIME POINTS MUST BE FORMATTED。All written in English except for the original language of white and graphic text。
9.1 Bones: three fields
[Shot 1]..
overall_sundscape:
no, no, no
- integrated_multimedia_description - Subject. It's time-lined, action, camera, talking person, confidant, inside sound。
- overall_sundscape - 1-4 sentence, sound of the environment and action (wind and rain, footsteps, clothing friction, breathing, laughter). Don't write: white, singing, painting music。
- no, no, no - 1 to 3 words, music (a character who can't hear, only an audience): instruments, speed, rhythm, power and power. Don't write: emotional words, "sorted" "historic" and explain the role of music。
Write the last two fields without content N/A(Soundscape only if the user expressly requests full silence。

9.2 Camera: [Shot N] and switch time
- [Shot 1] Initial answerWhole style + initial diagram, without a time stamp. Style words:Cinematic,live-action,2D-animatized,3D CG,playmation,watercolor,video field.
- Follow-up camera tape strict incremental transition time:[Shot 2] At 00:03.500, the camera cuts to..
- Toggle verb:i mean, the camera cups to / the shot transfers to / you know, the shot switches to;cross-dissolve, fade, wipe.
- When do you cut the camera: Switch should introduce new information (new subject, new space, new perspective, new time). Just pull or fine-tune the angle, use the mirror, don't cut it。
9.3 Mirror: Type + Range + Speed
Writes a natural English sentence in the lens and does not stack tags at the end of the sentence。

Range:with small amplitude / ♪ with broad angle ♪;velocity:at slow speed / i don't know, at lastI don't know. Medium range, normal speedDirectly omitted.
The camera pushes in with small movements at low speed told the fed better in her hands.
The game is right with broad expression at last, revealing the open door.
The game holds a status shot as the runner exits the fire.
9.4 CONVERSATION: SPEAKER ID + TAG
List of rules:
- ID:(S1),(S2);many-person (S1, S2);CROSS LENS ID REMAINS UNCHANGED; never a silent character is numbered。
- For the first time, the speaker is present at the foot anchor: type of role, age, sex, presence in the picture, sound, sound, speed, accent。
- IDENTITY, ID, ACTION, TONE Outside; Only language tags and lines themselvesNo change, no translation, no retention of original markers.
The young woman with a quiet, Breathy voice says:
[English] Wait for us

- An external voiceIt has to be fixed says in an off-screen voiceoverAnd the people's lips are not moving:
The man (S1) says in an off-screen voiceover: [English] I still remember that road.
- LineCross the cut point: to be placed in both sections of the connection and clearly write the audio series (e. g the carries over from the previous shot); lines are spokenSnippetsWith。
- Image Text(Signatures, subtitles, neon lights): Put in double quotation sign in English and leave the original untranslated -A red neon sign reading "in business" brings above the doorway.
9.5 A COMPLETE COPYABLE EXAMPLE (OFFICIAL T2VA CASE)
live-action: [Shot 1] Live-action, a medium-widge shot justices a bullet opening the system of a small street before sunrise.
wooden begins open over a quiet set as lessons clear only inside the bakery.
no, no, no, no, no.
REPLACED WITH YOUR SCENE, IT'S A GOOD H3 HINT。
9.6 Three variants with graphs: start with one additional line alignment command
I2V FAMILY (I2VA / FL2VA / L2VA) IN THREE FIELDSBeforeAn additional line of fixed format alignment commands, followed by an empty line:
I2VA ONLY- Fixed:
For the target video, at 0.00 seconds into the target video, is fully referred.
Description of the structure: First the style, the main body, the structure in the anchor map, then how the action will be carried out (first frame anchor embezzled action start, continuous development results). Role clothing, colour, key items, spatial relationship and graphic consistency。
FL2VA— Declare the point at which each of the two charts corresponds:
How the reference pictures aligns with the target video — locals with the 0.00-second mark of the target video; locals with the 8.00-second mark of the target video.
Don't describe the two images in the textWrite the path that connects them(THE INITIAL FRAME STATE IS THE VISIBLE INTERMEDIATE CHANGE THE GAP GRADUALLY NARROWS THE END FRAME STATE). FL2VA TRYS TO BE SINGLE AND ALLOWS THE MODEL TO INSERT A CONTINUOUS VALUE。
L2VA ONLY_Text of the image:
How the reference pictures align with the target video along with the 6.00-second mark of the target video.
Describe the structure: deduce a reasonable pre-state, a clear movement and a transition path, a last mirror gradually shrunk down on the map。
time S.SS It must be accurate to two decimals and consistent with your actual video time (converted against the 6-step frame table)。
9.7 Implicit laziness
This is how the official design is: the API version of H3 has a closed-source Context-IR module dedicated to "reforming people's words into the above-mentioned formats," which is not available locally, but you can have any LLM on your behalf. Throw the official guide link with your idea:
Learn how to write T2V papers from https://huggingface.co/MiniMaxi/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
then my idea into the three-field H3 policy format:
[Your Chinese thought]
The results are checked in three places: if the lines have been rewrited (must be in the original language), if the time stamp of the lens is increasing, and if there are mood words in the total length of the time and in the play field。
I DON'T WANT TO RELY ON THE LLM'S TOUCH, BUT THERE'S A CORRECT LAZY PATH:Cost-IR API(¥5.80/million token input) Throw in your big whites and material, and it returns the enhanced hint in standard format -- this is the closed-source module of the API version that rewrites your tip. Use the return result as a template, and change the words to a few words at a time, much faster than learning from zero, as is recognized by local ComfyUI。
Step 8: Accelerating - SageAttention 24%, stacking Cache up to 58%
Run for it and get this far. Accelerator in two layers, superseding:SageAttention counts every step faster, and the Cache node skips the unnecessary step。 First tier - same seed and workstream (1152 x 640/124 frame) A/B Measurement 24%:
- LoadSageAttention: download from SageAttention returns the wiel that matches your PyTorch/ CUDA version (torch and cu versions in the file name, match then down), and then pip install I don't know. Desktop user executes the Python environment that is used by ComfyUI in its built-in terminal。version approval 2.x- The back fp8 mode is only available in 2.x, 1.x has no kernel at all; these wiels are pre-compiled by 2.x and are directly the most economical. It really has to come from the source code itself, with a hard threshold: sm_89(40) requires CUDA ≥12.4 (written in setup.py), and we can't make it in 12.0, and we can't get it up to 12.8。
- Load KJNodes Node Pack:CommyUI Manager Search Photo by ComfyUI-KJNodes installing;or manual it's not like you're in love with me, it's not like you're in love with me until (a time) I'm sorryStart again。
- Connect: Workstream Riga Patch Page Attention KJ Node in UNETLoader and BasicGuider Between - UNETLoader 's model output → Patch node model input, Patch node model output → BasicGuider input. Only guider needs a patch, scheduler don't move。

- Select ModeDirect use autumn all right. it automatically falls on the speediest path by graphics -- 4090 on the auto is equal to qk_int8_pv_fp8_cuda++(checked source code: the sm89 branch returns exactly the same fp8 kernel, manual selection is not bad but unnecessary)30 is card (3090/3060 et al. Ampere) without FP8 hardware units, auto will reach the int8 path available to Ampere, and there is still a considerable acceleration. I really need to lock it manually qk_int8_pv_fp8_cuda++, provided that step 1 is indeed SageAttention 2.x-1.x without this kernel, the selection is a misstatement or a silent retreat。
- Normal Run。Lazy replacement: do not want node to start ComfyUI plus _other organiser Parameter global opening (the desktop version is added in the set startup parameter). The two pits must be clear:THIS PARAMETER AND KJ TWO AND ONE, NEVER OPEN AT THE SAME TIME— Both sides repeat patches, while community feedback slows; and no start-up mode is available and only the default path is followed. To control the nodes, to save things, to use parameters。
Second level: Cache Node, not hard to count
SageAttention is the result of faster counting every step, and Cache-like nodes are another idea: the adjacent output of the diffuse sample is often almost the same, and changes to a certain extent are directly replicating the previous stepIt doesn't countI DON'T KNOW. TWO LAYERS SUPERHEAVY, THE SAME 362 WITH A FULL LENGTH PIECE (1152 X 640, 4090 48G):
- SageAttention Only – 1183.6s (no acceleration baseline 21.9%; and 1158s in step 4 are measured in different batches and 2% inter- batch fluctuations are normal)
- Sage + EASYCache – 899.5s (reducing 24.0% based on Sage)
- Sage + TeaCache – 500.5s (sage save 57.7%More than just driving Sage)
EasyCache is the ComfyUI native nodeNo third-party bag: add one to the canvas EasyCache Node, serial to model link (and Sage patch node lined around), key parameter is i'm sorry(Other start_percent / end_percent / verbose Three, normally, do not move. Default 0.2 can be used directly. Note that the 899.5s above are measured under a more conservative 0.1 - the more the threshold jumps, the faster the 124 frame shorts are measured 0.2 run 215.7s, 0.1 run 260.0s, with a default value that will be faster than the figures in the text and a slightly greater loss. TeaCache H3 SPECIALIZEDIcyoung/CommyUI-Mini MaxH3-TeaCache Mini MaxH3TeaCache), the other option is lihaoyun6/CommyUI-MiniMaxH3-Cache Mini MaxH3Cache).Don't pretend to be Manager's universal ComfyUI-TeaCache (welltop-cn)– Its support list stops at the FLUX/ HiDream generation, without H3, and when you're finished, you find no point to connect to the H3 link. Key parameters are: rel_l1_threshThe more you jump, the more you jump。
There is no free lunch - Cache's acceleration comes from skipping real calculations, but the price depends on the place. We checked the frame by frame under R2V reference pattern:Faces don't change(Same face, same hair, same dress structure, steady from head to end) What really falls is the “activity” of the picture — the near-stilling “deep” frame ratio from 13.8% to 41.5%, which is easier to break. So when Cache is in a film, don't look at her face, see if the picture's up and down and down。
The other one is ahead:Cache changed to produce the results themselves, not “run faster on the same note”。 As seen, TeaCache output to SIM with no Cache baseline is only 0.78 (EasyCache is 0.96) - so the "draft on TeaCache pick-up, face-to-face reassembly" intuitive play is not working: instead of configuration, the image you picked changed. The right one:Sage + TeaCache, do not expect the draft image to be reproduced as it is; positive film that requires it, from selection to finalization, with the same conservative configuration(Sage + EASYCache or simply Sage). Thresholds are down, they're more conservative, they're smaller。
11. Inventory of frequently asked questions
Q: Tip missing Mini MaxH3ImageToVideo / EmttyMini MaxH3LatentAV A: ComfyUI version < 0.30.0, upgrade. Most high-frequency problems. Desktop/Integrator Packages follow the steady release and are updated and retested。
Q: Bomb / CUDA out of memory A: ORDERED — 1 pruned_int8_convrot DIT+ nvfp4_awq Encoders do not miss the smallest official combination, bf16;2 turn off other visible memory-eating programs (games, another ComfyUI, running model browser labels);3 saves in the system at above 32GB, 12GB cards are loaded by memory fractions;4 not yet, replaces community INT4 quantification (end of step 2). Remember the pattern:LOWER RESOLUTION WON'T SAVE OOMIt's the model itself。
Q: OUTPUT VIDEO WITHOUT SOUND A: BOTH VAES MUST BE LOADED..video_vae_fp16 and audio_vae_fp32And confirm that there's work in it VAEDECodeAudio Node attached SaveVideoI don't know. Official templates are usually not missing, and it is easier to delete errors when you change your own workflow。
Q: GENERATE COMPLETELY FAILED, RESOLUTION IS VERY SMALL A: H3 lowest 384p, 256p and the following are inevitable failures. Preset with the resolution of an official template。
Q: THE VIDEO IS DIFFERENT FROM THE ONE I FILLED OUT A: Frames are adsorbed to the 17k+5 grid (step 6) with a single cap of 15.08 seconds. Pick directly from the legal frame list and do not fill in any value。
Q: R2V TEMPLATE "NO MODEL FOUND" A: R2V ref2va WEIGHTS, AND T2V/I2V fl2va It's two, downloading alone (step 6)。
Q:R2V PARTICULARLY SLOW A: INSPECTION {\bord0\shad0\alphah3d}ref_image_size Did you choose max—refer token, take every step of the samplemax It's several times slower. Daily watch.
Q: REFERENCE VIDEO/INPUT CHART OUT OF SHAPE, DECORATION A: THE REFERENCE VIDEO WILL BE PRESSURIZED TO THE "SHORT EDGE 768, SIZE LIMIT 768 X 1344, ALIGNMENT 32" CANVAS, AND OUTPUT RESOLUTION IS TWO SETS OF RULES. THE MATERIAL WAS FIRST DESIGNED TO BE A WIDE-RANGING RATIO。
Q: Console brush dtype warning A: Normal phenomenon (step 5), partial retreat standard attitudinal, without affecting generation。
Q: MY 3090 IS SEVERAL TIMES BEHIND THE OTHERS' 4090 A: Two reasons for supersing - 1,390 is the Ampere architecture, without the FP8 hardware unit, INT8/FP8 power is important to turn back to high accuracy before computing, and the comparison between generations is naturally slow; 2 community members have observed that the dynamic H3 display is conservative, and that the 24GB card may only be used at around 18GB, and that the known underutilization is not your fault. Capable: OpenSageAttentionautumn Modes, INT8 Lean weight for mass-oriented, add memory to 64GB less layer to wait
Q: 16GB card (5070 Ti / 4080) goes off, "Device memory is nearly full" A: THIS IS A DISPLAY OF THE LOADING PHASE + MEMORY DOUBLE-TIGHTENING. SEQUENCED: 1 SYSTEM 32GB IS THE BOTTOM LINE, ADDED TO 64GB MOST EFFECTIVE; 2 PLUS START PARAMETERS –cache-none-disable-smart-memoory Force the model to be fully off-loaded;3 turn off the browser hardware accelerator (the browser itself will account for 1-2GB);4 replace the INT4/NVFP4 small weight at the end of step 2. There are also 5070 Ti cases of unusual fluctuations in the speed of user reporting, which are significantly slower than the same card and are updated to the latest stable version of ComfyUI and graphic card-driven comparison
Q: GENERATED PERSON DOES NOT SPEAK CHINESE ENOUGH / THE LINE HAS BEEN CHANGED A: CHECK LABEL - LANGUAGE TAG SYNTH[ Chinese ]) with the original line in the label and the identity and tone description outside the label. The lines are usually rewritten because they're written outside。
12. Sound effects, alignment ahead of expectations
The official campaign “Personal 32kHz stereo”, which measured decomposition: 32kHz, double-sounding, does have different left and right channels (L/R 0.2–0.9 dB, which is lower than the main signal 14-24 dB) and is not a one-channel copy of two. But it is expected to be right:It's a "space-sensitive" narrow field stereo, not a separation of powerful images, do not count on the right- and right-crossing sound positioning. It's one of the most valuable things: white, sound, music and images are generated in a front-to-face transmission -- lips are born, not lateral。
13. Local vs API, one account
OFFICIAL API PRICES (RMB):2K ~0.80/S, 768P ~0.50/S, 768P ~2K GENERATE ~0.30/S; audio input is free of charge, and pictures 5 are free of charge (over ¥0.20/)。
Note 0.50 + 0.30 = 0.80 - 768p draft first then 2K, and direct 2K One priceThe official pricing did not leave any discount on the cloud draft. This is exactly where the local deployment will be
- An iterative error, running draft local。 A 5-second 768p walk API wants 2.5, local electricity charges are ignored, and a 50-version tip saves more than $100。
- The final draft is clear, two paths。 ROUTE A:local overrated 1080p, freeI don't know. The community already has a ComfyUI workflow that adjusts the parameters, pulling 768p in pieces with LTX 2.3 or Wan 2.2 5B -- Note that sigmas are selected to lower the sigmas version, and that too high a parameter can change the face, damage the mouth; some of the earlier versions are planted on it. Route B:API L 2K, ¥0.30/S15 seconds into a piece of ¥4.5. It is not the same as normal overscores — recreated with the original context, the small words and details are not guessed, the value of the delivery-grade project. The ComfyUI contains a ready-made API template (starting with the three api_in the 3rd step table) and does not have to leave the table。
- N CARDS WITHOUT 12GB+, OR LESS THAN A FEW TIMES A YEAR, DO NOT MAKE 40 GB WEIGHTS FOR SEVERAL VIDEOS。

The local savings were mainly the biggest waste of the trial error phase; at the final stage, the 1080p was enough to save even the high-resolution money, and only the official 2K would pay 0.30/second. Think of this position, a consumer-class card is not just a drafter, it's the main machine that makes a film。
In the end, four categories of people died on the route:
- Just trying to play → Step 0 page is enough, don't go down。
- PRODUCTION CAPACITY, 2K, DELIVERY ~ DIRECT API, ONE 5 SECONDS 2K 4 DOLLARS, DON'T MESS AROUND。
- N CARD WITH 12GB+, LARGE VOLUME, LOVE → Local deployment + Mixing on top, this lesson is written for you。
- I want to fine-tune and study → The full bf16 weight of AdaLN, which is also in 13B, is officially recommended to SGLang Doc, which is beyond the scope of this paper。
ATTACH: IF YOU DECIDE TO LEAVE API, THIS SCRIPT WILL BE COPIED
Registered at platform.minimaxi.com (oversea platform.minimax.io), user centre full, create API key in account management. And then three steps: submit the mission → Rotation status → Download video:
i'm sorry
API_KEY = os.environ[“MINIMAX_API_KEY”] # 你的 key
BASE = “https://api.minimaxi.com” # 海外用 https://api.minimax.io
headers = {“Authorization”: f”Bearer {API_KEY}”}
# 1. SUBMISSION OF TASKS
payload ={
“model”: “MiniMax-H3”,
“content”: [
{“type”: “text”, “text”: “镜头拍摄一只橘猫趴在窗台上,”
“窗外下着雨,猫的尾巴慢慢摆动,画面是暖色调,”
“配上雨声和远处隐约的车流声。”}
],
“duration”: 5, # 4-15 的整数
“resolution”: “768P”, # 草稿用 768P,定稿再 2K
“ratio”: “16:9″, # 纯文生视频必填,不能写 adaptive
}
r = requests.post(f”{BASE}/v2/video_generation”, headers=headers, json=payload)
r.raise_for_status()
task_id = r.json()[“task_id”]
# 2. 轮询直到完成
while True:
time.sleep(10)
q = requests.get(f”{BASE}/v2/query/video_generation/{task_id}”, headers=headers).json()
status = q[“task”][“status”]
if status == “succeeded”:
url = q[“task”][“content”][“url”]
break
if status in (“failed”, “canceled”):
rice SystemExit(q)
# 3. DOWNLOAD (LINKS ARE TIME-LIMITED, DO NOT HOARD)
open (“out.mp4, “wb”).write(requests.get(url).content)
All you need to do is have a video listen Riga {"type": "image_url", "image_url": {"url": "https://your photo address.png"}, "role": "first_frame"}Add a frame to that i'm sorry, ratio Delete (which is automatically proportional to the picture)。
Four new walls. I'll tell you in advance:
- duration Only Integer number from 4 to 15(note and local 17k+5 frame grid are two sets of rules)。
- INDIVIDUAL LIMITATIONS ON MATERIAL: VIDEO 50MB, FIGURE 30MB, AUDIO 15MBREQUEST TOTAL 64MBOver and pass the URL。
- Audio cannot be entered as a separate input, must have a picture or a video。
- _Other Organiser 7 days, the video link is time-barred and downloads。
(b) The measured environment: RTX 4090 48GB (magic)/CommyUI 0.30.0/torch 2.11 + cu128. Time-consuming, visible and reversible data are measured in the text; power weights are based on the Hugging Face official warehouse; the hint is based on the Mini Max official guide。