{"id":55807,"date":"2026-08-23T09:15:49","date_gmt":"2026-08-23T01:15:49","guid":{"rendered":"https:\/\/www.1ai.net\/?p=55807"},"modified":"2026-08-11T13:24:21","modified_gmt":"2026-08-11T05:24:21","slug":"comfyui-%e4%bb%8e%e5%ae%89%e8%a3%85%e5%85%a5%e9%97%a8%e5%88%b0%e7%b2%be%e9%80%9a%ef%bc%8cltx-2-%e5%9c%a8-4060-8g-%e4%b8%8a%e8%b7%91-ai%e8%a7%86%e9%a2%91%ef%bc%88%e4%bd%8e%e6%98%be%e5%ad%98%e4%b9%9f","status":"publish","type":"post","link":"https:\/\/www.1ai.net\/en\/55807.html","title":{"rendered":"ComfyUI from installation entry to mastery, LTX-2 runs AI video on 4060 8G (low-visibility can also produce video)"},"content":{"rendered":"<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-55827\" title=\"3be49769jlazi00nvd000p000dup\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/08\/3be49769j00tjlazi00nvd000p000dup.jpg\" alt=\"3be49769jlazi00nvd000p000dup\" width=\"900\" height=\"498\" \/><\/p>\n<p>THE FIRST 13 ARTICLES COME DOWN, AND YOU'VE GOT A WHOLE SET OF STATIC MAPS \u2014 MACHINES, DRAWINGS, MODULATIONS, MAGNIFICATIONS, LOCKS. BUT THERE'S ONE THING I'VE BEEN HOLDING ON TO, BECAUSE IT'S THE EASIEST THING TO TALK ABOUT<strong>video<\/strong>.<\/p>\n<p>I gotta be honest, the first time I tried to touch the video, the mind was, \"Oh, forget it, 8G.\" Queue out of memory. And then I knew, not eight Gs, I couldn't play, I used the wrong model, I didn't quantify it. Now, LTX-2, this open-source video model, quantified by distilled version + GGUF, and we actually got 4060, 8G, and that's a low-profile set. This article speaks for itself\u3002<\/p>\n<p><strong>ONE, LTX-2. WHAT'S THE DIFFERENCE<\/strong><\/p>\n<p>One sentence:<strong>LTX-2 IS AN OPEN-SOURCE \"VANCTUARY\/TRUE VIDEO\" MODEL, YOU WRITE A SENTENCE (OR THROW A PICTURE) AND IT SPITS OUT A MOVING SHORT FILM\u3002<\/strong>The figure is \"one frame,\" and the video is \"a series of frames to move.\" Video is much heavier than the figure - a 49 frame video is the equivalent of repeating a similar process 49 times with a frame, so it is especially visible\u3002<\/p>\n<p><strong>Remember:<\/strong>Video = A lot of frame + frame and frame consistency, so it eats more visible than the picture, but the idea is the model + hint + sample\u3002<\/p>\n<p>Two main games:<br \/>\n\u00b7\u00a0<strong>VINCENT VIDEO (T2V)<\/strong>: Write a hint, generate a video in a vacuum\u3002<br \/>\n\u00b7\u00a0<strong>TUSHENG VIDEO (I2V)<\/strong>: Throw a static map as the front frame and let the things in the picture move\u3002<\/p>\n<p><strong>II. Models and plugins where to put them (the wrong drop-down frame is always empty)<\/strong><\/p>\n<p>LET'S START WITH THE STUFF, THE PIT IS THE MOST. LTX-2 IN <a href=\"https:\/\/www.1ai.net\/en\/tag\/comfyui\" title=\"_Other Organiser\" target=\"_blank\" >ComfyUI<\/a> This is not a self-contained node, and you want to load a custom node package + a quantitative model\u3002<\/p>\n<p><strong>Step 1: Load custom nodes<\/strong>(All found in ComfyUI Manager)<br \/>\n\u2022 CommyUI-LTXVideo: core node package, without which there would be no LTX node\u3002<br \/>\n\u2022 ComfyUI-GGUF: used to load .gguf quantitative models\u3002<br \/>\n\u2022 ComfyUI-KJNodes: Video VAE support node\u3002<br \/>\n\u2022 CommyUI-VideoHelperOffice: Video reading and writing, frame tool\u3002<\/p>\n<p><strong>Trail experience:<\/strong>For the first time, I only loaded LTXVideo with a red-screen \u201cmissing\u201d. It's actually GGUF and VideoHelperSite that aren't installed, and the LTX nodes depend on them. Reset, red nodes all\u3002<\/p>\n<p><strong>Step two: the following model, laid down by a fixed directory<\/strong><\/p>\n<p>Q3_K_M.gguf<br \/>\nComfyUI\/models\/text_encoders\/gemma-3-12b-it-qat-Q2_K.gguf<br \/>\nitx-2-19b-dev_embedings_conectors.safetensors<br \/>\nCommyUI\/models\/vae\/ \u2192LTX2_video_vae_bf16.safetensors (or ltx-2-19b-dev_video_vae.safetensors)<\/p>\n<p>The model goes under the LTX-2 section of Hugging Face\/Civitai, free of charge\u3002<\/p>\n<p><strong>Trail experience:<\/strong>It's not gonna come out of the bottom of the model -- it's old. It's the same pit as the Lora, the magnification model. I'm going to do it again\u3002<\/p>\n<p><strong>III. Hand-to-hand node (principle, how to do it)<\/strong><\/p>\n<p>LTX-2 HAS A NODAL CHAIN LIKE THIS<\/p>\n<p>Load LTX Model \u2192CFGGuider (with positive and negative) \u2192guider<br \/>\nLTXVScheduler sigmas samlerSelec<br \/>\nImptyLTXVLentVideo + LTXVEmptyLatentAudio<br \/>\nSamplerCustomAdvanced (noise+guider+sampler+sigmas+av_latent)<br \/>\nLTXVSeparateAVLatent<br \/>\nVHS_VideoCommine (Fit + Audio)<br \/>\nDualCLIPLoader (GUF, type=ltxv) CLAP Text Encode(\u00b1) \u2192 Popular \/ Negative<\/p>\n<p>How do you get the port<br \/>\n\u2022 Yellow of the DualCLIPLoader\u00a0<strong>CLIP PORT<\/strong>\u00a0crip mouth of the positive and negative node\u3002<br \/>\n\u00b7\u00a0<strong>critical pit: type must be manually set to ltxv<\/strong>i don't know. default may be sd3, unmodified video full of static snowflakes, and no moving frame. this is the first time i've been here, and it's snowflakes and snowflakes\u3002<br \/>\n\u00b7\u00a0<strong>Model does not go straight into the sampler<\/strong>: MODEL mouth of Load LTX Model \u2192 model mouth of CFGGuider; CFGGuider eat again Positive \/ Negative, output guider \u2192 to SamplerCustomAdvanced\u00a0<strong>guider mouth<\/strong>(It doesn't even have a model mouth, it's the biggest difference with the old KSampler\u3002<br \/>\n\u2022 LTXV Scheduler\u00a0<strong>SIGMAS<\/strong>\u00a0SamplerCustomAdvanced's sigmas mouth\u3002<br \/>\n\u2022 SamplerCustomAdvanced also eats noise (RandomNoise) and latent_image (av_latent above); noise + guider + sampler + sigmas + latent_image + five mouths to feed latent\u3002<br \/>\n\u2022 The sampler output is\u00a0<strong>AV Latet (video + audio pack)<\/strong>, the LTXVSeparateAVLatent must be removed and sent to VAEDecode; the decoding of VAE by feed to AVLent will be incorrect\u3002<br \/>\n\u2022 Video decoded by VAE + audio \u2192 VHS_Videocompine (export mp4)\u3002<\/p>\n<p>An additional step for the Tusheng Video (I2V): Add a Load Image LTXV Img To Video Inplace (inline node, category LTXVImgToVideoInplace). It encodes the picture into the latent frame (vae+image+latent \u2192 latent), which is attached to the latent chain before the sample (and the latent that came out of EmtyLTXVLentVideo, which joins the LTXVConcatAVLatent), and the first frame of the video locks you in. Note that it's going to be \"latent,\" not \"model.\" Don't use it as a model node\u3002<\/p>\n<p>One sentence:<strong>Text takes the GGUF double CLIP (type set ltxv) \u2192 model + hints into the CFGGuider change guider \u2192 scheduler to sigmas \u2192 sampler to remove the VV field \u2192 VideoCombin saver\u3002<\/strong><\/p>\n<p><strong>A 2-second dog video<\/strong><\/p>\n<p>LET'S START WITH THIS ONE. AFTER A LITTLE BIT OF MOVING PUPPY IN YOUR HAND, YOU CAN CONFIRM FOR YOURSELF THAT THE 8G CARD DOES RUN THE VIDEO\u3002<\/p>\n<p><strong>Preconditions<\/strong>: You have installed ComfyUI + four node bags above, and the model is placed in the corresponding directory by section II\u3002<br \/>\n<strong>Target<\/strong>: A 512 x 288 video of a dog running in the garden for about 2 seconds (49 frames) in distillation\u3002<\/p>\n<p>Step (sight, don't jump):<br \/>\n1.\u00a0<strong>Down Model<\/strong>:unet with ltx-2-19b-distilled_Q3_K_M.gguf, text encoders with text_encoders\/(gemma qat Q2 + connector), vae with LTX2_video_vae_bf16 and updated with section II, F5\u3002<br \/>\n2.\u00a0<strong>Plugin<\/strong>: Manager search ComfyUI-LTXVideo \/ GGUF \/ KJNodes \/ VidioHelperSuite, reset\u3002<br \/>\n3.\u00a0<strong>Load pattern workflow<\/strong>: node belt example_workflows\/video_ltx2_t2v_distilled.json (8 steps)\u3002<br \/>\n4.\u00a0<strong>DOUBLE CLIP<\/strong>: Plus DualCLIPLoader (GGF), crip_name1 for gema Q2, clip_name2 for connector, type manually to ltxv\u3002<br \/>\n5.\u00a0<strong>Size frames<\/strong>WIDTH 512, HIGH 288, FRAME NUMBER 49 (FORMULA (8XN)+1), FRAME RATE 24\u3002<br \/>\n6.\u00a0<strong>Sample<\/strong>: Steps fixes 1.0 with distillation of 8, CFG (retorted fixed value, full model 5~6)\u3002<br \/>\n7.\u00a0<strong>Write hints<\/strong>\"A gold retrever puppy playing in a sunny Garden, waking its hands.<br \/>\n8.\u00a0<strong>publish a video<\/strong>Queue Prompt, wait a few minutes, VHS_Videocompine spits out mp4\u3002<\/p>\n<p><strong>Checkpoint<\/strong>1 The mp4, the dog moves in the garden for about 2 seconds; 2 = type does not set ltxv, step back 4; 3 Full Black \/ Green = model file name or dtype is incorrect, check the catalogue; 4 Direct OOM = resolution or frame is too large, down to 41 frames 480 x 270\u3002<\/p>\n<p>Work (selection): Same hint, Steps 8 (distillation), Steps 25 (full model), contrasting the difference between \"quick but rough\" and \"slower but better\"\u3002<\/p>\n<p><strong>IV. Core knob\/ parameter details<\/strong><\/p>\n<table>\n<tbody>\n<tr>\n<td>\n<section>Crunch<\/section>\n<\/td>\n<td>\n<section>Sweet \/ Default<\/section>\n<\/td>\n<td>\n<section>clarification<\/section>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<section>Frames<\/section>\n<\/td>\n<td>\n<section>(8XN)+1, COMMON 49-881<\/section>\n<\/td>\n<td>\n<section>49 \u5e27\u22482 seconds@24fps; cap 193; 4060 with 49~61<\/section>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<section>Resolution<\/section>\n<\/td>\n<td>\n<section>Draft 512 x 288,4060 Upper limit 640 x 384<\/section>\n<\/td>\n<td>\n<section>Shortest edge 512<\/section>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<section>Steps<\/section>\n<\/td>\n<td>\n<section>Distillation fixed 8; full model 18-34<\/section>\n<\/td>\n<td>\n<section>Less than 18, consistent motion<\/section>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<section>CFG<\/section>\n<\/td>\n<td>\n<section>distillation fixed 1.0; full model 4.5 ~ 6.5<\/section>\n<\/td>\n<td>\n<section>DISTILLED VERSION CFG = 1.0; FULL MODEL TOO HIGH EDGE FLASH<\/section>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<section>frame rate<\/section>\n<\/td>\n<td>\n<section>24fps<\/section>\n<\/td>\n<td>\n<section>VHS_VideoCompine<\/section>\n<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><strong>Trail experience:<\/strong>I ONCE PULLED THE CFG TO 8 WITH A FULL-SCALE MODEL, AND I THOUGHT, \"BE A LITTLE MORE GOOD,\" AND IT TURNED OUT THAT THE EDGE OF THE VIDEO WAS A LOOP OF LUMINOUS, TEXTURE FLASH. DOWN TO 5.5, IMMEDIATELY NORMAL. AND THE DISTILLED VERSION, ON THE OTHER HAND, HAS A 1.0-- TWO-LINE CFG\u3002<\/p>\n<p><strong>V. 4060 \/ 8G SPECIAL PURPOSE: HONESTLY, IT'S \"HARD TO RUN\"<\/strong><\/p>\n<p>Let's start with the cold water<strong>4060's 8G running LTX-2 is \"hard to run,\" not silk 1080p\u3002<\/strong>MULTISOURCE MEASUREMENTS ARE CONSISTENT -- 8G CARDS HAVE TO BE QUANTIFIED + LOW RESOLUTION + LOW FRAME, WITH A FEW SECONDS OF LOW-CLEAN SHORT CLIPS. I CHECKED THIS OUT ON PURPOSE. I'M NOT FOOLING YOU\u3002<\/p>\n<p>\u00b7\u00a0<strong>Quantified<\/strong>:unet with Q3_K_M (GGF file approximately 10.3G, with low-visibility mode barely plugged into 8G), or with a FP8 distilled version; Q4_0 with approximately 10G, visible 11-12G, 8G card is in suspense, and it is recommended that a left-over selection of Q3_K_M be left. The text encoder maintains Q2_K (approximately 4G, authentication qat warehouse)\u3002<br \/>\n\u00b7\u00a0<strong>It's locked<\/strong>: 512 x 288, starting with a maximum of 640 x 384; minimum margin exceeding 512\u3002<br \/>\n\u00b7\u00a0<strong>Frame restraint<\/strong>: 49~61 frames (approximately 2-2.5 seconds), zoom in two stages if you want to\u3002<br \/>\n\u00b7\u00a0<strong>Display Memory Optimization<\/strong>: PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True\u3002<br \/>\n\u00b7\u00a0<strong>Patience<\/strong>: 512 x 288\/ 49 frame round approximately a few minutes\u3002<\/p>\n<p>I stepped through the pit and said it all at once: 1 big model at first turn, which mistook 8G and the video as unconnected; 2 type, which did not set ltxv, double-heavy, full of snow; 3 CFCG to the edges of 8; 4 frame numbers with 243 direct collapse; 5 without a visual optimisation variable, run the OOM at large\u3002<\/p>\n<p><strong>Unrunable retreat:<\/strong>IF THE LIMIT IS SET FOR YOUR 8G OR FALL, THREE WAYS TO GO -- ONE BY THE CLOUD GPU (HOURLY FOR A FEW MINUTES); TWO FOR THE SMALLER VIDEO MODEL; AND THREE FOR THE OLD VERSION OF LTX (EARLY AND SMALLER, LOW STORAGE THRESHOLD). DON'T PUSH, 8G HAS 8G GAME\u3002<\/p>\n<p><strong>VI. QUIREMENTS OF PARLIAMENTS<\/strong><\/p>\n<table>\n<tbody>\n<tr>\n<td>\n<section>take<\/section>\n<\/td>\n<td>\n<section>Model File<\/section>\n<\/td>\n<td>\n<section>Resolution<\/section>\n<\/td>\n<td>\n<section>Frames<\/section>\n<\/td>\n<td>\n<section>Steps<\/section>\n<\/td>\n<td>\n<section>CFG<\/section>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<section>4060 Rapid water test<\/section>\n<\/td>\n<td>\n<section>distilled Q3<\/section>\n<\/td>\n<td>\n<section>512 x 288<\/section>\n<\/td>\n<td>\n<section>49<\/section>\n<\/td>\n<td>\n<section>8<\/section>\n<\/td>\n<td>\n<section>1.0<\/section>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<section>4060. It's good<\/section>\n<\/td>\n<td>\n<section>distilled Q3<\/section>\n<\/td>\n<td>\n<section>640 x 384<\/section>\n<\/td>\n<td>\n<section>61<\/section>\n<\/td>\n<td>\n<section>8<\/section>\n<\/td>\n<td>\n<section>1.0<\/section>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<section>12G CARD GENERAL<\/section>\n<\/td>\n<td>\n<section>Q4_K_M<\/section>\n<\/td>\n<td>\n<section>768x432<\/section>\n<\/td>\n<td>\n<section>81<\/section>\n<\/td>\n<td>\n<section>24-28<\/section>\n<\/td>\n<td>\n<section>5.5<\/section>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<section>16G+ FULL MODEL<\/section>\n<\/td>\n<td>\n<section>Q5_K_M \/FP16<\/section>\n<\/td>\n<td>\n<section>1024 x 576<\/section>\n<\/td>\n<td>\n<section>97<\/section>\n<\/td>\n<td>\n<section>28-34<\/section>\n<\/td>\n<td>\n<section>5.5<\/section>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<section>TUSHENG VIDEO I2V<\/section>\n<\/td>\n<td>\n<section>distilled Q3<\/section>\n<\/td>\n<td>\n<section>512 x 288<\/section>\n<\/td>\n<td>\n<section>49<\/section>\n<\/td>\n<td>\n<section>8<\/section>\n<\/td>\n<td>\n<section>1.0<\/section>\n<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><strong>Common pits \/ mined areas<\/strong><\/p>\n<p>\u00b7\u00a0<strong>Still snowflakes<\/strong>: DualCLIPLoader type unchanged ltxv, back to step 4 in section III\u3002<br \/>\n\u00b7\u00a0<strong>Full Black \/ Green Screen<\/strong>: model file name or dtype error, check three sets of unet \/ text_encoders \/ vae (especially for gema to use qat repository)\u3002<br \/>\n\u00b7\u00a0<strong>OOM EXPLOSION<\/strong>: Too much resolution or frame, 4060 lock 512 x 288\/ 49 frames first\u3002<br \/>\n\u00b7\u00a0<strong>It's a very rough video<\/strong>:Steps is too low (detillation 8 is enough, the whole model is not below 18)\u3002<br \/>\n\u00b7\u00a0<strong>The hint was ignored<\/strong>: CFG TOO LOW (&lt;3), PULL BACK AROUND 5.5\u3002<br \/>\n\u00b7\u00a0<strong>Frame number set to 243<\/strong>: DIRECT COLLAPSE ABOVE THE CEILING, STRICT (8XN)+1\u3002<\/p>","protected":false},"excerpt":{"rendered":"<p>The first 13 articles come down, and you've got a whole set of static maps \u2014 machines, drawings, modulations, magnifications, locks. But there's one thing I've been keeping quiet about, because it's the easyest way to talk to 8G users: video. I gotta be honest, the first time I tried to touch the video, the mind was, \"Oh, forget it, 8G.\" Queue out of memory. And then I knew, not eight Gs, I couldn't play, I used the wrong model, I didn't quantify it. Now, LTX-2, this open-source video model, quantified by distilled version + GGUF, and we actually got 4060, 8G, and that's a low-profile set. This one says it<\/p>","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[149,144],"tags":[1989,6276,4749],"collection":[],"class_list":["post-55807","post","type-post","status-publish","format-standard","hentry","category-jiaocheng","category-baike","tag-comfyui"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts\/55807","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/comments?post=55807"}],"version-history":[{"count":0,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts\/55807\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/media?parent=55807"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/categories?post=55807"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/tags?post=55807"},{"taxonomy":"collection","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/collection?post=55807"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}