You have probably generated a thousand still images in ComfyUI. Your first video clip is a different animal: the models are ten times larger, the file layout is unfamiliar, one wrong VAE file leaves you with a silent render, and the VRAM math that worked for 1024x1024 images falls apart the moment motion enters the picture. That gap between "I know ComfyUI" and "I can run a video model locally" is exactly what this SOP closes.
To be precise about scope, this is not our ComfyUI resource roundup, which catalogs the ecosystem, and it is not our consistent-style image workflow, which locks a visual style across stills. This piece is about the thing those articles do not cover: running open-weight video models end to end inside ComfyUI, from the first download to a finished clip with sound. We use LTX-2.5 as the worked example because it currently has the lowest credible hardware bar among models that generate audio and video in one pass. Everything below reflects the state of play as of October 9, 2026.
Step 0: Check Your Hardware and Environment Before Downloading Anything
First, know what LTX-2.5 is. Lightricks released it on August 11, 2026 as the successor to LTX-2.3: a 22-billion-parameter model that generates video and 24kHz stereo audio in the same pass, built around a custom Gemma 4 12B text encoder and a diffusion video decoder. It outputs up to 4K HDR (3840x2176) at up to 50fps. That is the full-size architecture, and it is not what most of us will run: at full precision the model wants 80GB or more of VRAM, and all components together weigh in around 66GiB of downloads.
The ComfyUI path is different. The community runs LTX-2.5 through its distilled and int8-quantized builds, which drop the single-card requirement to 12GB of VRAM. That 12GB figure is the headline number of this entire workflow, and every file you download in Step 1 exists to make it true.
Before you download, verify three environment lines. Setup guides compiled by innfactory list CUDA 12.7 or newer, PyTorch around 2.7, and Python 3.12 or newer. If your stack is older, upgrade the toolkit before touching model files, because version mismatches here produce errors that look like model problems and waste an evening.
The honest triage:
- 12GB VRAM or more: you can run LTX-2.5 distilled/int8 as described in this SOP. This covers most RTX 3060/4070-class cards and anything newer.
- 8GB VRAM: do not force LTX-2.5. Load Wan 2.2's TI2V-5B model with ComfyUI's native offload instead. It is a smaller 5B model with the most permissive license in its class (Apache 2.0), and it will actually complete a render on your card.
- 80GB-class hardware (A100/H100 territory): you can chase the full-precision build and the 4K pipeline, but that is outside this SOP's scope.
One more disk-space note: even the ComfyUI-oriented file set in Step 1 adds up to tens of gigabytes. Clear the space now rather than discovering a full drive at 90 percent of a download.
Step 1: Download Four File Types and Put Them in Exactly the Right Folders
The weights live at huggingface.co/Lightricks/LTX-2.5 on Hugging Face (plain form: huggingface.co/Lightricks/LTX-2.5). The repo is gated, but in the friendliest way: it is set to gated=auto, which means you log in, accept the community license terms once, and access is granted automatically with no manual review. Be aware that even viewing the model card requires that login first, so do not mistake a login wall for a broken page.
You also want the ComfyUI custom node pack ComfyUI-LTXVideo from Lightricks (plain form: github.com/Lightricks/ComfyUI-LTXVideo). It supplies the LTX-specific nodes the workflows in Step 2 are built on. Install it through ComfyUI Manager or clone it into your custom_nodes folder.
Now the critical part. ComfyUI expects four distinct file types in four distinct directories, and mixing them up is the most common install failure. For the 12GB setup, download the int8_convrot builds, not the BF16 builds, for the transformer and text encoder:
- The diffusion transformer goes to
models/diffusion_models/. The exact file isltx-2.5-22b-distilled-transformer-comfy-int8-convrot.safetensors. This is the 22B model itself, distilled and quantized to int8. - The text encoder goes to
models/text_encoders/. The exact file isgemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot. This is the custom Gemma 4 12B encoder with the projection layer that maps its output into the diffusion model's input space. - The VAEs go to
models/vae/, and there are two of them:ltx-2.5-video-vae-bf16for the video latent space andltx-2.5-audio-vae-bf16for the audio latent space. These ship in BF16 because VAEs are small; the quantization savings are in the big models, not here. - The latent upscaler goes to
models/latent_upscale_models/. The exact file isltx-2.5-latent-spatial-upscaler-x2-bf16. This is the piece that makes the two-stage workflow in Step 2 possible.
Note the fourth category: latent_upscale_models is not a folder you may have used before with image models, and it is easy to create a wrong directory name out of habit. Copy the path exactly.
While you are on the model page, read the license terms you are agreeing to. LTX-2.5 ships under the LTX-2.x Community License, and Lightricks is explicit that it is not Apache 2.0. The short version: commercial use is free if your annual revenue is under 10 million US dollars (inclusive); at or above that threshold you need a Commercial Use Agreement. For individual creators and most studios, that means free, but it is a real contract you accept when you click through the gate.
Step 2: Pick One of Three Workflow Shapes
With files in place, the custom node pack gives you three workflow archetypes. Choose by intent, not by whichever template loads first.
Shape one: two-stage T2V or I2V, generate then upscale. This is the quality path. Stage one is a standard text-to-video or image-to-video pass at your base resolution, producing a latent clip. Stage two feeds that latent through the 2x spatial upscaler you placed in Step 1 and runs a refinement pass, roughly doubling effective resolution without re-rendering from scratch. The same skeleton handles both T2V and I2V; the only real difference is whether an image latent enters the first stage. Treat this as your default for anything you plan to keep.
Shape two: single-stage, fast preview. Same generation stage, no upscaler, no refinement. It is the fastest route from prompt to pixels and the right choice for testing prompt phrasing, checking composition, or iterating on camera motion. The output is visibly softer than the two-stage result, which is precisely why it is cheap: do not judge final quality on a single-stage preview, and do not ship one.
Shape three: A2V, audio-driven video. Here an audio track drives the generation, producing motion that follows sound. This is the workflow only audio-native models can support at all, and it is LTX-2.5's structural advantage: audio and video are generated in the same pass at 24kHz stereo, so there is no post-hoc dubbing step to synchronize.
How long can a single clip be? Two third-party numbers coexist and neither should be quoted as a hard spec. Pinggy's writeup puts the single-pass maximum at about 20 seconds, extendable through an Extend pipeline; innfactory's material lists 121 frames at 24fps, which is roughly 5 seconds. The discrepancy likely reflects different builds and settings, so treat duration as something to test on your own card rather than a number to promise a client.
Three Pitfalls That Quietly Ruin Renders
Pitfall one: the audio VAE is a separate file, and skipping it gives you a silent video. Because the video VAE and audio VAE are two different safetensors files, a partially completed download folder will still load, still render, and still produce a perfectly good-looking clip, just with no sound and no error message pointing at the cause. If your first render is silent, check that ltx-2.5-audio-vae-bf16 actually landed in models/vae/ before you rewire anything.
Pitfall two: the prompt enhancer is a hidden cost. The workflow supports an optional prompt enhancer model that rewrites your prompt into a richer description before generation. It works, but it is roughly a 5GB additional download and it adds one to two minutes to every single generation. It is off by default. Turn it on deliberately when you want prompt expansion, and leave it off when you are iterating, or your preview loop gets twice as slow for no reason.
Pitfall three: the instinct to render big in one pass. The reliable VRAM strategy on consumer cards is the opposite of what feels natural: generate a short, low-resolution clip, then let the dedicated upscaler carry the resolution. One giant render at high resolution and long duration is how 12GB turns into an out-of-memory crash at the last sampling step. Combine this with the defaults already built into your setup: run quantized/distilled builds first, and let ComfyUI's native offload move weights between VRAM and system RAM as needed. Short-plus-upscale reliably beats one-big-render on both success rate and total time.
Which Model Should You Actually Load?
LTX-2.5 is the worked example here, but the whole point of running models locally is that you can choose. For the full architecture story, including the diffusion video decoder that separates LTX-2.5 from its LTX-2.3 predecessor, see our LTX-2.5 open-source deep dive. For model selection, our open-video-model local comparison works through the four leading open-weight video models across five axes: license, VRAM, audio, duration, and regional availability. The compressed version for ComfyUI users:
- Want the most permissive license? Wan 2.2. Apache 2.0, runs on 8GB with offload via TI2V-5B, outputs 480p-720p short clips. No native audio.
- Want audio and video in one pass? LTX-2.5. 24kHz stereo in the same generation pass, 12GB distilled/int8 entry point, Community License with the 10-million-dollar revenue threshold.
- Have a 24GB card and no EU/UK/South Korea exposure? HunyuanVideo 1.5. Starts around 14GB VRAM, but Tencent's community license does not apply in those regions and very large user bases need separate terms.
- Need dialogue or reference audio? MiniMax H3, in the regions its license allows. Note its community license excludes the EU, UK, South Korea, and the US, and revenue above 20 million dollars needs written authorization.
For this SOP's audience, the decision is simpler: if your card has 12GB, load LTX-2.5 and follow the steps above; if it has 8GB, load Wan 2.2 TI2V-5B with offload and skip Step 1's int8 files entirely.
FAQ
Q1: What is the minimum hardware for ComfyUI video generation with LTX-2.5? 12GB of VRAM with the distilled/int8 builds is the working floor, alongside CUDA 12.7+, PyTorch around 2.7, and Python 3.12+. The full-precision model needs 80GB or more and is not the realistic target. On 8GB cards, run Wan 2.2 TI2V-5B with offload instead.
Q2: The Hugging Face page will not show me anything. Is the repo down?
No, the repo is gated with gated=auto. Log in to your Hugging Face account, accept the license terms once, and access is granted automatically. Even the model card is hidden until you log in, so a login wall is the expected behavior, not an outage.
Q3: My clip rendered fine but has no audio. What broke?
Almost certainly the audio VAE is missing. LTX-2.5 uses two separate VAE files in models/vae/, and the render completes silently without the audio one. Verify ltx-2.5-audio-vae-bf16 downloaded completely, then re-run the workflow.
Q4: Can I use LTX-2.5 output commercially? Yes, under the LTX-2.x Community License if your annual revenue is under 10 million US dollars (inclusive). At or above that threshold you need a Commercial Use Agreement from Lightricks. The license is custom, not Apache 2.0, so read the actual terms you accept when downloading.
Q5: Should I render long clips at high resolution in one pass? No. On consumer cards the reliable strategy is a short, low-resolution generation followed by the dedicated 2x latent upscaler. Single-pass big renders are the main cause of out-of-memory crashes near the end of sampling. Use single-stage output for previews only and the two-stage path for anything you keep.
References
- ComfyUI-LTXVideo GitHub Repository (plain: github.com/Lightricks/ComfyUI-LTXVideo)
- LTX-2.5 Weights on Hugging Face (plain: huggingface.co/Lightricks/LTX-2.5)
- Lightricks LTX-2 GitHub Repository