While cloud video models like Kling 4.0 and Seedance 2.5 race each other on specs and pricing, a second route has quietly laid its cards on the table: Lightricks has released the full weights of LTX-2 on GitHub, a 22B-parameter model that generates picture and sound together. Per the official README, this is the first DiT-based audio-video foundation model, with synchronized audio and video, multiple performance modes, production-ready outputs, and API access alongside open weights. As verified on 2026-10-02, the Lightricks/LTX-2 repo holds 9,569 stars, with the latest push also dated October 2 -- the project shows no sign of cooling off.
This article covers what is actually inside the model, how far its inference efficiency goes, and what the twelve pipelines each handle. The most important verification, however, concerns the license: the GitHub License field reads NOASSERTION, and behind it sits a custom community agreement that is not the same thing as Apache or MIT. Read it before any commercial use. For the cloud side of the story, see this site's coverage of Kling 4.0's October launch and the Seedance 2.5 release.
One model for everything: what sits inside 66GiB of weights
Start with positioning. Most mainstream video models generate picture only; voice and sound effects are stitched in afterwards by a separate toolchain. LTX-2's pitch is audio-video unity: one prompt goes in, video and matching audio come out together, with lip movements and sound naturally aligned to the picture. That removes the post-production sync pass at the architecture level, and it is what backs up the "foundation model" claim.
The recommended version is LTX-2.5, with weights published per component, so you only download the parts your pipeline needs. The core is a 22B distilled transformer shipped as a bf16 safetensors file; a full dev transformer also exists for the guided two-stage pipelines. The text encoder is a custom Gemma 4 12B tuned for LTX -- note that it is not interchangeable with Google's stock Gemma 4, since loading checks the encoder version against the one the checkpoint was trained with, and substituting the original weights fails outright. Beyond that: one video VAE and one audio VAE, where the video VAE comes in a diffusion-decoder build (higher quality, more VRAM) and a convolutional build (lighter, no extra dependencies), plus spatial and temporal upscalers. The full set is roughly 66 GiB.
A few engineering details are worth knowing in advance. The diffusion decoder of the video VAE runs fastest with the natten extra installed, but that accelerated backend is Linux and CUDA only -- on Windows and macOS decoding falls back to a Triton or eager implementation. The weights are gated on Hugging Face: if you hit a 401 or 403, accept the model terms on the model page first, then log in with a read token. There is also an optional duration head that lets you omit the frame count and have the clip length predicted from the prompt. The legacy LTX-2.3 checkpoints are still maintained, but they bundle transformer, VAEs, and text projection into single files with a separately downloaded Gemma 3 encoder; files are not interchangeable between the two generations, and a LoRA only works with the model it was trained on.
What does 66GiB mean in practice? The download alone is the first barrier, and this is not something a phone or a thin laptop will run. For the target users, though, the size buys complete self-deployment: the weights live on your machine and generation never touches a third-party server.
Eight-step inference and twelve pipelines
The key efficiency move is distillation. DistilledPipeline runs on only 8 predefined sigmas -- stage 1 takes eight steps, stage 2 takes four -- and is the fastest starting route. Default output is 1024x1536 at 24 fps. For 4K, the flags are width 3840 and height 2176; the README explicitly notes it is not 2160, so do not fill in the habitual number when scripting.
If VRAM runs tight, there are fallbacks: FP8 quantization via fp8-cast on the CLI, or fp8-scaled-mm on Hopper-class GPUs with native FP8 support, plus CPU or disk offload. Attention backends matter too: datacenter Blackwell B200 cards need a manual FlashAttention 4 install, and the vendor names the specific revision verified against torch 2.9.1+cu128; Hopper cards take the FlashAttention 3 wheel, and everything else uses PyTorch's built-in SDPA automatically. One more official trick, gradient estimation, cuts the standard pipelines from forty denoising steps to twenty or thirty with little quality loss. Consumer cards can still play; you just trade speed against quality.
There are twelve pipelines covering most of a video production workflow. The highlights:
- DFR (Diffusion Fidelity Rendering) is the production-quality text/image-to-video path. It uses the same distilled transformer plus a detailing IC-LoRA: extra generated keyframes, then a spatial detailing pass. The cost is longer runtime and more VRAM.
- DubIt handles rephrased dubbing: change the line while preserving the speaker's voice identity and lip movements. For teams producing multilingual content, this is a genuine necessity.
- Retake regenerates a specific time region of an existing video, so a flawed segment does not force a full re-render.
- HDR IC-LoRA converts SDR video to HDR, writing a BT.2020/HLG master plus an ACEScct EXR sequence that plugs straight into professional grading; native EXR and HDR input and output are standard capabilities.
- A2Vid runs the reverse direction, generating video conditioned on an input audio file. The rest include IC-LoRA video-to-video transforms, keyframe interpolation, and three TI2Vid variants (two-stage, HQ, and one-stage).
The training side is not left empty either: the ltx-trainer package supports LoRA, full fine-tuning, and IC-LoRA, meaning you can tune the model on your own footage into a proprietary style -- a freedom that pay-per-generation cloud models do not offer. Workflow users can go through the official ComfyUI plugin, ComfyUI-LTXVideo. On prompting, the official guide advises keeping prompts under 200 words, describing actions chronologically with concrete camera language. The method transfers to other video models as well; this site's AI video generation SOP generalizes it.
License verification: NOASSERTION is not decoration
Open the GitHub repo and the License field reads NOASSERTION rather than the familiar Apache or MIT. That is not a hosting-platform glitch: the repo genuinely uses no OSI-recognized open-source license. LTX-2 ships under Lightricks' own "LTX Community License," applicable to LTX-2.5 versions released since August 11, 2026. Downloading counts as acceptance. The key clauses, each traceable to the license text:
First, the revenue threshold. Any entity with annual revenues of at least 10 million USD (aggregated across subsidiaries, affiliates, and companies under common control) must purchase a Commercial Use Agreement from Lightricks before any commercial use of LTX-2 or its derivatives. Violations are deemed a material breach, with fees owed for the period of use, payable within thirty days of written demand at the licensor's standard rates.
Second, what stays free. Individual non-commercial use -- research, learning, hobby projects, personal entertainment, with no direct or indirect connection to any commercial activity -- is free. Commercial entities are also exempt for testing, evaluation, and non-commercial R&D in non-production environments. Students, hobbyists, and teams validating the technology can run it freely; the moment it touches revenue, clause one applies.
Third, the definition of derivatives explicitly includes distillation. Any model that acquires LTX-2's capabilities by transferring weights, parameters, activations, or outputs -- including models distilled with intermediate representations or synthetic data -- is a derivative and bound by the same agreement. Want to use LTX-2 as the teacher for your own commercial model? Buy the agreement first.
Fourth, distribution obligations. Derivatives (fine-tuned weights and LoRA adapters included) must be distributed under the same terms with a complete copy of the agreement attached; before transferring a derivative to a commercial entity, that entity must first obtain its own paid license, and the transferor must notify it in writing. The training data is not licensed. Attachment A lists twenty use restrictions, including: no training competing models for commercial use, no use in products that directly compete with Lightricks' offerings, no removal of watermarking, provenance, or latent-disclosure features, and a requirement to label content as machine-generated. The licensor claims no rights in your outputs; responsibility is yours.
A few clauses are easy to skim past. The licensor reserves the right to remotely restrict usage that violates the agreement, and expects users to make reasonable efforts to stay on the latest version -- running an outdated release is at your own risk. On termination for breach, you must stop using the model, delete all copies in your possession, and notify downstream recipients. Disputes are governed by New York law and settled under ICC arbitration rules.
One telling detail: the license states that for purposes of the EU AI Act, the licensor intends LTX-2 to be treated as a free and open-source general-purpose AI model under Article 53(2). The vendor is itself walking a tightrope between "open" and "commercial." For contrast, look at the reference points: Wan2.2 in the same space is Apache-2.0, and the previous-generation LTX-Video is also Apache-2.0. Moving from the previous generation's fully permissive license to a revenue-threshold model says something on its own -- vendors are no longer willing to give away 22B-class video weights for free.
The appeal and the price of self-deployment
Pull back to model selection. Today's flagship video models -- Kling 4.0 (launching in October), Seedance 2.5 (already live), and Alibaba's Wan3.0 (officially launched August 24, see the Wan3.0 launch article) -- are all cloud services: open a web page, type a prompt, burn compute in the vendor's datacenter. LTX-2 is the only option in this tier that lets you pull the weights onto your own machine.
The appeal is concrete: data stays inside your network, a hard requirement for compliance-sensitive teams; there is no per-generation fee, so marginal cost approaches the electricity bill at batch scale; and the model is fine-tunable, meaning a LoRA-trained style is an asset you exclusively own. The benefits land differently by scenario: multilingual content teams save reshoots with DubIt's lip-synced line changes; archive-restoration teams get grading-ready EXR sequences straight from the SDR-to-HDR pipeline; high-volume short-video teams run the distilled pipeline locally with no quota anxiety. The barriers are equally concrete: you need a capable GPU, you swallow a 66GiB download first, production quality means running DFR with a higher VRAM bill, and commercial use hits the revenue threshold -- companies above 10 million USD cannot avoid the paid agreement.
One-line conclusion: for personal learning, hobby creation, and technical validation, LTX-2 is currently free to run without limits. For startup teams below the revenue line, it is the only audio-video unified model on the market with weights in hand. Above the line, buy the Commercial Use Agreement or stay on cloud flagships. An official API and web playground also let you test the output before committing to self-hosting. For how the four models shaping the current landscape compare, see this site's open-vs-cloud video model comparison.
FAQ
Q1: Is LTX-2 genuinely open source?
A1: Strictly speaking, no. The GitHub License field reads NOASSERTION, and the actual terms are the custom LTX Community License, not an OSI-recognized license like Apache or MIT. Individual non-commercial use is free, but any commercial use by entities with annual revenues of 10 million USD or more requires a Commercial Use Agreement, and derivatives (including distilled models) are bound by the same agreement.
Q2: What hardware does self-hosting require?
A2: A capable NVIDIA GPU, plus the roughly 66GiB full weight set downloaded first. If VRAM is short, enable FP8 quantization (fp8-cast, or fp8-scaled-mm on Hopper and above) with CPU or disk offload. The production-grade DFR pipeline is more VRAM-hungry.
Q3: What is the actual 4K output resolution?
A3: 3840x2176 -- the README explicitly notes it is not 3840x2160, so do not fill in the habitual value in scripts. Default output is 1024x1536 at 24 fps, and the temporal upscaler can raise the frame rate further.
Q4: Can I use it for commercial projects?
A4: It depends. Entities below 10 million USD in annual revenue may use it commercially, but any distribution of derivatives must include the complete agreement. At or above that line, every commercial use requires a purchased Commercial Use Agreement. Additionally, commercial use may not train competing models or power products that directly compete with Lightricks.
Q5: How should I choose between it and Kling 4.0 or Seedance 2.5?
A5: It depends on your needs. For convenience, top-tier quality, and pay-as-you-go billing, pick a cloud flagship. For data privacy, batch generation without per-call fees, and LoRA fine-tuning of proprietary styles, LTX-2 is the only self-deployable option. The two are not mutually exclusive -- many teams generate on the cloud and customize privately on their own hardware.
Discussion
Would you pull 66GiB onto your own GPU? Tell us your pick in the comments: cloud flagship, or the freedom of self-hosting.