Hardcore Reviews
Hardcore Reviews

Open Video Model Local Deployment: Five-Axis Comparison Review

A four-way open-source video model comparison on the five axes that matter for local deployment: license, VRAM, audio, duration/resolution, and regional restrictions - not generation quality, no cross-model benchmarks, no ranking. LTX-2.5 (22B, distilled/FP8 from 12GB, 24kHz stereo in-pass, LTX Community license with a $10M threshold, gated=auto); Wan 2.2 (TI2V-5B, 8GB with offload, Apache 2.0 and the most permissive, short clips 480p-720p, no native audio); HunyuanVideo 1.5 (8.3B, from 14GB, ~75s per clip on a 4090 step-distilled build, Tencent community license not valid in EU/UK/South Korea with 100M-MAU renegotiation, no native audio); MiniMax H3 (33B dense, 32kHz stereo with dialogue in 11 languages, 15s/768p default while 2K needs an unreleased regenerator, community license excludes EU/UK/South Korea/USA with written authorization required above $20M, VRAM unpublished so not invented). Sources labeled per vendor (bestfreewebresources/pinggy compilations plus official model cards); fal's H3 Max #1 on i2v-with-audio is third-party. Verdicts by scenario, no arbitration: most permissive license picks Wan 2.2, audiovisual-in-one picks LTX-2.5, a 24GB card outside restricted regions picks HunyuanVideo 1.5, reference audio and multilingual dialogue pick H3 (within licensed regions).

Published October 10, 20269 min read
<!-- open-video-model-local-comparison-review | review | Open Video Model Local Deployment: Five-Axis Comparison Review -->

Open-weight video generation stopped being a curiosity in 2026. Four models now ship weights you can download and run on your own GPU: Lightricks' LTX-2.5, Alibaba's Wan 2.2, Tencent's HunyuanVideo 1.5, and MiniMax's H3. The catch is that "open" means four very different things here, and the difference is not picture quality. It is whether you can legally and practically run the model at all, on your hardware, in your country.

This review compares the four contenders on five local-deployment barriers: license terms, VRAM entry price, native audio support, clip duration and resolution, and territorial licensing restrictions. It deliberately does not rank generation quality. We did not run cross-model benchmarks, and quality-first comparisons already exist in our coverage.

If you want the cloud-versus-open story instead, including how the 30-second generation club stacks up on price and output quality, read our earlier open vs cloud video model review. That piece asked which model makes the better video. This one asks which model will actually run on your machine, under your laws, on your power bill.

What Local Deployment Actually Gates

Running a video model locally is a chain of gates, and each gate can stop you cold:

  • License: can you use the outputs commercially, and at what revenue threshold?
  • Territory: does your country's users fall inside or outside the license grant?
  • VRAM: does your GPU clear the model's memory floor, possibly with quantization?
  • Audio: does the model produce synchronized sound, or video only?
  • Duration and resolution: what does a single pass give you before post-processing?

Data sources for this review are the official model cards and license texts for each model, plus third-party roundups and deployment writeups from bestfreewebresources, pinggy, and innfactory. Where two sources disagree, we say so. Where a figure is unpublished, we say that too, rather than guessing.

The Four Contenders at a Glance

AxisLTX-2.5Wan 2.2HunyuanVideo 1.5MiniMax H3
Parameters22B (distilled/FP8 variants lower the bar)TI2V-5B and siblings8.3B33B dense
LicenseLTX Community License, USD 10M revenue threshold, gated downloadApache 2.0, most permissiveTencent Community License, excludes EU/UK/South KoreaCommunity License, excludes EU/UK/South Korea/US
VRAM12GB with distilled/FP8; 80GB+ for full weights8GB with 5B model and offloadFrom 14GBNot published as a unified figure
Native audio24kHz stereo, same passNoneNone (separate Avatar series)32kHz stereo, dialogue in 11 languages
Duration / resolutionUp to about 20s per pinggy; innfactory reports 121 frames at 24fps (about 5s); 4K HDRShort clips; 480p to 720p for the 5BAbout 75s rendered per clip15s at 768p default; 2K needs a regenerator that is not open-sourced

Keep this table in view as we walk each axis. The single biggest surprise is in the territory row.

Axis 1: License Terms and Territorial Restrictions

This is the headline finding of the review, because it is the axis most coverage ignores.

Wan 2.2 is the outlier in the good direction. Apache 2.0 is the standard open-source license most people imagine when they hear "open weights." You can use it commercially, modify it, ship products on it, and no revenue threshold applies. If legal simplicity is your top priority, the decision ends here.

LTX-2.5 uses a custom community license. Lightricks' own materials state plainly that it is not Apache 2.0. The LTX-2.x Community License permits free commercial use for organizations with annual revenue below USD 10 million (inclusive); at or above that line you must negotiate a Commercial Use Agreement. Downloads are also gated on Hugging Face with auto-approval: you must log in and accept the terms at huggingface.co/Lightricks/LTX-2.5 before you can pull the weights, and that gate applies even to the model card. The reference code ships from the official repo, github.com/Lightricks/LTX-2 (plain path: github.com/Lightricks/LTX-2). For most individuals and small studios this is a formality, but it is a formality with consequences for CI pipelines and automated provisioning. We covered the license mechanics and the 66GB download footprint in our LTX-2.5 open-source resource guide.

HunyuanVideo 1.5 carries the Tencent Community License, and its territory clause is the first gate that can simply say no. The license does not apply in the EU, the UK, or South Korea. There is also a scale clause: reaching roughly 100 million monthly active users requires a separate agreement with Tencent. If you or your end users are in the excluded regions, no amount of hardware upgrades gets you to yes.

MiniMax H3 is the strictest of the four on territory. Its community license excludes the EU, the UK, South Korea, and the United States. Above USD 20 million in relevant revenue, written authorization from MiniMax is required. That is a large share of the global market for AI tooling locked out of local deployment, and it means the model's impressive audio capabilities are only legally reachable to a subset of readers of this review.

The pattern worth internalizing: "open weights" and "open to you" are different claims. Two of the four models are legally off-limits to entire continents.

Axis 2: VRAM Entry Price

The second barrier is the memory floor, and here the spread runs from humble to datacenter.

Wan 2.2 wins on accessibility. Its reference implementation is public at github.com/Wan-Video/Wan2.2 (plain path: github.com/Wan-Video/Wan2.2), and the TI2V-5B variant runs in about 8GB of VRAM with CPU offload, which means a gaming laptop or a mid-range desktop card qualifies. Third-party roundups, including the bestfreewebresources compilation we cross-checked, consistently put 5B-class models at the bottom of the memory ladder.

LTX-2.5 is a 22B model with a clever escape hatch. The full-precision deployment wants 80GB or more of VRAM, which is A100/H100 territory. But the distilled and FP8 variants run from 12GB on a single consumer card. The full component set is roughly a 66GiB download, so plan disk as well as memory. The 12GB floor is what makes a same-pass audio-video model viable for prosumers rather than labs.

HunyuanVideo 1.5 starts at around 14GB, and the pinggy deployment writeup documents roughly 75 seconds of rendered clip per run on an RTX 4090 using the step-distilled variant. That positions it squarely for owners of high-end 24GB cards rather than entry-level hardware.

MiniMax H3's VRAM requirement is not published as a unified figure. At 33B dense parameters it is the largest model in this comparison, and local memory needs will depend on quantization choices, but MiniMax has not put an official number on the card. We will not invent one. Treat H3 as the model where you should check community deployment reports for your specific quantization before committing disk space.

A practical note that cuts across all four: the community consensus, echoed in roundups and in our own deployment testing, is that generating short at low resolution and then upscaling with a dedicated model is the most reliable VRAM-saving pattern.

Axis 3: Native Audio

Two of the four models generate sound with the video. The other two generate silence.

LTX-2.5 produces 24kHz stereo audio synchronized in the same generation pass as the video. There is no separate dubbing step, and the architecture ties the audio to the video latents, which matters for lip-sync-adjacent coherence. For anyone building a local pipeline that outputs finished-sounding clips, this is the single differentiating capability.

MiniMax H3 goes further on the audio spec: 32kHz stereo with dialogue generation across 11 languages. It also accepts rich reference inputs through its omni-reference system, which supports up to nine images, three videos, and three audio references per generation. If your goal is dialogue-driven content, H3 has the deepest audio feature set, subject to the territorial restriction from Axis 1. Its cloud API prices at USD 0.08 per second at 768p and USD 0.13 per second at 2K, and a third-party fal leaderboard placed H3 Max first for image-to-video with audio in August, which we report as a third-party data point rather than our own benchmark.

Wan 2.2 and HunyuanVideo 1.5 have no native audio generation. HunyuanVideo has a separate Avatar series for talking-head work, but that is a different model with its own deployment overhead, not a feature of the 1.5 weights. For these two, plan on a post-generation audio pipeline or silent output.

Axis 4: Duration and Resolution

LTX-2.5 has a duration discrepancy we must present honestly. Pinggy's roundup states a maximum of about 20 seconds per pass, extendable through an Extend pipeline. Innfactory's deployment writeup states 121 frames at 24fps, which is about 5 seconds. Both figures are in circulation, the official promotional materials and third-party summaries disagree, and we will not pick a winner between them. Plan your workflows against the conservative figure until you have tested your own build. On the resolution side, LTX-2.5 targets 4K (3840 by 2176) HDR with up to 50fps, which is the highest resolution ceiling in this comparison.

Wan 2.2's 5B variant is built for short clips at 480p to 720p. That pairs with its 8GB memory floor: the small model makes fast, low-resolution drafts, and quality comes from iteration speed rather than single-shot spectacle.

HunyuanVideo 1.5 renders roughly 75 seconds per clip on the 4090 step-distilled setup cited above, which is the longest output window in this comparison and changes what kinds of content are practical locally.

MiniMax H3 defaults to 15 seconds at 768p. The path to 2K output runs through a regenerator component that is not open-sourced, so local deployments are effectively capped at the default resolution. That is a meaningful constraint if 2K was your target.

Scenario Verdicts

No model wins every axis, so the honest verdict format is a mapping from scenario to choice:

  • You want the most permissive license and the lowest hardware bar. Choose Wan 2.2. Apache 2.0, 8GB of VRAM, no territory exclusions, no revenue thresholds. The trade is modest resolution and no audio.
  • You want audio and video generated in one pass. Choose LTX-2.5. The 12GB distilled floor makes it runnable on a single consumer card, 4K HDR is a real ceiling, and the USD 10M license threshold covers most small teams. Accept the gated download and verify the duration figure on your own build.
  • You have a 24GB card and no territorial conflict. Choose HunyuanVideo 1.5. The 75-second output window on a 4090 step-distilled setup is the longest local render in this field, and 8.3B parameters keep the footprint sane. Confirm your jurisdiction is outside the EU/UK/South Korea exclusion list first.
  • Your content is reference-driven audio and dialogue. Choose H3, where the license applies. Eleven-language dialogue generation and a nine-image reference system are unmatched in this comparison, but the EU/UK/South Korea/US exclusions mean a large fraction of readers cannot legally self-host it at all.

Before You Download

Three operational notes from the deployment writeups worth heeding:

  1. The Hugging Face gated-auto pattern is spreading. Expect to log in and accept terms before any download script works, and account for that in automated environments.
  2. Quantized and distilled variants are the difference between "runs on my desk" and "runs in a datacenter." Start there before concluding a model is out of reach.
  3. If you are turning any of this into a repeatable ComfyUI pipeline, our ComfyUI local video workflow SOP walks the LTX-2.5 file layout, the audio-VAE trap that produces silent videos, and VRAM-saving patterns end to end.

And if your interest is specifically in how audio-capable video models price out in the cloud, our audio-video model cost review covers that economics in depth.

The five-axis picture, in one sentence: Wan 2.2 opens the door widest, LTX-2.5 offers the most capability per gigabyte of VRAM, HunyuanVideo 1.5 renders the longest clips for well-equipped machines, and H3 is the audio specialist whose license map determines who gets to use it. Check the territory row before anything else, because it is the one gate no GPU upgrade can fix.

FAQ

Q1: Which open video model has the most permissive license? A1: Wan 2.2, released under Apache 2.0. It has no revenue thresholds, no territory exclusions, and runs in about 8GB of VRAM with offload, making it the lowest-friction local deployment in this comparison.

Q2: Can I run LTX-2.5 on a 12GB GPU? A2: Yes, using the distilled or FP8 variants, which run from 12GB of VRAM. The full-precision model needs 80GB or more. Note the Hugging Face download is gated: you must log in and accept the community license terms first.

Q3: Can I use HunyuanVideo 1.5 in the European Union? A3: No. The Tencent Community License explicitly does not apply in the EU, the UK, or South Korea. There is also a separate agreement requirement once usage reaches roughly 100 million monthly active users.

Q4: How long can LTX-2.5 clips be? A4: Two figures coexist in current sources. Pinggy's roundup reports up to about 20 seconds per pass with an Extend pipeline for longer output, while innfactory reports 121 frames at 24fps, about 5 seconds. Test your own build against the conservative figure.

Q5: How much VRAM does MiniMax H3 need for local deployment? A5: MiniMax has not published a unified VRAM figure for H3. At 33B dense parameters it is the largest model in this comparison, so check community deployment reports for your specific quantization before planning a build.

This article is AI-assisted and human-edited. Last updated: 2026-10-10

FAQ

Which open video model has the most permissive license?
Wan 2.2, released under Apache 2.0. It has no revenue thresholds, no territory exclusions, and runs in about 8GB of VRAM with offload, making it the lowest-friction local deployment in this comparison.
Can I run LTX-2.5 on a 12GB GPU?
Yes, using the distilled or FP8 variants, which run from 12GB of VRAM. The full-precision model needs 80GB or more. Note the Hugging Face download is gated: you must log in and accept the community license terms first.
Can I use HunyuanVideo 1.5 in the European Union?
No. The Tencent Community License explicitly does not apply in the EU, the UK, or South Korea. There is also a separate agreement requirement once usage reaches roughly 100 million monthly active users.
How long can LTX-2.5 clips be?
Two figures coexist in current sources. Pinggy's roundup reports up to about 20 seconds per pass with an Extend pipeline for longer output, while innfactory reports 121 frames at 24fps, about 5 seconds. Test your own build against the conservative figure.
How much VRAM does MiniMax H3 need for local deployment?
MiniMax has not published a unified VRAM figure for H3. At 33B dense parameters it is the largest model in this comparison, so check community deployment reports for your specific quantization before planning a build.

Related

Hardcore Reviews

Embedding Model Comparison: 4 RAG Contenders, One Honest Table

A four-way RAG embedding model comparison (pricing snapshots October 9, 2026): EmbeddingGemma 2, Qwen3-Embedding (0.6B/4B/8B), BGE-M3 and OpenAI text-embedding-3-large, with gemini-embedding-001 as a reference row. All licenses API-verified on Hugging Face (2026-10-09): google/embeddinggemma-2 and Qwen3-Embedding-8B are apache-2.0 and ungated; BAAI/bge-m3 is mit (33.7M downloads). Engineering ledger aligned item by item: context 8192 (EG2/BGE-M3/OpenAI) vs 32768 (Qwen3); dimensions 768 (MRL to 128) / 1024-4096 / 1024 / 3072 (MRL to 256); three self-host free vs OpenAI at $0.13 per million tokens. Discipline first: each vendor's benchmark numbers cited from their own tables only (EG2's MTEB Code +9.92 and gemini-embedding-001's three board leads are official self-reports; Qwen3-8B's ~70.6 multilingual is third-party relayed) - no cross-table ranking; production popularity (relayed from LangChain/LlamaIndex telemetry): BGE-M3 first, Qwen3 second, OpenAI third. Verdicts by scenario, no arbitration: on-device multimodal picks EG2, long documents and Chinese pick Qwen3, hybrid retrieval and mature ecosystem pick BGE-M3, zero-ops picks OpenAI; and since switching means a full re-embed, decide once, decide carefully.

Oct 9, 202610 min read
Hardcore Reviews

Small Model Price War: Haiku 5.5 Ties Luna Only Below 100K

A small-model price-war comparison (October 8, 2026 price snapshots; five columns, four rows): Claude Haiku 5.5, GPT-6 Luna, Sonnet 5.5, and Qwen3.8-27B as the open self-hosting row. Pricing: Haiku 5.5 at $0.10/$0.50 per million (under-100K prompts), Luna $0.10/$0.50 (up to 272K), Sonnet 5.5 at $2/$10 list with cache reads halved to $0.10, and Qwen3.8-27B free weights (Apache-2.0) on your own GPU. Core finding: Haiku 5.5 matches Luna only in the short-prompt tier - above 100K tokens its $0.50 input faces Luna's $0.20, so long-document batch bills invert the picture. Benchmarks cite each model's own official numbers, labeled and never mixed across suites: Haiku 5.5 Terminal-Bench 4.0 at 39.2% and OSWorld 72.4%; no fresh Luna numbers shipped with the release; Sonnet 5.5 at 70.6% on TB4; Qwen3.8-27B has no cross-suite table. Effort settings compared (Haiku 5.5 is the first Haiku-class with the dial; Luna runs none-to-max), contexts (1M for both hosted models; Qwen 262K extensible to 1M). Verdicts by use case - high-volume subagents, batch classification, main-line coding, private deployment - with the site's standing no-stress-test-no-arbitration stance. Four same-topic deep dives linked inline.

Oct 8, 202610 min read
Hardcore Reviews

17 hours vs 16 days: which AI can actually work overnight

A continuous-autonomy comparison (October 7, 2026 basis; a different axis from the site's October 5 single-output-capacity review, which it links rather than repeats): how long can one AI work on a single task, and what does it deliver. Four rows: Ant's Ling-3.1-flash wrote a Lua-to-x86-64 compiler from scratch in about 17 hours (178/182 tests passing, self-reported) and ported a C image library to Rust in about 20 hours with an 8.015x speedup (30/30 checks); Alibaba's Qwen3.8-Max ran a 16-day autonomous build of oh-my-cli (265 commits/127 PRs/151 issues, with a public GitHub trace) plus a ~125-hour paper reproduction that beat the original method (AIME24 49.58% to 52.29%, 7,600 lines of code, 33 training rounds); Google's Argon handled the 800K+ line Zircon migration and a 2.7x libgav1 rewrite (DeepSWE v1.1 77.9%, general public still locked out); reference row DeepSeek-V4-Flash-Vision-Exp scores 59.3 on the same benchmark (MIT, self-hostable). Table columns: model / longest case / verifiability / access and cost / sourcing. An architecture section answers "Ling vs Qwen": Ling uses 7 KDA + 1 Gated MLA layers, Qwen 3 Gated DeltaNet + 1 Gated Attention, both 512-expert MoE. License red lines reported faithfully: Qwen3.8-Max open weights ship under a custom qwen3.8-max license (revenue share, threshold not finalized), not Apache 2.0 (only the 27B sibling is), and the open build is text-only with forced thinking - not the API version with vision, 1M context and tools; the "half-open" controversy is presented from both sides; 4.89TB of weights need 24 GPUs (BF16) or 16 (FP8). All numbers self-reported; DeepSWE v1.1 compared only within v1.1.

Oct 7, 202610 min read