Key Takeaways
- Wan 2.2 is the only top-tier local video model with zero license restrictions. Apache 2.0, unrestricted commercial use, no revenue caps, no territory exclusions — and the highest verified open-source VBench quality score (~84.7%).
- InVideo bundles 200+ models — including Kling 3, Veo 3.1, and Seedance 2.5 — into one browser-based pipeline starting at $17/month (Plus plan, billed annually), with script, voiceover, music, and subtitles handled automatically.
- HunyuanVideo 1.5's license explicitly excludes the EU, UK, and South Korea — for both the model and its outputs. Readers in those regions should use Wan 2.2 or LTX-2 instead.
- LTX-2 is the fastest of the local trio and the only one with built-in synchronized audio, free commercially for companies under $10M annual revenue.
- 12GB VRAM is the realistic floor for serious local video generation. Below that, InVideo becomes the more practical option.
- Local models generate raw silent clips of 5–20 seconds, not finished videos. Script, voiceover, music, subtitles, and editing are separate tools you assemble yourself — InVideo does all of this in one pass.
- There is no "Wan 2.7." Download pages offering it are SEO scams — official Wan releases stop at 2.2.
Why 2026 Is a Strange Moment for AI Video
The proprietary video market has been chaotic. OpenAI shut down the Sora consumer app in March 2026, less than six months after launch, after downloads fell roughly 66% from their peak (the API remains live separately). ByteDance's Seedance 2.0 ran into Hollywood lawsuits and a paused global rollout the same month, after cease-and-desist letters from Disney, Paramount, and Warner Bros. — it remains accessible in China but carries legal risk for international commercial use. Alibaba's HappyHorse model topped the quality leaderboards in April 2026 — and never opened to the public.
That chaos is exactly what makes both doors attractive. Open local models make you independent of vendor drama. And InVideo absorbs the drama for you: its subscription bundles access to 200+ models — including Kling 3, Veo 3.1, and Seedance 2.5 — so when one model disappears or gets sued, your workflow doesn't notice.
The Local Door: Three Free Models on Your Own GPU
Three open-weights systems dominate local video generation right now, measured by downloads, community activity, and benchmark results. All three run through ComfyUI, a node-based interface installed on your own machine — not a chat-style tool like Ollama. These are diffusion models, not LLMs.
📍 In One Sentence
Wan 2.2 is the best all-around local video model in 2026 — Apache 2.0, highest quality, no restrictions — while LTX-2 wins on speed and synchronized audio, and HunyuanVideo 1.5 offers the most cinematic look but excludes EU/UK/South Korea users by license.
💬 In Plain Terms
If you just want one answer: get a 12GB+ GPU and run Wan 2.2. It has the best quality, the simplest license, and no fine print.
| Model | License | VRAM | Output | Standout feature |
|---|---|---|---|---|
| Wan 2.2 (Alibaba) | Apache 2.0 — unrestricted | 6–8GB (5B) / 15–25GB (14B) | 480p/720p, ~5s clips | Highest verified VBench quality (~84.7%) |
| LTX-2 (Lightricks) | LTX Community License — free under $10M revenue | 18–20GB quantized, 32GB+ full | 480p–1080p, 5–20s, with audio | Only model with synchronized audio+video in one pass |
| HunyuanVideo 1.5 (Tencent) | Tencent Community License — excludes EU/UK/South Korea | 14GB minimum, 24GB comfortable | 480p/720p, up to 10s | Community favorite for cinematic lighting; lightest on VRAM |
⚠️ Scam alert: there is no "Wan 2.7." Download pages claiming to offer "Wan 2.7 open weights" are SEO scams. Official Wan releases stop at 2.2 — only download from the official GitHub or Hugging Face repositories linked below.
Wan 2.2 (Alibaba) — the quality king, truly free
Wan 2.2 is the most widely deployed open video model: its I2V-A14B repository alone recorded roughly 4.24 million Hugging Face downloads in a single month, with hundreds of community derivatives built on top. It ships in three variants — T2V-A14B and I2V-A14B (mixture-of-experts, 27B total / 14B active parameters), plus a compact TI2V-5B that handles both text- and image-to-video on as little as 6–8GB VRAM. The 14B tier needs 15GB (GGUF Q3) to 25GB (FP8); the official unquantized command asks for 80GB. Its license is Apache 2.0 — genuinely free, unrestricted commercial use, no revenue thresholds, no territory exclusions.
Speed, concretely: a single 5-second clip takes roughly 4–9 minutes on an RTX 4090 (one independently reported figure — Wan 2.2 doesn't natively output longer clips in one pass). To build a 20-second sequence, you'd generate 4 separate 5-second clips and stitch them — call it 16–36 minutes of raw generation time, plus manual editing to join them smoothly. That range is an extrapolation from the per-clip figure, not a directly measured 20-second benchmark.
LTX-2 (Lightricks) — speed plus synchronized sound
LTX-2 is the only open model in this trio that generates synchronized audio and video in a single pass — footsteps, ambience, and effects arrive with the picture. It is also the fastest of the three and the most forgiving on hardware. The architecture is a 22B diffusion transformer; LTX-2.3 (March 2026) remains fully supported alongside the current LTX-2.5 release. The license is the LTX Community License — free for commercial use if your company's total revenue is under $10M per year, with a paid commercial license required above that threshold. (Some third-party write-ups incorrectly call it Apache 2.0 — the official license page is the only reliable source.) Hardware needs run 18–20GB VRAM quantized, 32GB+ at full precision; on 12GB cards, the older LTX-Video 0.9.5 remains the practical choice.
Speed, concretely: LTX-2 is qualitatively the fastest of the trio, with near-real-time previews on high-end consumer cards — but no independently verified minutes-per-clip figure on an RTX 4090 exists as of this writing, so we won't invent one. The one hard number available is from Lightricks' own benchmark on datacenter-class "Nvidia superchips" (not a consumer GPU): a 10-second clip in about 6.8 seconds. Treat that as a ceiling for what the architecture can do on serious hardware, not what your home rig will see.
HunyuanVideo 1.5 (Tencent) — the cinematic look, with a legal catch
Tencent's 8.3B model, released November 2025, is a community favorite for cinematic lighting and texture, and the lightest of the trio on VRAM: 14GB minimum with offloading, 24GB comfortable, at roughly 75 seconds per 480p clip on an RTX 4090. It generates 480p/720p natively, up to 1080p via built-in super-resolution, clips up to 10 seconds.
Speed, concretely: at ~75 seconds per 5-second 480p clip, that's roughly 15 seconds of render time per second of video. Its native max clip length is 10 seconds, so a 20-second sequence means two generations at max length — extrapolating the per-second rate, call it roughly 5 minutes of raw generation for 20 seconds of footage, before stitching. This is an extrapolation from the sourced 5-second figure, not a directly measured 10-second or 20-second benchmark.
⚠️Warning: License warning — read before downloading. HunyuanVideo 1.5 uses the Tencent Hunyuan Community License, not Apache 2.0. The license does not apply in the European Union, the United Kingdom, or South Korea — users in those regions are not authorized to use the model or its outputs. It also caps use at 100 million monthly active users and prohibits training competing models on its outputs. If you're in the EU, UK, or South Korea, skip this model: Wan 2.2 covers the same quality tier with zero restrictions.
One to Watch: MiniMax H3
Released August 3, 2026, MiniMax H3 is a 33.1B omni-modal model with native stereo audio, day-one ComfyUI support, and quantized versions that run on an RTX 3060. Two caveats before treating it as a fourth pick: the local release caps at 768p (the full 2K pipeline stays hosted-only), and its Community License reportedly carries its own geographic restrictions and a $20M revenue threshold — check the official model card before committing. Early signs are strong, but three weeks old and production-ready are different things.
The Hardware Gate
Local video generation is free the way a puppy is free: the model weights cost nothing, but the GPU is the real price of entry. Skip local generation entirely if your GPU has under 12GB VRAM and you're not planning to upgrade — none of the three models above run at usable quality below that tier, and a cloud platform will get you better output faster.
Not sure what any of this means for your machine? These guides break it down: VRAM Calculator for exact requirements per model, How Much VRAM Do You Need? for charts across model sizes, Best GPUs for Local AI and Best Budget GPUs for hardware picks, and GPU vs CPU vs Apple Silicon for platform comparisons. One honest caveat: those guides use the LLM VRAM formula (parameters × bits ÷ 8). Video diffusion models also scale VRAM with resolution and clip length, so treat their numbers as a floor, not a ceiling, for video workloads.
| Your GPU | What you can run |
|---|---|
| 6–8GB VRAM | Wan 2.2 TI2V-5B (quantized) — usable, entry quality |
| 12GB VRAM | LTX-Video 0.9.5 — the only serious option at this tier |
| 16GB VRAM | HunyuanVideo 1.5 (license permitting), Wan 2.2 14B at GGUF Q3 |
| 24GB+ VRAM | Everything: Wan 2.2 14B at high quality, LTX-2 quantized |
Rough hardware cost as of August 2026: a used RTX 3060 12GB runs about $170–220, a used RTX 3090 stack about $900–1,100. GPU prices move — verify current pricing before buying rather than trusting these figures past a few months.
What Running Local Video Generation Actually Involves
With local models, you are not installing a video tool — you are assembling a pipeline.
The generation setup. ComfyUI is node-based: you build, or import and debug, a workflow graph of loaders, samplers, and decoders. Expect CUDA version mismatches, PyTorch pins, and the occasional flash_attn install error before your first frame renders.
The prompting. Video models need structured prompts — shot type, camera movement, lighting, subject action — not one-liners. There is no built-in prompt helper and no system-prompt layer; you write the full structure yourself. Our guides on system prompts vs. user prompts and prompt engineering for local models cover fundamentals that transfer directly to video prompting.
Everything around the clip. Local models output raw, silent (LTX excepted) clips of 5–20 seconds. Script, voiceover, music, stock footage, subtitles, and editing are each separate tools you choose, install, and wire together yourself.
Weak (one-liner)
“A dog on a beach”
Structured (what video models need)
“Golden retriever sprinting along a wet shoreline at golden hour, low tracking shot following from the side, shallow depth of field, warm backlight, gentle slow motion, cinematic 24fps”
The Cloud Door: What InVideo Bundles
Want to create AI videos without the local setup? If you don't have a powerful GPU — or simply don't want to spend hours installing and configuring local AI video tools — InVideo is worth trying. Try InVideo's free version →
InVideo is one example of the cloud door — not the only one, and it's worth knowing how it differs from the others before assuming "cloud" means one thing. Runway integrates directly into professional editors (Premiere Pro, Final Cut, DaVinci Resolve), aimed at hybrid AI-plus-editor workflows rather than a finished, assembled video. Luma AI's Dream Machine specializes in native 16-bit HDR output for VFX compositing pipelines (After Effects, Nuke) — a different audience entirely. Pika stays lightweight: fast raw clip generation with no built-in script, voiceover, or stock-footage assembly, so you still need separate tools for everything around the clip — the same DIY-pipeline problem as running a local model, just without the GPU requirement. What sets InVideo apart from all three is that it isn't primarily a raw-generation tool: it's a script-to-finished-video assembler that also gives you access to raw generation models (Kling, Veo, Seedance) when you need them.
InVideo is not a video model — it's the whole production pipeline as a service. You type a topic or paste a script; its v4 agent returns a finished video of up to 30 minutes: AI-generated script, scenes assembled from a 16M+ asset stock library or freshly generated clips, AI voiceover in 50+ languages (including voice cloning), music, subtitles, and brand-kit styling. It runs in the browser — your GPU is irrelevant.
For anyone who wants to start making videos today rather than researching GPUs and quantization formats, InVideo is the practical choice: no local hardware requirement, no ComfyUI installation or CUDA troubleshooting, and a single workflow that already includes the script, voiceover, music, and subtitles most people actually need. It's particularly well suited to creators who care more about the finished video than about controlling the underlying generation model — and since the free tier exists, you can find out whether that fits before spending anything.
Three things stand out for this comparison:
- Model chaos, absorbed. All paid plans include access to 200+ models — Seedance 2.5, Veo 3.1, and Kling 3 among them. When a model gets sued or shut down, InVideo swaps it; your workflow continues.
- Automation is built in, not bolted on. There's an official MCP server, so the entire prompt → script → footage → subtitles pipeline can be triggered programmatically — the kind of harness you'd otherwise build yourself around ComfyUI.
- The free tier is a real test drive. Watermarked and minute-limited, but enough to judge output quality before paying.
Speed, concretely — and the honest catch: a single raw generation is fast, typically minutes. But InVideo's own FAQ puts full end-to-end production of a short film at 2–5 days, not minutes — because choosing and assembling among multiple generated options, not the generation itself, is what takes the time. Treat "2 days as a realistic floor" for a 1–3 minute finished film as the fair comparison point against the local door's 16–36 minutes of raw generation for 20 seconds of unedited footage: InVideo trades your setup and editing time for its own production time, it doesn't eliminate time entirely.
Current plans, starting at $17/month (Plus plan, billed annually, verified August 2026 — check InVideo's pricing page for live figures):
| Plan | Price | Credits/mo | Best for |
|---|---|---|---|
| Free | $0 | limited | Testing the waters (watermarked) |
| Plus | $17/mo ($200/yr) | 75 | Regular creators — all AI models, 4 avatars & voice clones, 100 iStock assets, unlimited watermark-free exports |
| Max | $85/mo ($1,000/yr) | 390 | High-volume channels, 16 avatars |
| Generative | $170/mo ($2,000/yr) | 800+ | Short-film / production volume |
| Elite | $900/mo ($10,800/yr) | 4,250+ | Episodic and commercial scale |
All prices above are annual-billing rates as of August 2026 — paying month-to-month costs more (InVideo's own FAQ cites Plus $20, Max $100, Generative $200, Elite $1,000 per month). Check InVideo's live pricing page before relying on any figure here; plans and prices change.
Cloud or Local: Which Door Is Yours?
The short version, mapped to common situations:
| Your situation | Recommendation |
|---|---|
| No GPU, or under 12GB VRAM | InVideo (cloud) — no local model runs well below this tier |
| Want a finished video with voiceover, not raw clips | InVideo (cloud) — local models don't assemble a full production |
| Deadline-driven, zero setup tolerance | InVideo (cloud) |
| 12GB+ GPU, comfortable with setup, want privacy and $0 marginal cost | Local: LTX-Video (12GB) or Wan 2.2 (24GB for full quality) |
| In the EU, UK, or South Korea | Local = Wan 2.2 or LTX-2 only (HunyuanVideo's license excludes you) |
| Need automation/API at scale without building it | InVideo (cloud, MCP server) |
Who Should Choose InVideo?
Not sure which route is right for you? If you want to avoid the hardware and technical setup, the easiest experiment is simply to try InVideo and see whether its workflow fits your needs. Try InVideo for free →
InVideo is probably the better choice if you:
- Don't own a powerful GPU
- Want to start creating videos immediately
- Don't want to install and configure ComfyUI, CUDA, models, or Python environments
- Want an integrated workflow rather than assembling multiple local tools
- Need scripts, voice, music, subtitles, and video generation in one workflow
- Care more about finished videos than experimenting with the underlying models
Local AI is probably the better choice if you:
- Already own suitable GPU hardware
- Want maximum control
- Want to experiment with models and workflows
- Have strong technical skills
- Prioritize keeping generation locally controlled
- Expect to generate very large volumes and want to optimize marginal generation cost
See Them in Action
- 4 Open Source AI Video Models Compared — Which One's Actually Free? — side-by-side output of LTX 2.3, Wan 2.2, HunyuanVideo 1.5, and MiniMax H3, including the license fine print.
- InVideo Agent One Review — the full prompt-to-finished-video workflow.
- Wan 2.2 Full Local Demo — honest render times on consumer hardware (launch week, July 2025).
- Low-VRAM Wan 2.2 Tutorial — running the 14B model on a 6GB laptop (2025).
FAQ
Can I run AI video generation on 8GB of VRAM?
Barely. Wan 2.2's TI2V-5B variant runs on 6–8GB quantized, at reduced quality and short clip lengths. For the serious models, 12GB is the real floor — and below that, a cloud tool like InVideo is the practical answer.
Is Wan 2.2 really free for commercial use?
Yes. It's Apache 2.0 — unrestricted commercial use, no revenue caps, no territory exclusions, no rights claimed over your outputs. It's the only one of the top local models with zero license fine print.
Can I use HunyuanVideo in the EU or UK?
No. The Tencent Hunyuan Community License explicitly does not apply in the EU, UK, or South Korea — that covers both the model itself and its outputs. Use Wan 2.2 or LTX-2 instead.
Do I need a GPU to use InVideo?
No. InVideo runs entirely in the browser; all generation happens on their infrastructure. A five-year-old laptop works fine.
Can local models produce a complete YouTube video with voiceover?
Not by themselves. Local models generate raw clips of 5–20 seconds (LTX-2 includes synchronized audio; the others are silent). Script, voiceover, music, subtitles, and editing each require separate tools that you assemble into a pipeline yourself.
What's the actual catch with "free" local AI video?
Hardware cost (a capable GPU), setup time (ComfyUI and its dependencies), and the DIY pipeline required around the raw output clips. The model weights themselves genuinely cost $0 per generation, forever.
Is there a Wan 2.7 or newer Wan model?
No. Official Wan releases stop at 2.2. Any site offering "Wan 2.7 weights" is a scam — download only from the official GitHub or Hugging Face repositories.
I'm a complete beginner. Where should I start?
InVideo's free tier — you'll have a finished, narrated video in minutes and can judge whether AI video serves your goals at all. If you later buy a capable GPU and want full control and privacy, the local door stays open.
What's different about running these local models on Mac vs Windows?
ComfyUI runs on Apple Silicon (M1–M4) via PyTorch's MPS backend, but expect roughly 3–5x slower generation than an equivalent NVIDIA GPU — usable, not competitive on speed. The bigger practical issue is software support: CUDA-specific optimizations these models lean on (flash-attention, GGUF/FP8 quantization tooling) are far less mature on Mac, so several community workflows and installation guides assume Windows or Linux with an NVIDIA card and may need adjustment, or simply won't run as documented. One upside: Apple Silicon's unified memory can let you fit a larger model in memory than a discrete GPU with equivalent VRAM would allow, even though it runs slower. If you're buying hardware specifically for local video generation, Windows or Linux plus NVIDIA is the well-supported path; a Mac you already own is fine for experimenting, not the recommended target for serious throughput.
Can I keep the same character consistent across multiple local video clips?
Yes, with extra work — none of the three models guarantee this out of the box across separate generations. The two working approaches: feed the same reference image into image-to-video mode (all three support I2V), or train a small LoRA on your character. Wan 2.2 and LTX-2 both have documented LoRA workflows for this — LTX-2's version is called IC-LoRA (in-context LoRA) and explicitly supports multi-character consistency. Community guidance is consistent on one point: a trained LoRA gives far more reliable results than prompting or a reference image alone. InVideo's brand-kit and AI avatar features solve the same underlying problem differently — a fixed avatar and voice profile you configure once and reuse, no training required.
Try Before You Decide
You don't need to commit to a local GPU setup — or a paid subscription — just to evaluate the cloud workflow. Before buying hardware or spending a weekend on ComfyUI, it's worth spending five minutes the other way first:
1. Try InVideo's free version. 2. Create one short video. 3. Evaluate the output quality and how the workflow felt. 4. Compare that experience against the setup effort a local install would take.
That turns the comparison from something you read about into something you can test yourself in less time than it takes to read the rest of this article.
The Verdict
Go local if you have (or will buy) a 12GB+ GPU, enjoy building your own tools, and value privacy and unlimited $0 generations over convenience. Wan 2.2 is the safest foundation — top quality, Apache 2.0, no fine print — with LTX-2 as the speed-and-sound specialist.
Go cloud if you don't have the hardware, don't want the setup, or need finished videos rather than raw clips. For most people who simply want to make AI-generated videos, the cloud route is the easier starting point: if you don't already have the hardware and technical interest local generation requires, InVideo removes most of that complexity in one prompt, with every model and asset bundled and automation included — starting at $0 to test and $17/month (billed annually) to remove the watermark. The simplest way to find out whether it fits your workflow is to try the free version.
Both doors lead to AI video. The question was never which technology is better — it's which workflow fits your machine, your patience, and your goals.
Sources
- Wan 2.2 on GitHub — official repository, license, and setup instructions.
- Wan 2.2 on Hugging Face — official model card and download.
- LTX model license — official LTX Community License terms.
- LTX-2 model page — official architecture and release details.
- HunyuanVideo 1.5 on GitHub — official repository and LICENSE file, including the EU/UK/South Korea exclusion.
- VBench-2.0 leaderboard — independent benchmark used for quality and physics-faithfulness figures.
- InVideo pricing — official plan and pricing details.
- InVideo MCP server — official automation documentation.
- MiniMax H3 on GitHub — official repository.
- MiniMax H3 on Hugging Face — official model weights.
- InVideo: How Long Does It Take to Make an AI Short Film? — InVideo's own end-to-end production timeline figures (2–5 days).
- ComfyUI system requirements — official Mac/Apple Silicon MPS support documentation.
- LTX Blog: How to Use IC-LoRA in LTX-2 — official character-consistency (IC-LoRA) guide.
