Skip to main content
oneneural

Production

The best open video models in 2026, if you actually have to ship

Wan2.2 is the safest general starting point, and it is not the answer to every video product. License territory and cost per accepted clip decide more of this than output quality does.

By Subodh Jena24 min readintermediate

A product team watches twenty generated clips, picks the one people liked best, and starts the integration. Then procurement reads the license. Infrastructure measures the wall time for a single five-second clip in minutes rather than seconds. The product lead discovers that the "image control" in the demo meant a starting frame, not the character consistency the roadmap assumed.

The demo answered whether the model makes a good clip. That is a different question from whether the model can become a product.

The best open video model is the exact checkpoint whose license and controls fit the thing being built, at a cost per accepted clip the business survives. That standard produces a default rather than a winner.

Wan2.2 is the safest general starting point in this comparison. LTX-2.3 is the stronger technical fit for synchronized audio and video. Kandinsky 5 Lite fits when permissive weights and camera control both matter. HunyuanVideo-1.5 and Helios each win a narrower argument. So do LongCat-Video and Bernini.

Start here, then look for the condition that reverses it

Product requirementStart withWhat can reverse the decision
General text-to-video and image-to-videoWan2.2The A14B variants cost far more compute than the 5B model
Synchronized generated audio and videoLTX-2.3Its license requires a paid agreement at USD 10M revenue
Camera-directed generation, permissive weightsKandinsky 5 LiteNo optimized production server exists for it
Compact T2V or I2V with first-party trainingHunyuanVideo-1.5The licensed territory excludes the EU, UK, and South Korea
Progressive or long-form generationHeliosStreaming is text-only, and the speed numbers are author-reported
Native video continuationLongCat-VideoThe surrounding ecosystem is close to empty
Instruction-based video editingBerniniThe full system adds a planner model and a second stack

The table is conditional on purpose. Model families hide different checkpoints behind one marketing name, and product requirements hide different meanings of the word "control." A first frame is not a first-and-last-frame constraint. Pose sequences and reference characters solve a third problem, camera adapters a fourth, editable source video a fifth.

Two corrections come before any of that.

The open lane and the frontier lane have separated

Alibaba's newest video model is not downloadable. Model Studio documents wan2.7-t2v and wan2.7-i2v as API models with a create-task-then-poll flow, 720P and 1080P output, and no weights on offer. The newest Wan family with published checkpoints is still Wan2.2. Lightricks moved the same way from the other direction, having deprecated its LTX-2 API tier on 2026-07-15 with removal set for 2026-08-15 and those requests auto-served on LTX-2.3 in the meantime.

The preference gap is measurable. On Artificial Analysis's with-audio video arena, read on 2026-08-01, Gemini Omni Flash leads at 1244 Elo and the best open-weight entry, LTX-2.3 Fast, sits at 977. That is a preference score across a voting pool rather than a capability measurement, and it is still a gap worth naming out loud before anyone commits a quarter to self-hosting.

None of this makes open weights the wrong call. It changes what the argument for them has to be. Deployment control and exact version pinning are real reasons. So are fine-tuning and a specialist control surface. Beating the frontier on general output quality is not one of them right now.

Open weights are not open source

People search for open-source video models. Almost everything they find is better described as open-weight.

The Open Source AI Definition asks for four freedoms: use for any purpose without permission, study and inspect the components, modify for any purpose including changing the output, and share with or without modifications. Exercising those freedoms requires the preferred form for making modifications, which the definition breaks into three parts: data information detailed enough that a skilled person could build a substantially equivalent system, the complete training and inference code, and the parameters. The OSI's own page on the topic states that open weights alone fall short, because they carry neither the training process nor the data details the definition asks for. No model in this comparison has been shown to clear the full bar.

The practical distinction survives anyway:

  • Permissively licensed open weights use Apache 2.0 or MIT for the exact checkpoint. Five models here qualify: Wan2.2, Kandinsky 5, LongCat-Video, Helios, Bernini.
  • Custom-license open weights are downloadable under an agreement that can restrict territory, company size, competitive use, and what happens to the outputs. LTX-2.3 and HunyuanVideo-1.5 sit here.
  • Public API only means the model is callable and its weights are not available for inspection. Wan2.7 sits here.

A permissive checkpoint license still does not settle training-data rights or output ownership. Likeness questions survive it. So do trademark and indemnity. Multi-model pipelines also inherit the licenses of their text encoders and upscalers, and of every adapter, safety model or prompt enhancer bolted onto the graph. The license audit follows the dependency graph, not the marketing name.

The leaderboard cannot choose the foundation

The Arena.ai open-source text-to-video board still self-reports a 2026-07-05 update. It places Kandinsky 5 Pro at 1172 with a 20-point interval on 2,018 votes, HunyuanVideo-1.5 at 1169 on 4,271 votes, LTX-2 19B at 1143 on 53,873 votes, and Wan2.2 A14B at 1131 on 10,418 votes.

Read the intervals before the ranks. Kandinsky Pro spans 1152 to 1192 and Hunyuan spans 1153 to 1185, so the three-point lead is not a lead. LTX-2 19B touches both. The entire top four form one overlapping chain, and the model at the top has a quarter of the votes of the model in third, which is why its interval is twice as wide. Six open models appear on that board. Helios, LongCat-Video, Bernini and LTX-2.3 are all absent from it.

Artificial Analysis shows a different failure mode. Its with-audio open-weight pool contains four entries and every one of them is an LTX variant, so "leading open-weight model with audio" describes a comparison against itself. Remove audio and LTX-2 Pro at 1127, LTX-2.3 Fast at 1125, and LTX-2 Fast at 1122 collapse into a five-point spread across three checkpoints. That board also moves a point or two between reads, so any Elo quoted in an article is a timestamp rather than a fact.

Preference tests answer whether evaluators liked one rendered clip more than another. Product selection needs something else: instruction adherence on its own prompt distribution, identity consistency across shots, editable controls, unsafe-output rates, retry behavior, p95 latency, cost per accepted clip.

The automated benchmarks close part of that gap and carry their own expiry date. Zheng and collaborators built VBench 2.0 because earlier metrics rewarded what they call superficial faithfulness, and it scores five dimensions those metrics skipped: human fidelity, controllability, creativity, physics, commonsense. T2VWorldBench runs 1,200 prompts across 60 subcategories and reports overall scores generally below 0.70. T2VPhysBench tests twelve physical laws, finds every model below 0.60 in every law category, and reports that law-specific prompt hints fail to repair the violations.

All three evaluate a 2025 cohort built around Sora and Wan2.1, not the checkpoints on the current boards. For current models the newer work is more useful. Wu and collaborators built WorldReasonBench to stress-test generators as future world-state predictors across 436 curated test cases, with a companion preference set of roughly 6,000 expert-annotated pairs. EntityBench measures cross-shot entity identity across 140 episodes and 2,491 shots, behind a fidelity gate that admits only entities the model rendered correctly in the first place.

A plausible clip can still be wrong. Production finds out either way.

Wan2.2 is the safest general default

The safest default combines a permissive license with broad task coverage and the deepest serving ecosystem. The Wan2.2 repository and its Hugging Face model metadata both state Apache 2.0. The family covers text-to-video, image-to-video, a unified text-and-image path, speech-driven video through S2V, and character animation and replacement through Animate.

One family name conceals two operational classes. TI2V-5B is a dense 5B model that unifies T2V and I2V behind a high-compression VAE. The A14B models are mixture-of-experts, about 27B total parameters with roughly 14B active at each denoising step. S2V-14B and Animate-14B are dense 14B, not MoE. A benchmark row that says "Wan2.2" without naming the variant has dropped the fact that matters most.

The stack around it is the broadest here. Diffusers exposes the Wan2.2 T2V, I2V, TI2V and Animate pipelines. ComfyUI ships native Wan2.2 workflows including a first-and-last-frame path. vLLM-Omni serves asynchronous video jobs through a /v1/videos endpoint, with knobs for the MoE two-stage split. SGLang-Diffusion and LightX2V add further optimized paths.

Support is not uniform across that surface. Diffusers is explicit that Animate expects preprocessed pose input, and that integrating the preprocessing steps is planned for a future release. The first-and-last-frame path is a ComfyUI workflow, not a separate Alibaba checkpoint. There is no first-party trainer either: the README points at ModelScope's DiffSynth-Studio for LoRA and full training, which is community work rather than an Alibaba deliverable. The one genuine first-party editing model, Wan2.1 VACE, is reference-conditioned and control-conditioned rather than instruction-driven.

The compute boundary is public and blunt. The Wan efficiency table reports 534.7 seconds and 22.9 GB peak memory for TI2V-5B at 720p on one RTX 4090, with that memory figure measured under model offload and CPU text-encoder flags. The A14B text-to-video path reports 1041.5 seconds and 59.8 GB on a single H100 or H800. Eight GPUs cut wall time to 155.1 seconds and raise aggregate GPU time by about 19%.

One new checkpoint arrived after that release cycle and is easy to miss, because the README does not mention it. Wan-Dancer-14B, published in July 2026 under Apache 2.0, does music-driven minute-scale dance video through global keyframe planning and local temporal refinement. It is a specialist, not a replacement.

Wan2.2 is the default because it creates the fewest immediate disqualifiers. It is neither cheap nor trainable out of the box, and it is not the best model at any single task.

LTX-2.3 wins on capability and charges for it in the license

The most integrated multimodal system here is LTX-2.3. The model card describes a DiT-based audio-video foundation model that generates synchronized video and audio from one network, conditioned on text, images, video or audio. Current checkpoints are named ltx-2.3-22b-dev and ltx-2.3-22b-distilled, replacing the 19B LTX-2 family. First-party IC-LoRA adapters cover LipDub, Motion Track Control, Union Control and HDR.

The training story is unusually complete. The trainer quick-start ships recipes for forward and backward video extension, V2V IC-LoRA, inpainting and outpainting, audio-to-video, video-to-audio, joint audio-video transformation, plus full fine-tuning. A recipe is not a polished pretrained adapter. It is a credible path to building one.

Quality comes from a graph rather than a checkpoint. The recommended production path is a two-stage pipeline pairing the base model with a spatial upscaler and a distilled LoRA, both marked required. The temporal upscaler is supported by the model and reserved for future pipelines, so it is not part of today's recommended path. Prompt handling runs through Gemma 3, which doubles as the multilingual text encoder backbone and the built-in prompt enhancer. Memory planning, model caching, dependency pinning and failure isolation all have to cover the whole graph.

Version drift is a real trap in this repository. Union Control, Motion Track, HDR and LipDub carry the current LTX-2.3-22b naming. The camera-control adapters for dolly, jib and static moves, along with pose control and the detailer, all carry LTX-2-19b naming and sit in the same list. An adapter cannot be promoted to 22B support by proximity.

The LTX-2 Community License is the harder boundary, and its wording is stricter than the summaries suggest. Entities with annual revenues of at least USD 10,000,000 must obtain a paid commercial license, so a company at exactly ten million is captured rather than exempt. Revenue is calculated on an aggregative basis across subsidiaries, affiliates and companies under common control, with control defined as more than fifty percent of voting securities. Breach carries liquidated damages at double the fee that would otherwise have been paid. Attachment A prohibits use in any product that directly competes with or substitutes for Lightricks offerings, applies its restrictions to outputs and derivatives, treats models trained on outputs as derivatives bound by the same agreement, and requires the restrictions to be carried into downstream agreements as enforceable provisions.

For generated audio-video or learned multimodal control, LTX-2.3 is the better technical fit by a wide margin. Legal review belongs before the prototype, not after it.

Five models that win narrower arguments

Kandinsky 5 is more than one product

Kandinsky 5 pairs a 2B Lite family with a 19B Pro family under MIT, alongside image models the video comparison usually ignores. It covers T2V and I2V at five and ten seconds, ships official camera-motion LoRAs for named trajectories, and provides a first-party LoRA trainer.

The official runtime table makes the split concrete. On an H100 after a warm run, Pro at 100 NFE takes 1,241 seconds for a five-second HD clip and 560 seconds for five-second SD. Lite takes 139 seconds at the same NFE, and the 16-step distilled Lite path takes 35 seconds. Those are vendor measurements across different checkpoints and inference configurations, and the five-second numbers depend on Flash Attention 3, which is not a default-install assumption.

Two caveats change how the ecosystem reads. ComfyUI's Kandinsky 5 coverage is Lite only, with no documented Pro workflows. And Kandinsky is absent from the vLLM-Omni supported video model list that carries Wan, LTX, Helios and HunyuanVideo-1.5. Pro is a patient batch-quality candidate. Lite is the plausible product foundation when explicit camera control and permissive weights both matter, and its serving story is the weakest of any model here that is otherwise recommendable.

HunyuanVideo-1.5 is technically compact and legally narrow

HunyuanVideo-1.5 is an 8.3B T2V and I2V family with 480p and 720p checkpoints, distilled variants, a separate super-resolution module to 1080p, LoRA and full training, three caching schemes, offload, plus distributed inference. Tencent documents a 14 GB minimum GPU memory path with offload enabled. Diffusers, ComfyUI, vLLM-Omni and LightX2V all support it.

The 75-second figure attached to Hunyuan needs its conditions restated every time it is quoted. Tencent reports that a single RTX 4090 generates video within 75 seconds using the step-distilled 480p image-to-video model. It says nothing about text-to-video, 720p, cold starts or concurrent service.

The Tencent Hunyuan Community License is where global products stop. It defines the Territory as worldwide excluding the European Union, United Kingdom and South Korea, then prohibits using, reproducing, modifying, distributing or displaying the works or their outputs outside that Territory. The 100 million monthly-active-user threshold is measured across all products or services made available by or for the licensee, in the calendar month preceding the version release date. The agreement separately bars using the works or outputs to improve any other AI model.

For an eligible regional batch product, Hunyuan is compelling on cost and footprint. For a global SaaS product, the territory clause ends the evaluation before latency ever gets measured.

Helios is a streaming experiment worth a bake-off

Helios is a 14B autoregressive family under Apache 2.0, with Base, Mid and Distilled checkpoints. It generates 33 frames per chunk and supports T2V, I2V, V2V, minute-scale generation, plus an interactive mode the maintainers mark as still under development. Diffusers merged its pipelines, and both SGLang-Diffusion and vLLM-Omni carry integrations.

The unusual part is vLLM-Omni's WebSocket path, which exposes progressive Helios-Distilled generation over /v1/realtime/video. The documented endpoint is text-only, and the page says image and reference input are intentionally excluded for now. Progressive delivery of image-conditioned work is not on that path today.

Two performance numbers circulate together and should not. The maintainers report 19.5 generated frames per second in end-to-end inference on a single H100, and separately report roughly 6 GB of VRAM under low-VRAM mode with leaf-level group offloading. Those are different runs at different speeds. Both are author-reported model-path figures with no matched end-to-end service configuration, no first-chunk delay measurement, no evidence about behavior under concurrency.

Helios earns a bake-off when progressive delivery or duration defines the product. It has not earned a production real-time claim.

LongCat-Video owns continuation and almost nothing around it

LongCat-Video is a dense 13.6B model whose weights are MIT-licensed, unifying text-to-video, image-to-video and video continuation in one network, with scripts for segmented long video and interactive generation. Native continuation matters because it is part of the model's published task surface rather than a community graph assembled around a start-frame checkpoint.

The ecosystem is the price. There is no upstream Diffusers video pipeline, no ComfyUI core workflow, no vLLM-Omni path, no SGLang-Diffusion entry, no general quantization guide, no official trainer. Anyone evaluating this should be careful not to credit it with LongCat-Image's tooling, which does exist in most of those places and belongs to a different model.

LongCat is a valid foundation for a continuation-heavy product whose team is ready to own the wrapper and the queue, and to validate cancellation behavior, memory ceilings and boundary drift themselves. It is not the low-risk general choice.

Bernini is an editing worker, not a platform

Bernini exposes a shared task interface covering text-to-video, video-to-video, reference-video editing and reference-to-video generation, plus image tasks, all under Apache 2.0. The full system places a Qwen2.5-VL-7B semantic planner in front of a Wan2.2-derived 14B renderer. Renderer-only packages ship separately at 14B and at 1.3B, the smaller one fine-tuned from Wan2.1 rather than Wan2.2.

That separation is the whole decision. The planner is the reason to run full Bernini on complex instructions, and it is a second model to deploy, version and pay for. ByteDance's published training documentation is renderer-only and full-finetune only, with no planner training and no LoRA path. The optional prompt enhancer calls an external OpenAI-compatible vision-language endpoint through its own API key and base URL, which adds credentials, latency, a privacy boundary and a new failure mode to a video pipeline.

Put Bernini behind a task-specific editing API. Do not make the general text-to-video queue inherit its complexity.

Three more worth watching

SkyReels V3 covers reference-to-video, video extension and a 19B talking-avatar model under the Skywork Community License, which does permit commercial use subject to its terms. Lance is a 3B unified multimodal model under Apache 2.0, which its authors describe as a research project rather than a polished product model, trained at 768px images and 480p video. NVIDIA Cosmos 3, released at the end of May 2026 under OpenMDW 1.1, is an omnimodal world model at 16B and 64B that emits action commands alongside video, which puts it in a robotics conversation more than a content one. All three belong on a specialist watchlist rather than in a general ranking.

Control belongs in the model decision

Required controlBest documented fitWhat the label actually means
Joint generated video and audioLTX-2.3One multimodal model produces synchronized outputs
Speech-driven character videoWan2.2 S2V or LTX LipDubAudio drives a human performance, not a soundtrack
Pose-driven animationWan2.2 AnimatePreprocessed pose and face sequences drive a reference character
First and last frameWan2.2 through ComfyUIAn upstream workflow constrains both temporal boundaries
Learned camera movementKandinsky 5Official LoRAs target named trajectories
Tracked object movementLTX-2.3Motion Track and Union Control supply structural signals
Native continuationLongCat-VideoA source clip continues through a first-party task
Progressive long generationHeliosAutoregressive chunks can be delivered incrementally
Instruction-based editingBerniniSource and reference media feed named editing tasks

This matrix is more useful than a generic controllability score. A product that needs identity preservation should evaluate identity. One that needs camera choreography should evaluate exact trajectories, because prompting "dolly in" is weaker evidence than an adapter trained for that motion.

Customization draws another boundary. LTX-2.3 and HunyuanVideo-1.5 offer the broadest first-party LoRA and full-training paths. Kandinsky offers official LoRA training. Helios offers full distributed training with no official LoRA recipe. The Wan community ecosystem contains many adapters while the base repository provides no general trainer. LongCat and the Bernini planner offer nothing at all.

Training code is where the work starts. Dataset rights, caption quality, filtering, evaluation, checkpoint lineage and rollback all stay with the product team.

Managed pricing and raw GPU time answer different questions

Managed endpoints are the fastest way to learn whether users accept the output. Comparing their prices is harder than it looks, because providers pick different durations, resolutions, frame rates, checkpoints, speed modes and billing rules.

Endpoint, read 2026-08-01Published priceThe boundary that matters
LTX-2.3 FastUSD 0.06 per second at 1080pPro is 0.08; 1440p doubles the rate and 4K quadruples it
Wan2.2 A14B on falUSD 0.04 per second at 480p580p is 0.06 and 720p is 0.08; seconds are billed at 16 fps
HunyuanVideo-1.5 on falUSD 0.075 per second at 480pThe default 121 frames is roughly a five-second clip at 0.375
Kandinsky 5 Lite distilled on falUSD 0.05 per five-second clipFlat regardless of resolution, and not the 19B Pro model
LongCat-Video on falUSD 0.04 per second at 720pSeconds are calculated at 30 fps
Bernini-R on falUSD 0.08 per second at 848pxLongest-edge scaling, and this is the renderer without the planner

Self-hosting removes the endpoint markup and adds an operating system.

At Runpod's published on-demand rates, an RTX 4090 is USD 0.69 per hour and an H100 PCIe is USD 2.89 on Secure Cloud, with Community Cloud materially cheaper for the same card. Combine the Secure Cloud rate with the Wan efficiency table and the raw GPU cost lands near USD 0.10 for one TI2V-5B 720p job on a 4090, USD 0.84 for one A14B 720p job on a single H100, and USD 1.00 for the eight-H100 run. Eight GPUs buy latency, not cheaper aggregate GPU time.

None of those numbers is a product unit cost. The real denominator is:

cost per accepted clip = total generation and operating cost / accepted outputs

The numerator absorbs cold starts and idle capacity, CPU and RAM, storage and egress, input processing and moderation, failed jobs and retries, observability, engineering time. The denominator falls every time a user rejects a clip for identity drift, broken text, a physics error or weak instruction adherence. A model that costs less per second and gets rejected twice as often is the expensive one.

Start on a managed endpoint while demand and acceptance are unknown. Keep the job schema, artifact storage, model configuration and evaluation portable from day one. Move a measured, stable workload to self-hosting when the accepted-output economics justify the operational burden, not before.

The model is one worker in the product

A reliable video product wraps inference in fairly ordinary engineering.

The durable job needs idempotency and cancellation, progress and timeouts, retry limits and quotas, backpressure, a dead-letter path. The generation record needs the model and checkpoint hash, the inference-code version and precision, the sampler and step count, the seed and any prompt rewrite, the controls, the host. Output checks need both a policy layer and a product-specific evaluation.

Provenance is now part of the product surface rather than a compliance afterthought. Article 50 of the EU AI Act applies from 2026-08-02 and requires providers of generative systems to mark outputs in a machine-readable format, detectable as artificially generated, as far as is technically feasible. The Digital Omnibus postponed the high-risk deadlines and left Article 50 standing, with reporting of a grace period into December 2026 for machine-readable marking of systems already on the market. C2PA specification 2.4 is the current technical provenance standard, published as-is with the usual warranty disclaimers and no promise that a manifest survives every downstream transformation. NIST AI 100-4, final since November 2024, recommends combining provenance, labeling, watermarking and detection rather than betting on any one of them.

Moderation and consent are product features. So are audit and provenance. A checkpoint ships none of them.

A closed API can be the responsible choice

Open weights earn their engineering cost at stable utilization, once a real workload has been measured rather than guessed. A closed API often wins when the team needs the best available general output, global operational support, burst handling or a launch date that does not include hiring GPU operations. A team with volatile traffic may pay less for provider margin than it would for idle accelerators and an inference platform it has to staff.

The compromise that survives both directions is a portable product layer over a managed model. Keep requests and artifacts, evaluations and policies, model configuration under product control, so the provider can change before the product has to.

Choosing open weights by ideology is as careless as choosing a closed API by demo quality.

Run a product bake-off, not a beauty contest

The responsible selection sequence is short:

  1. Define the output mode, the control surface, the latency target and the acceptance test.
  2. Eliminate ineligible licenses and dependencies before anything gets benchmarked.
  3. Run exact named checkpoints against a versioned product prompt set.
  4. Measure accepted-output rate, end-to-end latency, retries and safety failures.
  5. Price managed, self-hosted and hybrid operation at expected utilization.
  6. Build cancellation, observability, provenance and rollback around the winner.

Wan2.2 remains the safest general starting point, and it holds that position by creating the fewest immediate disqualifiers rather than by being the best model on the list. Every other model here beats it at something narrower. The teams that ship are the ones who worked out which narrow thing their product actually is before they opened a single demo reel.

oneneural™ builds production AI systems around the checkpoint: model routing and serving, evaluation and controls, safety, and the operating economics that decide whether a video product survives its own success. Discuss your AI project.

References

Model releases and licenses

Benchmarks and evaluation

Serving, pricing, governance

Last updated: August 1, 2026