Submit a Tool

Wan Video

by Alibaba / Alibaba Cloud · 2025
No reviews yet
Paid · Wan 3.0 API pricing is calculated primarily per second of generated video and varies by resolution, model variant, and region. Wan 2.1 and Wan 2.2 are open-source releases, although running them requires suitable computing resources. Free plan Open source
Visit Website ↗ Docs
Wan Video

Wan Video is Alibaba’s AI video-generation platform and model family for creating and editing video from text, images, video, audio, and other reference inputs. The official Wan platform at wan.video is associated with Alibaba’s Wan AI ecosystem, while the underlying Wan models are developed by the Wan team and made available through Alibaba Cloud Model Studio as well as open-source releases. The current Wan 3.0 generation expands the platform beyond conventional text-to-video by combining multimodal input, reference control, video editing, and native audio generation in a single video-generation workflow.

Wan 3.0 supports text-to-video, image-to-video, and reference-based video generation. Users can provide text prompts, images, videos, audio, documents, and web-page references depending on the workflow and deployment. The model supports first-frame and first/last-frame control, allowing creators to establish how a generated sequence begins and ends. This makes the system useful for creating more controlled sequences rather than relying only on a single text description.

A major feature of the current Wan 3.0 generation is its ability to produce videos up to 30 seconds in a single generation. Output resolutions include 480P, 720P, and 1080P, while the model can generate at up to 30 frames per second. The model is also designed to handle longer-form visual instructions and can recommend an appropriate duration for a prompt.

Wan 3.0 also expands video generation into audio-visual production. The official documentation states that the model can natively generate dialogue, background music, and sound effects. It accepts audio as an input modality as well, allowing creators to incorporate reference audio into supported workflows. This makes Wan more suitable for complete short-form video concepts than models that generate silent video only.

Reference-based generation is another important capability. Wan 3.0 can work with multiple reference materials in one request, including images, video clips, audio, documents, and web pages. The official documentation allows up to 20 multimodal reference materials per request, subject to the individual input limits. This can be useful when a creator needs to maintain the appearance of a character, product, location, or visual style across generated content.

The platform is also relevant to video editing. Wan’s broader model ecosystem includes dedicated models for video editing, camera-movement replication, effect replication, character animation, and character replacement. Alibaba’s current model guide lists Wan 2.7 video-editing models as well as Wan 2.2 animation models alongside Wan 3.0.

Wan has an extensive open-source history. Wan 2.1 was released as an open suite of video foundation models and supports text-to-video, image-to-video, video editing, text-to-image, and video-to-audio tasks. The Wan 2.1 project also provides checkpoints and inference code and integrates with tools such as ComfyUI and Hugging Face Diffusers.

Wan 2.2 continued this open model approach, with the official Wan-Video GitHub organization listing Wan2.2 as a public repository under the Apache-2.0 license. The organization also maintains projects such as Wan-Animate-2, Wan-Dancer, and Wan-skills. This provides developers and researchers with options for running and experimenting with parts of the Wan ecosystem locally instead of relying exclusively on a hosted interface.

For developers, Wan can be accessed through Alibaba Cloud Model Studio. The current Wan 3.0 API supports asynchronous video-generation workflows and model-specific parameters for prompts, source media, resolution, aspect ratio, duration, callbacks, and generation options. This makes Wan suitable for applications that need to incorporate video generation into a larger software workflow.

The current Wan 3.0 API pricing is usage-based. Alibaba Cloud lists different per-second prices according to resolution, model variant, and deployment region. The standard wan3.0-video model is less expensive than the speed-optimized wan3.0-video-prime model. For example, in several global regions, standard Wan 3.0 is listed at $0.041256 per second for 480P, $0.082513 per second for 720P, and $0.165025 per second for 1080P, while Singapore pricing is listed separately.

Wan therefore occupies two related positions in the AI video market: an open-source model ecosystem through projects such as Wan 2.1 and Wan 2.2, and a newer hosted/API generation platform built around newer Wan models such as Wan 3.0. Users interested in experimentation, research, or self-hosting can explore the open releases, while businesses and developers can use Alibaba Cloud’s managed inference services for production-oriented workflows.

Pricing

Wan’s pricing depends on how the service is accessed. The current Wan 3.0 API through Alibaba Cloud Model Studio uses pay-as-you-go pricing based on video duration and resolution. Standard Wan 3.0 is currently listed at:

  • 480P: $0.041256 per second in several global regions
  • 720P: $0.082513 per second
  • 1080P: $0.165025 per second

In Singapore, the standard model is listed at $0.05/second for 480P, $0.10/second for 720P, and $0.20/second for 1080P. The speed-optimized Wan 3.0 Prime model has higher per-second rates.

Wan 2.1 and Wan 2.2 are also available as open-source model releases, meaning the model itself can be obtained without a subscription fee. However, users running these models locally still need compatible hardware and may incur infrastructure or cloud-compute costs.

Pricing Note

Wan does not have one universal price because its open-source models and managed API are different offerings. API costs depend on model, resolution, duration, region, and input type. Alibaba Cloud’s documentation also states that limited-time promotions may change the effective price displayed in the Model Studio console.

Review

Wan Video provides a broad video-generation ecosystem rather than a single narrowly focused generator. Its current Wan 3.0 model combines text, image, video, audio, document, and web references with video generation, editing, and native audio capabilities. The 30-second generation limit and 1080P support also make it suitable for more substantial video concepts than very short clip-generation systems.

One of Wan’s notable advantages is the combination of managed services and open-source releases. Wan 2.1 and Wan 2.2 provide developers and researchers with access to model weights and source code, while Alibaba Cloud provides a managed API for newer models. This gives the ecosystem flexibility for both experimentation and production development.

The main consideration is that different Wan generations have different availability, licensing, hardware requirements, and pricing structures. The open-source Wan 2.1/2.2 models should not be treated as identical to the current Wan 3.0 managed model. Users should also account for cloud inference costs when using the API instead of running an open model locally.

For developers, researchers, and creative teams looking for multimodal video generation with reference control, audio generation, longer clips, and an established open-source ecosystem, Wan offers a wide range of capabilities across its different model generations.

Key Features

  • Text-to-video generation
  • Image-to-video generation
  • Reference-to-video generation
  • First-frame control
  • First/last-frame control
  • Video extension
  • Video editing
  • Up to 30-second video generation with Wan 3.0
  • 480P, 720P, and 1080P output
  • Up to 30 FPS output
  • Native dialogue generation
  • Background music generation
  • Sound-effect generation
  • Multimodal reference inputs
  • Image, video, audio, document, and webpage inputs
  • Up to 20 reference materials per request
  • Character consistency workflows
  • Video-to-video and editing capabilities through the broader Wan ecosystem
  • Character animation
  • Character replacement
  • Camera-movement replication
  • ComfyUI integration
  • Hugging Face Diffusers support
  • Open-source Wan 2.1 and Wan 2.2 models
  • Alibaba Cloud Model Studio API
  • Asynchronous API generation
  • Local inference for supported open models

Pros & Cons

Pros

  • Supports multiple video-generation workflows.
  • Wan 3.0 can generate videos up to 30 seconds in one generation.
  • Supports 1080P output.
  • Accepts multiple types of multimodal references.
  • Supports native dialogue, music, and sound effects.
  • Provides first-frame and first/last-frame control.
  • Has an established open-source ecosystem through Wan 2.1 and Wan 2.2.
  • Supports ComfyUI and Diffusers.
  • Offers an API for application and business integration.
  • Provides models for video generation, editing, and character animation.

Cons

  • Pricing varies depending on region, model, resolution, and duration.
  • API usage can become expensive for high-volume 1080P generation.
  • Different Wan versions have different capabilities and licensing arrangements.
  • Open-source models require suitable GPU hardware or paid computing infrastructure.
  • The latest Wan 3.0 experience is distinct from the older open-source Wan 2.1/2.2 releases.
  • Some advanced features are distributed across different models rather than one universally available model.

Reviews

No reviews yet

No reviews yet. Be the first to review this tool.

Sign in to write a review

Related Tools

DeepBrain Studios

DeepBrain AI

DeepBrain Studios, officially presented as AI Studios by DeepBrain AI, is an all-in-one AI video creation platform for producing videos with realistic…

Free plan API Freemium
No reviews yet

Hour One Character was an AI video creation platform designed around realistic virtual presenters and digital characters. It enabled users and businesses…

Free plan API Subscription
No reviews yet

HeyGen Interactive Avatar

HeyGen Technology Inc.

HeyGen Interactive Avatar is a real-time AI avatar technology that allows businesses and developers to create interactive digital humans that can listen,…

Free plan API Subscription
No reviews yet