The new Wan 3.0 model is set to transform AI video creation! It handles multiple tasks simultaneously, from PDFs and PPTs to audio.
Alibaba has announced a major update in the field of AI video generation with the official launch of its new Wan 3.0 model. This model goes beyond simply creating videos from text or images; it can utilize a wide range of inputs, including text, images, audio, and video, as well as PDFs, PPTs, documents, spreadsheets, and web pages. Its standout feature is the ability to generate videos up to 30 seconds long in a single pass.
With Wan 3.0, users can create videos lasting up to 30 seconds in one generation cycle—double the 15-second limit of the previous Wan 2.7 model. According to Alibaba, this extended duration allows the model to convey a more complete narrative within a single video, eliminating the need to stitch together multiple short clips.
The model is specifically designed to produce longer and more complex videos. It possesses the capability to manage elements such as characters, scenes, camera movements, dialogue, and audio based on a single creative input.
One of Wan 3.0’s most intriguing features is its support for multimodal inputs. Instead of relying solely on text prompts, users can utilize various file types—such as PDFs, PowerPoint presentations, Word documents, and Excel spreadsheets—as references. Additionally, web pages, images, audio files, and existing videos can also serve as reference material.
For instance, a promotional video could be generated based on a product presentation or a planning document. This expands the potential applications of AI video generation beyond mere creative experimentation, opening up possibilities for use in marketing, product demonstrations, and other professional tasks.
Wan 3.0 also features native support for audio-visual generation. In other words, the model does not merely create the visuals for a video; it can also generate the accompanying audio and sound design. According to Alibaba, it can be used to produce videos that include dialogue, background sounds, and other audio elements.
Another key focus of the model is maintaining consistency regarding characters and objects. Many AI video tools suffer from issues where a character’s face, clothing, or a product’s design changes across different frames. Wan 3.0 has been designed to preserve reference details across frames with greater accuracy.
Wan 3.0 supports video output up to 1080p resolution. According to Alibaba Cloud Model Studio, users can generate videos in 480p, 720p, and 1080p quality. API pricing is determined by video duration and resolution; costs start at $0.05 per second for 480p output, $0.10 for 720p, and $0.20 for 1080p.
However, a crucial point regarding Wan 3.0 is that API access is currently subject to approval, as Alibaba Cloud is offering it as a public preview. This means that while the model’s capabilities are available, API access is not immediately open to every user.
The launch of Wan 3.0 comes at a time of rapidly intensifying competition among AI companies in the video generation space. Alibaba has already been active in this field with its Wan series, and with Wan 3.0, the company has focused on integrating video duration, input formats, and audio-visual generation into a single system.