Alibaba Releases Wan3.0 Video Model: 30-Second Generations, Documents to Video in One Click, "Everything Becomes Video"

8.7 On the evening of August 6, Alibaba Cloud officially announced that its next-generation video generation model Wan3.0 has entered public beta. The model features comprehensive upgrades in generation length, multimodal input, reference consistency, and visual realism. Its most disruptive highlight is native support for converting office documents—including doc, xls, ppt, and pdf—directly into video, expanding AI video generation from content creation into broader productivity scenarios.

On the evening of August 6, Alibaba Cloud officially announced that its next-generation video generation model Wan3.0 has entered public beta. The model features comprehensive upgrades in generation length, multimodal input, reference consistency, and visual realism. Its most disruptive highlight is native support for converting office documents—including doc, xls, ppt, and pdf—directly into video, expanding AI video generation from content creation into broader productivity scenarios.

2026080701

30-Second One-Shot: Narrative Becomes Possible

Wan3.0 extends single-generation length to 30 seconds, matching ByteDance's Seedance 2.5 and making it one of the few models capable of reaching this duration.

Longer generations unlock narrative space. Continuous tracking shots, one-take sequences, and complex camera movements can now be completed in a single generation, moving AI video from "producing a stunning shot" to "telling a complete story"-
. In real-world tests, Wan3.0 maintained consistent facial features, clothing details, and spatial positioning across multiple dancers in a complex one-shot music video. The model also features an intelligent duration recommendation function that suggests optimal lengths based on prompts, along with a video extension tool for continuing storylines.

Documents to Video: Turning Office Materials into Visual Content

Wan3.0's most disruptive upgrade lies in its input layer. Beyond text, images, audio, and video, the model now supports doc, xls, ppt, pdf, txt, key, pages, numbers, and md—a total of nine office document formats, with a 100MB file size limit and up to 50 pages.

Upload a product introduction PPT with a prompt, and the model automatically generates a complete promotional video. Upload a financial report or business presentation, and it transforms into a dynamic data visualization. This means training materials, product demonstrations, and business reports can now be directly "translated" into video.

Unique Characters and Consistency: Moving Beyond the "AI Plastic Look"

To address the common issues of "oily" textures and generic faces in AI video, Wan3.0 has specifically optimized character generation. In text-to-video mode, the model more accurately renders facial features and skin details, preserving real characteristics like pores and smile lines, giving each character a distinct appearance and emotional expression within the same frame.

In multi-reference tasks, Wan3.0 can simultaneously process reference information across characters, props, audio, spatial relationships, and styles, maintaining consistency throughout the generation. Character faces, hairstyles, clothing, and accessories remain stable; prop logos, materials, and structures stay consistent across angles; different styles—cinematic, documentary, etc.—remain distinct without cross-contamination. This capability significantly reduces continuity errors in multi-shot commercial advertisements and short dramas.

Pricing and Availability

Starting today, Wan3.0 is available for public beta on Alibaba Cloud Bailian, Wanjing Yike, Wanxiang's official website, Qianwen Creation for PC, as well as IF STUDIO and DuYou. It is also rolling out gradually on the Qianwen App. API pricing is per second of generated video at 0.3/0.6/1.2 RMB per second for 480P, 720P, and 1080P, respectively—approximately 40% to 50% of Seedance 2.5's cost. The API interface will be fully opened soon.

The release of Wan3.0 signals that AI video generation is transitioning from a "tool for creators" to a "productivity platform for knowledge workers." The 30-second generation capability is already close to the basic unit of TV commercials and short dramas, while document-to-video functionality opens a more commercially compelling space—enterprise training, product marketing, data reporting, and many other daily office tasks could be redefined by AI video. From 3-second clips to 30-second narratives, video generation has advanced rapidly in just 18 months. The next step might be enabling AI to understand entire scripts.

Qwen Creation Platform  https://create.qianwen.com/ 


Alibaba Releases Wan3.0 Video Model: 30-Second Generations, Documents to Video in One Click, "Everything Becomes Video" | AIYXL