Wan 3.0
Wan 3.0 (also known as Tongyi Wanxiang 3.0) is a multimodal video generation model developed by Alibaba's Tongyi Lab. Released in public beta in August 2026, it represents a significant update to the Wan series by unifying previously separate text-to-video and image-to-video pipelines into a single architecture. Its primary technical advancement is the native generation of 30-second single-shot clips, allowing for continuous camera movements and long-form cinematic storytelling without the need for stitching or temporal post-processing.
Omni-Reference and Document Integration
A central feature of Wan 3.0 is its Omni-Reference capability, which extends input modalities beyond text and images. The model can ingest and interpret documents (PDF, PPT, XLS, DOC), spreadsheets, and even web pages to transform structured data or presentations into video content. This "document-to-video" functionality is aimed at enterprise workflows, such as converting a marketing deck into a promotional video or a spreadsheet into an animated data visualization.
Technical Specifications and Audio
The model generates synchronized audio natively in the same generation pass as the visual content. This includes ambient sound, sound effects, and dialogue that matches on-screen action. Wan 3.0 supports output resolutions up to 1080p and provides "spot editing" tools that allow users to regenerate specific time intervals within a video using instruction-based prompts while maintaining character and scene consistency. It is currently available as a closed model via Alibaba Cloud's API and the Wanxiang web interface.
Helpful Prompting Tips
- Cinematic Language: The model responds accurately to director-level commands. Use terms like "unbroken push," "slow aerial pull-back," or "tracking shot" to guide the 30-second duration.
- Multi-Reference Locking: You can provide multiple reference images (up to 20) to ensure consistent character and prop appearance across long takes.
- Intelligent Duration: If you are unsure of the optimal timing for a scene, the model's "smart duration" setting can automatically determine the clip length based on the complexity of the prompt descriptions.
Run Wan 3.0 in Crafiq
Ready to use in the studio. No API keys, no setup.
Open the studioHow Wan 3.0 ranks
Wan 3.0 is highlighted in the table below. Switch the metric to see how the ordering changes.