Cosmos3-Super-Image2Video-4Step
Cosmos3-Super-Image2Video-4Step is a 64-billion parameter video generation model developed by NVIDIA as part of its Cosmos 3 family of omnimodal world foundation models. Designed specifically for Physical AI, the model is engineered to simulate and interact with the physical world by generating temporally coherent video sequences from a single input image and optional text prompts. This "Super" variant is optimized for high-fidelity output and is a distilled version of the base Cosmos 3 architecture.
The model's architecture is built on a Mixture-of-Transformers design featuring two distinct components: a Reasoner Tower and a Generator Tower. The Reasoner Tower—an autoregressive vision-language model—interprets the visual and textual context to predict future world states, which the Generator Tower (a diffusion transformer) then translates into video frames. This specific 4-step variant was created using Improved Distribution Matching Distillation (DMD2). By distilling the knowledge of a 50-step teacher model into just four denoising steps, the model achieves an inference speedup of approximately 25x while eliminating the need for classifier-free guidance (CFG).
Key capabilities of the model include the generation of physics-aware videos that maintain high visual consistency with the source image, making it suitable for applications in robotics, autonomous vehicle simulation, and industrial digital twins. It is released under the NVIDIA Open Model License, which allows for both commercial use and the creation of derivative works. The model is designed to run on high-end hardware, typically requiring significant VRAM for local inference of the full 64B parameter weights.
To achieve the best results, users should provide high-resolution input images with clear subjects. Since the model is distilled to operate in four steps, motion prompts should be concise, as the model relies heavily on the reasoning capabilities of its integrated VLM to interpret the intended physical dynamics of the scene.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Cosmos3-Super-Image2Video-4Step ranks
Cosmos3-Super-Image2Video-4Step is highlighted in the table below. Switch the metric to see how the ordering changes.