Cosmos3-Super-Text2Image-4Step
Cosmos3-Super-Text2Image-4Step is a high-performance image generation model developed by NVIDIA as part of the Cosmos 3 omnimodal world foundation model platform. It is a distilled version of the larger Cosmos3-Super-Text2Image model, utilizing Improved Distribution Matching Distillation (DMD2) to achieve high-fidelity image generation in significantly fewer iterations. By reducing the required sampling from approximately 50 denoising steps to just 4, the model provides up to a 25x increase in inference speed compared to its non-distilled counterpart.
The model is built upon a Mixture-of-Transformers (MoT) architecture, which employs a unified 3D multi-dimensional rotary position embedding (mRoPE) to encode spatial and temporal structures. A key technical advancement in the 4-step version is the elimination of the need for classifier-free guidance (CFG). Since CFG typically requires two forward passes per inference step, its removal further accelerates the generation process while maintaining structural consistency and visual quality.
Optimized for Physical AI applications, Cosmos3-Super-Text2Image-4Step is designed to serve as a building block for world simulation, robotics training, and synthetic data generation. It allows developers to generate high-quality visual assets and environments that adhere to physical and causal laws. While primarily a text-to-image model, it is part of a larger ecosystem of models that share a common architecture to facilitate reasoning and action across text, images, and video modalities.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Cosmos3-Super-Text2Image-4Step ranks
Cosmos3-Super-Text2Image-4Step is highlighted in the table below. Switch the metric to see how the ordering changes.