Logocrafiq.ai

An AI-powered assets creation platform. Generate, edit & ship content faster.

Explore

  • Home
  • Contact
  • Pricing
  • Blog

Features

  • 2D Assets Generator
  • Text to 3D
  • Video Generator
  • Sound Effects
  • All Features

Rankings

  • Image generation
  • Image upscaling
  • Video generation
  • 3D generation
  • Text generation
  • Music generation
  • Speech generation

© 2026 Crafiq. All rights reserved.

Privacy PolicyTermsImpressum
Models/Image/Cosmos3-Super-Text2Image-4Step
NVIDIA logoNVIDIA·Image GenerationOpen weights

Cosmos3-Super-Text2Image-4Step

View rankingsHugging Facenvidia.com
AA Text→Image#62
Parameters64B
ReleasedJul 2026

Cosmos3-Super-Text2Image-4Step is a high-performance image generation model developed by NVIDIA as part of the Cosmos 3 omnimodal world foundation model platform. It is a distilled version of the larger Cosmos3-Super-Text2Image model, utilizing Improved Distribution Matching Distillation (DMD2) to achieve high-fidelity image generation in significantly fewer iterations. By reducing the required sampling from approximately 50 denoising steps to just 4, the model provides up to a 25x increase in inference speed compared to its non-distilled counterpart.

The model is built upon a Mixture-of-Transformers (MoT) architecture, which employs a unified 3D multi-dimensional rotary position embedding (mRoPE) to encode spatial and temporal structures. A key technical advancement in the 4-step version is the elimination of the need for classifier-free guidance (CFG). Since CFG typically requires two forward passes per inference step, its removal further accelerates the generation process while maintaining structural consistency and visual quality.

Optimized for Physical AI applications, Cosmos3-Super-Text2Image-4Step is designed to serve as a building block for world simulation, robotics training, and synthetic data generation. It allows developers to generate high-quality visual assets and environments that adhere to physical and causal laws. While primarily a text-to-image model, it is part of a larger ecosystem of models that share a common architecture to facilitate reasoning and action across text, images, and video modalities.

Create with Crafiq

Generate images, 3D models, video and audio in one studio.

Explore the studio

How Cosmos3-Super-Text2Image-4Step ranks

Cosmos3-Super-Text2Image-4Step is highlighted in the table below. Switch the metric to see how the ordering changes.