Logocrafiq.ai

An AI-powered assets creation platform. Generate, edit & ship content faster.

Explore

  • Home
  • Contact
  • Pricing
  • Blog

Features

  • 2D Assets Generator
  • Text to 3D
  • Video Generator
  • Sound Effects
  • All Features

Rankings

  • Image generation
  • Image upscaling
  • Video generation
  • 3D generation
  • Text generation
  • Music generation
  • Speech generation

© 2026 Crafiq. All rights reserved.

Privacy PolicyTermsImpressum
Models/Video/Cosmos3-Super-Image2Video-4Step
NVIDIA logoNVIDIA·Video GenerationOpen weights

Cosmos3-Super-Image2Video-4Step

View rankingsHugging Facenvidia.com
AA Image→Video#20
Parameters64B
ReleasedJul 2026

Cosmos3-Super-Image2Video-4Step is a 64-billion parameter video generation model developed by NVIDIA as part of its Cosmos 3 family of omnimodal world foundation models. Designed specifically for Physical AI, the model is engineered to simulate and interact with the physical world by generating temporally coherent video sequences from a single input image and optional text prompts. This "Super" variant is optimized for high-fidelity output and is a distilled version of the base Cosmos 3 architecture.

The model's architecture is built on a Mixture-of-Transformers design featuring two distinct components: a Reasoner Tower and a Generator Tower. The Reasoner Tower—an autoregressive vision-language model—interprets the visual and textual context to predict future world states, which the Generator Tower (a diffusion transformer) then translates into video frames. This specific 4-step variant was created using Improved Distribution Matching Distillation (DMD2). By distilling the knowledge of a 50-step teacher model into just four denoising steps, the model achieves an inference speedup of approximately 25x while eliminating the need for classifier-free guidance (CFG).

Key capabilities of the model include the generation of physics-aware videos that maintain high visual consistency with the source image, making it suitable for applications in robotics, autonomous vehicle simulation, and industrial digital twins. It is released under the NVIDIA Open Model License, which allows for both commercial use and the creation of derivative works. The model is designed to run on high-end hardware, typically requiring significant VRAM for local inference of the full 64B parameter weights.

To achieve the best results, users should provide high-resolution input images with clear subjects. Since the model is distilled to operate in four steps, motion prompts should be concise, as the model relies heavily on the reasoning capabilities of its integrated VLM to interpret the intended physical dynamics of the scene.

Create with Crafiq

Generate images, 3D models, video and audio in one studio.

Explore the studio

How Cosmos3-Super-Image2Video-4Step ranks

Cosmos3-Super-Image2Video-4Step is highlighted in the table below. Switch the metric to see how the ordering changes.